Threat Event Detection Method, Device, Equipment and Medium Based on Traceability Graph Matching

Through the threat event detection method based on traceability map matching, using isomorphic distance minimization calculation and permutation matrix optimization, the problem of not being able to effectively identify unknown attack patterns in the prior art is solved, and high accuracy detection and accurate positioning of network threat data are achieved.

CN119945798BActive Publication Date: 2025-06-10PENG CHENG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510398821.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-10
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The prior art cannot effectively identify and detect unknown attack patterns when detecting cyber threats, resulting in reduced detection accuracy.

Method used

Using the threat event detection method based on traceability map matching, the query map constructed by network threat intelligence and the traceability map generated based on the preset kernel audit log is used to perform isomorphic distance minimization calculations, a target permutation matrix is ​​generated, and the nodes in the traceability map are rearranged through this matrix to match the query map.

Benefits of technology

It improves the accuracy of detection of network threat data, can capture the topological pattern similarity of unknown attacks in the traceability map, and accurately traces the specific occurrence nodes in the log through node mapping relationships, achieving accurate positioning of network threat data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945798B_ABST
    Figure CN119945798B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a threat event detection method, device, equipment and medium based on traceability graph matching. The method includes: obtaining a query graph and a traceability graph; inputting the query graph and the traceability graph into a target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph; through the target model, performing an isomorphic distance minimization calculation on the objective function to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix; through the target permutation matrix, permuting the second node embedding matrix into a target node embedding matrix, and determining a target isomorphic distance between the target node embedding matrix and the first node embedding matrix; when the target isomorphic distance is less than a preset warning distance threshold, based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix, outputting a threat event detection result. In this way, the accuracy of threat event detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technologies, and in particular, to a threat event detection method, apparatus, device, and medium based on traceability graph matching. Background Art

[0002] With the rapid development of information technology and the wide popularization of the Internet, network attack means have become increasingly diversified and complex. Network attacks carry out malicious behaviors against computer systems, network infrastructures, or data assets through means such as illegal intrusion, data theft, and service disruption, posing a serious threat to network security, enterprise operations, and personal privacy. Therefore, it is necessary to detect network threat data to enhance network security protection capabilities.

[0003] In related technologies, generally, a signature database of known attack patterns is maintained. Each attack signature in the signature database contains the characteristics of a specific attack, such as specific strings, packet structures, or behavior patterns, etc., and the data in network traffic or system logs is matched with these signatures to identify potential attack behaviors. However, attackers will continuously evolve their attack means, resulting in new attack patterns that may not be covered by the signature database. At this time, these unknown attack patterns cannot be identified and detected through matching, reducing the accuracy of detection. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a threat event detection method, apparatus, device, and medium based on traceability graph matching, which can improve the accuracy of detecting network threat data.

[0005] To achieve the above object, the first aspect of the embodiments of this application proposes a threat event detection method based on traceability graph matching, and the method includes:

[0006] Obtain a query graph constructed according to network threat intelligence and a traceability graph generated according to preset kernel audit logs;

[0007] Input the query graph and the traceability graph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph;

[0008] Through the target model, perform isomorphic distance minimization calculation on a preset target function to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, where the target function is used to represent the isomorphic distance relationship between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph under different permutation matrices;

[0009] Through the target permutation matrix, permute the second node embedding matrix into a target node embedding matrix aligned with the element positions in the first node embedding matrix, and determine the target isomorphism distance between the target node embedding matrix and the first node embedding matrix;

[0010] When the target isomorphism distance is less than a preset warning distance threshold, based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix, output the threat event detection result of the query graph.

[0011] Correspondingly, a second aspect of the embodiments of the present application proposes a threat event detection device based on traceability graph matching. The device includes:

[0012] An acquisition module, configured to acquire a query graph constructed according to network threat intelligence and a traceability graph generated according to preset kernel audit logs;

[0013] An input module, configured to input the query graph and the traceability graph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph;

[0014] A calculation module, configured to perform an isomorphism distance minimization calculation on a preset target function through the target model to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, where the target function is used to represent the isomorphism distance relationship between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph under different permutation matrices;

[0015] A permutation module, configured to permute the second node embedding matrix into a target node embedding matrix aligned with the element positions in the first node embedding matrix through the target permutation matrix, and determine the target isomorphism distance between the target node embedding matrix and the first node embedding matrix;

[0016] An output module, configured to output the threat event detection result of the query graph based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix when the target isomorphism distance is less than a preset warning distance threshold.

[0017] In some embodiments, the threat event detection device based on traceability graph matching further includes a construction module, configured to:

[0018] For each query node in the query graph, obtain a corresponding query neighborhood sequence;

[0019] Perform sequence matching in the traceability graph based on the query neighborhood sequence. When the query neighborhood sequence is a subsequence of any traceability node in the traceability graph, add the any traceability node as a candidate node to the candidate node set corresponding to the corresponding query node;

[0020] Perform global optimization on the candidate node set according to the relevance between each candidate node in the candidate node set and the corresponding query node, and obtain the corresponding target node set;

[0021] Construct a corresponding traceability subgraph based on the multiple target nodes included in the multiple target node sets and the connection relationships of the multiple target nodes in the traceability graph;

[0022] Then, inputting the query graph and the traceability graph into the pre-trained target model to obtain the first node embedding matrix corresponding to the query graph and the second node embedding matrix corresponding to the traceability graph includes:

[0023] Input the query graph and the traceability subgraph into the pre-trained target model to obtain the first node embedding matrix corresponding to the query graph and the second node embedding matrix corresponding to the traceability subgraph.

[0024] In some embodiments, the construction module is further configured to:

[0025] Obtain the candidate adjacent nodes of each candidate node in the candidate node set and the query adjacent nodes of the corresponding query node;

[0026] Construct a bipartite graph between each candidate node and the corresponding query node based on the candidate adjacent nodes and the query adjacent nodes;

[0027] When the subsequence of the query adjacent node is the subsequence of the candidate adjacent node, add a connection edge between the query adjacent node and the candidate adjacent node in the bipartite graph to obtain a target bipartite graph;

[0028] Perform global optimization on the candidate node set according to the number of connection edges of each query adjacent node corresponding to the query node in the target bipartite graph, and obtain the corresponding target node set.

[0029] In some embodiments, the threat event detection device based on traceability graph matching further includes a conversion module, configured to:

[0030] Obtain a preset permutation matrix variable, and perform a multiplication operation between the permutation matrix variable and the second node embedding matrix to construct a permutation adjustment term;

[0031] Taking the minimization of the difference between the first node embedding matrix and the permutation adjustment term as the constraint objective, a first initial function is established based on the difference between the first node embedding matrix and the permutation adjustment term;

[0032] Obtain the Lagrange multiplier matrix variable, and perform dual transformation on the first initial function through the Lagrange multiplier matrix variable to obtain a second initial function; wherein, the dimensions of the Lagrange multiplier matrix variable, the first node embedding matrix, and the second node embedding matrix are the same;

[0033] By calculating the Lagrange multiplier matrix variable, obtain the Lagrange multiplier matrix, and adjust the second initial function based on the Lagrange multiplier matrix to obtain the objective function.

[0034] In some embodiments, the calculation module is further configured to:

[0035] Input the first node embedding matrix and the second node embedding matrix into the bilinear activation network of the target model respectively to obtain corresponding first low-dimensional embedding representations and second low-dimensional embedding representations;

[0036] Perform an inner product operation on the first low-dimensional embedding representation and the second low-dimensional embedding representation to generate a corresponding node similarity matrix;

[0037] Through the operator network, perform differentiable continuous relaxation processing on the node similarity matrix, so as to approach the permutation matrix variable of the objective function through the soft permutation matrix, and use the soft permutation matrix as the target permutation matrix corresponding to the permutation matrix variable.

[0038] In some embodiments, the threat event detection device based on traceability graph matching further includes a training module, configured to:

[0039] Obtain a preset positive sample pair and negative sample pair, wherein each positive sample pair includes a sample traceability graph and a sample first query graph whose isomorphism distance from the sample traceability graph is less than the training warning distance threshold, and each negative sample pair includes the sample traceability graph and a sample second query graph whose isomorphism distance from the sample traceability graph is greater than the training warning distance threshold;

[0040] Through a preset model, obtain the first isomorphism distance between the positive sample pairs and the second isomorphism distance between the negative sample pairs;

[0041] Based on the difference between the first isomorphism distance and the second isomorphism distance, construct a hinge loss;

[0042] Based on the hinge loss, adjust the network parameters of the preset model to obtain the target model.

[0043] In some embodiments, the input module is further configured to:

[0044] Input the query graph and the traceability graph into a pre-trained target model;

[0045] Through each neural network layer in the target model, for each query node in the query graph, based on the adjacent query node features updated by adjacent query nodes in the previous neural network layer, update the query node features of each query node, and generate a first node embedding matrix based on the updated query node features corresponding to multiple query nodes in the last layer;

[0046] Through each neural network layer, for each traceability node in the traceability graph, based on the adjacent traceability node features updated by adjacent traceability nodes in the previous neural network layer, update the traceability node features of each traceability node, and generate a second node embedding matrix based on the updated traceability node features corresponding to multiple traceability nodes in the last layer.

[0047] Correspondingly, a third aspect of the embodiments of the present application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the threat event detection method based on traceability graph matching according to any one of the embodiments of the first aspect of the present application.

[0048] Correspondingly, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the threat event detection method based on traceability graph matching according to any one of the embodiments of the first aspect of the present application.

[0049] In an embodiment of the present application, a query graph constructed based on cyber threat intelligence and a tracing graph generated according to preset kernel audit logs are obtained; the query graph and the tracing graph are input into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the tracing graph; through the target model, an isomorphic distance minimization calculation is performed on a preset target function to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, where the target function is used to represent the isomorphic distance relationship between the second node embedding matrix of the tracing graph and the first node embedding matrix of the query graph under different permutation matrices; through the target permutation matrix, the second node embedding matrix is permuted into a target node embedding matrix with element positions aligned with those in the first node embedding matrix, and the target isomorphic distance between the target node embedding matrix and the first node embedding matrix is determined; when the target isomorphic distance is less than a preset warning distance threshold, based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix, a threat event detection result of the query graph is output. In this way, by generating a tracing graph from kernel audit logs, the dynamic changes during the system operation can be fully presented, new emerging attack behaviors can be comprehensively covered, the limitations of the static feature library are broken through, which is beneficial to improving the detection accuracy. At the same time, by introducing isomorphic distance minimization calculation and permutation matrix optimization, node-level alignment mapping is synchronously achieved when measuring the graph structure similarity. It can not only capture the topological pattern similarity of unknown attacks in the tracing graph to improve the accuracy of threat event detection, but also accurately trace the specific occurrence nodes of attacks in the logs through the node mapping relationship, realizing the accurate positioning of network threat data. In summary, the present application can improve the accuracy of threat event detection while achieving the accurate positioning of network threat data. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 FIG. 1 is a schematic architecture diagram of a threat event detection system based on tracing graph matching provided by an embodiment of the present application;

[0051] Figure 2 FIG. 2 is a flowchart of a threat event detection method based on tracing graph matching provided by an embodiment of the present application;

[0052] Figure 3 FIG. 3 is a general flowchart of a threat event detection method based on tracing graph matching provided by an embodiment of the present application;

[0053] Figure 4 FIG. 4 is a schematic diagram of functional modules of a threat event detection device based on tracing graph matching provided by an embodiment of the present application;

[0054] Figure 5 FIG. 5 is a schematic hardware structure diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0056] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0058] With the rapid development of information technology and the wide popularity of the Internet, network attack means have become increasingly diversified and complex. Through means such as illegal intrusion, data theft, and service disruption, network attacks carry out malicious acts against computer systems, network infrastructure, or data assets, posing a serious threat to network security, enterprise operations, and personal privacy. Therefore, it is necessary to detect network threat data to enhance network security protection capabilities.

[0059] In the related art, generally, a signature database of known attack patterns is maintained. Each attack signature in the signature database contains the characteristics of a specific attack, such as specific strings, packet structures, or behavior patterns, etc., and the data in network traffic or system logs is matched with these signatures to identify potential attack behaviors. However, attackers will continuously evolve their attack means, resulting in new attack patterns that may not be covered by the signature database. At this time, these unknown attack patterns cannot be identified and detected by matching, reducing the accuracy of detection.

[0060] Based on this, the embodiments of the present application provide a threat event detection method, device, equipment, and medium based on traceability graph matching, which can improve the accuracy of detecting network threat data.

[0061] The threat event detection method, device, equipment, and medium based on traceability graph matching provided by the embodiments of the present application will be specifically described through the following embodiments. First, the threat event detection system based on traceability graph matching in the embodiments of the present application will be described.

[0062] Please refer to Figure 1, in some embodiments, the embodiments of the present application provide a threat event detection system based on traceability graph matching, including a terminal 11 and a server side 12.

[0063] Exemplarily, the terminal 11 can be a personal computer, a dedicated security device (such as a network security gateway, an intrusion detection system, etc.), a mobile device, and so on. The terminal 11 can collect data such as system logs and network traffic in real time through a kernel audit module or a network interface, provide raw data for the construction of the traceability graph, and preprocess the data. At the same time, technicians can also construct a query graph on the terminal 11 according to threat intelligence, and define the characteristics and patterns of threat behaviors. The terminal 11 can also provide a graphical interface for displaying the threat event detection results corresponding to the query graph and allowing users to perform further analysis and operations.

[0064] Furthermore, the server side 12 can be deployed on a computer device. For example, the server side 12 can be a high-performance server, a distributed computing cluster, and so on. The server side 12 can generate training samples according to the traceability graph and train the parameters of a preset model. Moreover, the server side 12 can also receive the query graph input by the terminal 11, combine it with the traceability graph constructed in real time, perform inference matching through the model, and return the obtained threat event detection results to the terminal 11.

[0065] The threat event detection method based on traceability graph matching in the embodiments of the present application can be illustrated by the following embodiments.

[0066] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing that needs to be carried out according to data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0067] In the embodiments of the present application, a description will be made from the dimension of a threat event detection device based on traceability graph matching, and this threat event detection device can be specifically integrated in a computer device. Refer to Figure 2 , Figure 2The figure is a flowchart of the steps of the threat event detection method based on traceability graph matching provided by the embodiments of the present application. In the embodiments of the present application, taking the threat event detection device being specifically integrated on a terminal or a server as an example, when the processor on the terminal or the server executes the program instructions corresponding to the threat event detection method based on traceability graph matching, the specific process is as follows:

[0068] Step 101, obtain a query graph constructed according to network threat intelligence and a traceability graph generated according to preset kernel audit logs.

[0069] In some embodiments, in order to identify potential security threats in the system, key data structures for network threat hunting, namely the query graph and the traceability graph, can be obtained, so as to facilitate subsequent network threat detection and positioning by comparing the graph structures of the query graph and the traceability graph.

[0070] Among them, network threat intelligence can be information about potential or current network attacks collected and analyzed from various sources. Specifically, network threat intelligence includes but is not limited to attack techniques, malware characteristics, attacker identities, etc., which helps to predict network threats.

[0071] Among them, the query graph can be a graphical representation constructed according to network threat intelligence, used to describe the behavior patterns of potential or known threats, and reflect the causal relationships and information flows among all entities within the system. The query graph is used to search for matching items in the traceability graph and can determine the corresponding threats existing in the system.

[0072] Among them, the kernel audit log can be an activity log recording the operating system kernel level, containing real-time interaction information among all entities such as processes, files, and network connections in the system.

[0073] Among them, the traceability graph can be a directed graph converted from the kernel audit log, where nodes represent entities in the system (such as files, processes), and edges represent the causal relationships or information flows among the entities. The traceability graph can be used to characterize the causal relationships and information flows among all entities within the system.

[0074] Exemplarily, network threat intelligence can be obtained from a threat intelligence platform and a query graph can be constructed. For example, if the threat intelligence platform reports the attack pattern of an Advanced Persistent Threat (APT) organization, and the pattern involves the following steps: creating a malicious process (entity A); downloading a malicious file (entity B) through this process; establishing a network connection (entity C) to upload data back, then these entities and their relationships can be constructed into a query graph for searching for matching patterns in the traceability graph.

[0075] Exemplarily, the query graph can also be manually constructed by a security analyst or generated by an automated tool. The present application does not limit the specific manner of generating the query graph.

[0076] In some embodiments, the traceability graph can be obtained from kernel audit logs, system monitoring tools, network traffic analysis tools, and log management platforms. Taking the acquisition from kernel audit logs as an example, in the Linux system, the auditd (audit daemon) tool can be used to record kernel audit logs. The logs can include process creation events, file access events, network connection events, etc. These pieces of information can be parsed and constructed into the traceability nodes and edges of the traceability graph, and finally, the constructed traceability graph can be obtained.

[0077] By generating the query graph and the traceability graph, it can facilitate the subsequent effective detection and response to the possible complex attack patterns in network threat data.

[0078] Step 102: Input the query graph and the traceability graph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph.

[0079] In some embodiments, in order to provide basic data for subsequent graph matching and threat localization, node embedding matrices corresponding to the query graph and the traceability graph can be generated to capture the structural and feature information of each node in the graph, efficiently and accurately calculate the similarity between the query graph and the traceability graph, and then identify potential security threats.

[0080] Among them, the target model can be a pre-trained model based on the Interpretable Node-Preserving Graph Submatching Network (INPGS), which combines GraphQL (graph filtering algorithm) and a graph message passing framework, and can efficiently learn and calculate the distance value and the best matching scheme of the graph isomorphism relationship between the query graph and the traceability graph, so as to locate the query graph in the traceability graph.

[0081] Among them, the first node embedding matrix can be a node embedding matrix generated by the target model according to the query graph. It contains the feature vectors of each query node in the query graph, and these feature vectors contain the context semantic information of the nodes, such as the neighborhood structure and its features of the nodes.

[0082] Among them, the second node embedding matrix can be a node embedding matrix generated by the target model according to the traceability graph. It contains the feature vectors of each traceability node in the traceability graph, and also contains the context semantic information of the nodes.

[0083] In some embodiments, a pre-trained target model can be used to process the query graph and the traceability graph to generate corresponding node embedding matrices. Taking the query graph as an example, the target model can first initialize the embedding vectors of each node, and then in each round of iteration, in each corresponding neural network layer, for each query node, collect the adjacent query node features of its adjacent query nodes in the previous round of iteration. After receiving the adjacent query node features, the query node combines its own query node features in the previous round of iteration, and updates its own features through a state update function. After the last round of update is completed, the first node embedding matrix corresponding to the query graph is output through the last neural network layer.

[0084] Similarly, the generation method of the second node embedding matrix of the traceability graph is the same as that of the first node embedding matrix. The process of generating the first node embedding matrix according to the query graph can be referred to and will not be elaborated here.

[0085] Through the above method, the target model can effectively extract the key information in the query graph and the traceability graph, and convert it into a form convenient for calculation and analysis, laying a foundation for further threat detection.

[0086] In some embodiments, in order to provide high-quality input data for subsequent graph matching, each node in the query graph and the traceability graph can be respectively feature-updated through the neural network layer in the pre-trained target model to generate their respective first node embedding matrices and second node embedding matrices. This process utilizes the neighborhood information of the nodes to obtain rich context information and improve the accuracy of threat event detection. Exemplarily, step 102 may include:

[0087] (102.1) Input the query graph and the traceability graph into the pre-trained target model;

[0088] (102.2) Through each neural network layer in the target model, for each query node in the query graph, update the query node features of each query node based on the adjacent query node features updated by the adjacent query nodes in the previous neural network layer, and generate the first node embedding matrix based on the multiple query node features corresponding to the updated multiple query nodes in the last layer;

[0089] (102.3) Through each neural network layer, for each traceability node in the traceability graph, update the traceability node features of each traceability node based on the adjacent traceability node features updated by the adjacent traceability nodes in the previous neural network layer, and generate the second node embedding matrix based on the multiple traceability node features corresponding to the updated multiple traceability nodes in the last layer.

[0090] Among them, the neural network layer can be a hierarchical structure used to process and update node features in the target model. Each neural network layer can receive information from adjacent nodes through a message passing mechanism and update the feature vector of the current node.

[0091] Among them, the query node can be a node in the query graph that represents a specific entity or behavior pattern, and is used to characterize key elements in network threat intelligence, such as the behavior characteristics of malware, the operation steps of attackers, etc.

[0092] Among them, the adjacent query node can be other query nodes in the query graph that are directly connected to the query node whose features need to be updated currently.

[0093] Among them, the adjacent query node features can be the feature vectors of adjacent query nodes after being updated in the previous neural network layer (i.e., the previous iteration).

[0094] Among them, the query node features can be the feature vectors of the query nodes themselves in the query graph that need to be updated currently. The query node features can be continuously updated during the message passing process and finally form the corresponding node embeddings.

[0095] Among them, the tracing node can be a node in the tracing graph that represents a specific entity or behavior pattern, such as a file, a process, etc.

[0096] Among them, the adjacent tracing nodes can be other tracing nodes in the tracing graph that are directly connected to the tracing node whose features need to be updated currently.

[0097] Among them, the adjacent tracing node features can be the feature vectors of adjacent tracing nodes after being updated in the previous neural network layer (i.e., the previous iteration).

[0098] Among them, the tracing node features can be the feature vectors of the tracing nodes themselves in the tracing graph that need to be updated currently. The tracing node features can be continuously updated during the message passing process and finally form the corresponding node embeddings.

[0099] In some embodiments, the process of inputting the query graph into the target model, performing message passing and state update, and finally obtaining the first node embedding matrix is introduced. Exemplarily, after the query graph is input into the target model, each query node corresponds to an initial embedding .

[0100] For each query node, the message from its adjacent query nodes can be calculated :

[0101] ;

[0102] Among them, among them is the query node The historical query node features in the previous iteration, is the adjacent query node The adjacent query node features in the previous iteration, is the edge (that is and the edge between) the feature vector, is in the t-th iteration, the query node The message received from the adjacent query node, is the message function, used to calculate from the adjacent query node to the query node the message.

[0103] Furthermore, the target model can update the query node according to the message corresponding to the received adjacent query node features The query node features. The specific process can be described by the following formula:

[0104] ;

[0105] Among them, is the query node feature to be updated currently (that is, the node embedding corresponding to this query node), is a state update function, which can be implemented by a neural network, such as a Multi-Layer Perceptron (MLP).

[0106] Furthermore, after the update of the last neural network layer, the query node features of all query nodes in the query graph can be used to obtain the final first node embedding matrix , .

[0107] So far, the processing process of generating the first node embedding matrix of the query graph by the target model has been described. For the processing process of generating the second node embedding matrix of the traceability graph by the target model, it is the same as the processing process of the query graph, only different in the graph input to the model. The embodiments of the present application do not elaborate on the processing process of generating the second node embedding matrix, and can specifically refer to the aforementioned processing process of generating the first node embedding matrix.

[0108] In some embodiments, when the network threat intelligence indicates that there are multiple edges between any two query nodes (each edge corresponding to a different connection relationship), in the query graph, multiple edges can be added between the corresponding two query nodes (for example, one edge indicates that query node A directly invokes the kinetic energy of query node B, and the other edge indicates that query node B invokes some interfaces of query node A in reverse). Subsequently, in the process of generating the first node embedding, in each iteration, for each query node, the features of the adjacent query nodes of the adjacent query nodes are respectively passed as messages through each edge to obtain multiple messages, and different weights are set based on the messages passed through each edge. After adjusting and weighting the multiple messages passed based on the weights, the features of the query node to be updated currently are updated based on the finally obtained messages, and finally, a first node embedding matrix corresponding to multiple query nodes is obtained. In this way, the structure and behavior of the complex system can be described more precisely. In practical applications, the target model is allowed to capture more dimensional information, thereby improving the accuracy of analysis and prediction. Similarly, the above update method is also applicable to the traceability graph, and the specific process can refer to the query graph and will not be elaborated here.

[0109] By generating high-quality node embedding matrices for the query graph and the traceability graph respectively, not only the local and global context information of the nodes in the graph is fully captured, but also the semantic richness and distinctiveness of the node representations are enhanced through the deep feature extraction of the multi-layer neural network. The finally obtained node embedding matrix can more accurately reflect the relationships and characteristics of the nodes in the graph structure, thereby providing more reliable data support in subsequent tasks such as graph matching, anomaly detection, or classification, and significantly improving the accuracy and efficiency of network threat detection.

[0110] In some embodiments, in order to improve the efficiency and accuracy of threat event detection, nodes and edges irrelevant to the query in the traceability graph can be filtered and pruned, so as to construct a refined traceability subgraph closely related to the query graph, so as to reduce the interference of these nodes and edges on the detection results. Exemplarily, before step 102, it may further include:

[0111] (A.1) For each query node in the query graph, obtain the corresponding query neighborhood sequence;

[0112] (A.2) Perform sequence matching in the traceability graph based on the query neighborhood sequence. When the query neighborhood sequence is a subsequence of any traceability node in the traceability graph, add any traceability node to the candidate node set corresponding to the corresponding query node;

[0113] (A.3) According to the relevance between each candidate node in the candidate node set and the corresponding query node, globally optimize the candidate node set to obtain the corresponding target node set;

[0114] (A.4) Construct a corresponding traceability subgraph based on multiple target nodes included in multiple target node sets and the connection relationships of the multiple target nodes in the traceability graph;

[0115] Then, input the query graph and the traceability graph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph, including:

[0116] Input the query graph and the traceability subgraph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability subgraph.

[0117] Among them, the query neighborhood sequence can be a label sequence arranged in lexicographical order for the current query node itself and its jump neighborhoods (i.e., adjacent nodes within a certain distance, and the specific distance can be set according to the actual situation) in the query graph. The query neighborhood sequence can be used for matching in the traceability graph to identify potential traceability nodes that may be related to the query node.

[0118] Among them, a subsequence can be a partial sequence that appears continuously in a sequence. For example, if the query neighborhood sequence of a query node in the query graph is a subsequence of the traceability neighborhood sequence of a certain traceability node in the traceability graph, it is considered that there is similarity or relevance between the corresponding query node and the traceability node.

[0119] Among them, the candidate node set can be, for each query node in the query graph, if the query neighborhood sequence of this query node can be used as a subsequence of the traceability neighborhood sequence of a traceability node, then the set composed of these traceability nodes is the candidate node set corresponding to this query node.

[0120] Among them, a candidate node can be a node in the traceability graph that satisfies the condition that the query neighborhood sequence is a subsequence of its traceability neighborhood sequence. Each candidate node is a potential matching object for the corresponding query node.

[0121] Among them, the target node set can be a set composed of traceability nodes that are confirmed to be related to the query nodes in the query graph after global optimization.

[0122] Among them, a target node can be a single node in the target node set.

[0123] Among them, the traceability subgraph can be a subgraph constructed based on multiple target nodes and their connection relationships in the traceability graph. The traceability subgraph can only contain nodes and edges related to the query graph, removing the interference of irrelevant information, making subsequent subgraph matching and threat positioning more efficient and accurate.

[0124] Exemplarily, for each query node in the query graph , its query neighborhood sequence can be obtained. For example, for a query node , its query neighborhood sequence can be , indicating that is adjacent to . After that, in the traceability graph, for the query neighborhood sequence of each query node, it is checked whether it is a subsequence of any traceability node in the traceability graph. If so, the traceability node is added to the candidate node set corresponding to the query node as a candidate node. For example, there is a traceability node in the traceability graph. Assume that , its traceability neighborhood sequence is . The query neighborhood sequence in the query graph is 's subsequence. Therefore, can be added to the candidate node set of .

[0125] Further, in order to globally optimize the candidate node set, candidate nodes can be screened by constructing a bipartite graph between each query node and its corresponding candidate nodes to obtain multiple target nodes. Specifically, a bipartite graph can be constructed, where one side contains the query adjacent nodes of the query nodes in the query graph, and the other side contains the candidate adjacent nodes of each candidate node in the candidate node set. Next, it is detected whether the subsequence of the query adjacent node in the bipartite graph is the subsequence of the candidate adjacent node. If so, an edge is added between the corresponding query adjacent node and candidate adjacent node in the bipartite graph. Further, when all the query adjacent nodes corresponding to a query node in the bipartite graph composed of all query nodes and candidate nodes all have at least one edge, that is, when there is a semi-perfect match in the bipartite graph where each query node adjacent node is at least matched with one candidate node adjacent node, the candidate node can be determined as a target node. Otherwise, the candidate node corresponding to the bipartite graph is removed from the candidate node set, and this process can be repeated to ensure that the candidate node set is highly relevant to the query nodes. Finally, through this optimization process, the obtained target node set can more accurately reflect the corresponding relationship of the query nodes in the traceability graph, thus providing a solid foundation for constructing the traceability subgraph.

[0126] Further, the target nodes in all target node sets and their connection relationships in the traceability graph can be extracted to construct a traceability subgraph. For example, assume that there are two query nodes and in the query graph, and the corresponding target nodes are and respectively. In the traceability graph, there is an edge between and . Then the constructed traceability subgraph contains the traceability nodes and , and the edges between the two.

[0127] Through the above steps, the traceability graph can be effectively simplified, reducing the computational complexity while improving the performance of the subgraph matching network, thereby achieving more accurate detection of network threat data.

[0128] In some embodiments, to further improve the matching accuracy and efficiency of the query graph, a bipartite graph can be constructed and optimized to screen out target nodes highly relevant to the query nodes in the query graph from the candidate node set, and eliminate candidate nodes that do not exactly match the query nodes in the candidate node set, finally generating a target node set to improve the quality and reliability of graph matching and analysis. Exemplarily, (A.3) may include:

[0129] (A.3.1) Obtain the candidate adjacent nodes of each candidate node in the candidate node set, and the query adjacent nodes of the corresponding query nodes;

[0130] (A.3.2) Based on the candidate adjacent nodes and the query adjacent nodes, construct a bipartite graph for each candidate node and the corresponding query node;

[0131] (A.3.3) When the subsequence of the query adjacent node is the subsequence of the candidate adjacent node, add a connection edge between the query adjacent node and the candidate adjacent node in the bipartite graph to obtain the target bipartite graph;

[0132] (A.3.4) According to the number of connection edges of each query adjacent node corresponding to the query node in the target bipartite graph, globally optimize the candidate node set to obtain the corresponding target node set.

[0133] Among them, the candidate adjacent node can be other traceability nodes directly connected to the corresponding candidate node in the traceability graph.

[0134] Among them, the query adjacent node can be other query nodes directly connected to the query node corresponding to the candidate node in the query graph.

[0135] Among them, the bipartite graph can be a graph structure constructed in the candidate node set based on the candidate adjacent nodes of each candidate node and the query adjacent nodes of the corresponding query nodes, and each candidate node and the corresponding query node correspond to a bipartite graph.

[0136] Among them, the target bipartite graph can be an adjusted bipartite graph. In the target bipartite graph, connection edges are added only between node pairs where the subsequence of the query adjacent node matches the subsequence of the candidate adjacent node.

[0137] Exemplarily, for each query node u in the query graph, each candidate node v in its candidate node set CS(u) can be obtained, and the adjacent nodes of the query node u and each candidate node v can be obtained respectively. Taking the query node u1 and the corresponding candidate node set CS(u) as an example, for the query node u1 and the candidate node v1 in the candidate node set, a corresponding bipartite graph can be constructed. Specifically, the set N(u) composed of multiple query adjacent nodes corresponding to the query node u1, and the set N(v) composed of multiple candidate adjacent nodes of the candidate node v1 can be obtained.

[0138] Furthermore, a bipartite graph Bvu can be constructed based on N(u) and N(v). In this bipartite graph, on one side are all the query adjacent nodes in N(u), and on the other side are all the candidate adjacent nodes in N(v). And it is detected whether each query adjacent node in N(u) is a subsequence of each candidate adjacent node in N(v). If so, a corresponding edge is added between the corresponding query adjacent node and candidate adjacent node in the bipartite graph Bvu. After all the nodes in N(u) and N(v) are traversed, the target bipartite graph corresponding to the query node u1 and the candidate node v1 can be generated.

[0139] Furthermore, it can be checked whether there is a semi-perfect matching in the target bipartite graph, that is, whether all the query adjacent nodes in N(u) have at least one edge in the target bipartite graph. If not, the corresponding candidate node is removed from the candidate node set corresponding to the query node u1 and the candidate node v1; otherwise, there is no need to remove. Through multiple optimizations (the number of optimizations can be set according to the actual situation, such as 3 times, 5 times, 10 times, etc.), finally the optimal target node set can be obtained.

[0140] It can be understood that the above example is only an embodiment. In fact, the above optimization method can be adopted for the candidate node sets of any query nodes, which will not be elaborated here.

[0141] Through the above method, the target node set in the traceability graph that is most relevant to the query graph can be more accurately screened out. Thereby, not only the efficiency of graph matching is improved, but also the quality of the matching result is significantly enhanced, so that threat event detection can be more accurately performed in practical applications.

[0142] Step 103, through the target model, perform an isomorphic distance minimization calculation on the preset target function to obtain the target permutation matrix between the first node embedding matrix and the second node embedding matrix, where the target function is used to characterize the isomorphic distance relationship between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph under different permutation matrices.

[0143] In some embodiments, in order to precisely match the nodes in two graph structures to locate potential threat behaviors, the target model can be used to optimize and calculate a preset target function to find a target permutation matrix that minimizes the isomorphism distance between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph, so as to improve the accuracy and efficiency of network threat hunting.

[0144] Among them, the target function can be a mathematical expression used to measure the difference between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph.

[0145] Among them, the isomorphism distance can be a metric used to quantify the structural similarity between the traceability graph and the query graph. The isomorphism distance can be calculated by comparing the first node embedding matrix and the second node embedding matrix. A smaller isomorphism distance indicates that the traceability graph and the query graph are more similar in structure, while a larger isomorphism distance indicates a greater difference between the traceability graph and the query graph.

[0146] Among them, the target permutation matrix can be used to rearrange the nodes in the traceability graph to maximize the match with the query graph. Specifically, the target permutation matrix is a doubly stochastic matrix, which can rearrange the second node embedding matrix of the traceability graph to minimize the isomorphism distance from the node embedding matrix of the query graph. In this way, the existence of the query nodes in the query graph in the traceability graph can be determined, and when the existence is confirmed, the corresponding positions of the query nodes in the traceability nodes can be determined, so as to achieve the specific positioning of network threat behaviors.

[0147] In some embodiments, if the traceability graph and the query graph have node embedding matrices

[0148] and If there exists a permutation matrix such that , it indicates the existence of similarity between the traceability graph and the query graph. It can be understood that since the query graph is usually constructed based on known threat intelligence and represents specific attack patterns or malicious behaviors, if there is a subgraph in the traceability graph that is similar to the query graph, it can indicate that a similar attack activity has occurred in the system or there is a potential security threat. Therefore, in the process of network threat hunting, a permutation matrix that minimizes the distance between the traceability graph and the query graph should be found as much as possible to conduct threat hunting in the traceability graph as accurately as possible without missing any abnormal behaviors.

[0149] In some embodiments, the target function can be constructed through the following process:

[0150] First, construct a first initial function (i.e., distance function) :

[0151] (1)

[0152] Among them, is a permutation matrix variable, is a permutation adjustment term, is the first node embedding matrix, is the second node embedding matrix. The smaller the value of and is, the higher the correlation of the graph isomorphism relationship between . Therefore, it is necessary to solve the permutation matrix variable to maximize the structural similarity degree between the traceability graph and the query graph, facilitating matching and positioning.

[0153] Furthermore, since S is a hard permutation matrix (only allowing 0 / 1 elements), the calculation is relatively complex, and it can only represent the case of complete or incomplete matching between nodes, resulting in the inability to fully utilize the potential information in the image. Therefore, S needs to be relaxed to a doubly stochastic or "soft" permutation matrix (allowing continuous values). In some embodiments, two neural networks can be used through the target model to calculate to approximate , where the neural network with parameter can provide a differentiable solution for the optimization problem in the node alignment representation , that is, the "soft" permutation matrix, to adapt it to the node arrangement alignment requirements of different graph structures. The neural network with parameters can be used to learn the node embedding vector matrices of the traceability graph and the query graph.

[0154] In some embodiments, a neural network with parameter can be used to learn the "soft" permutation matrix, obtaining the following distance function:

[0155] (2)

[0156] Furthermore, since there is a ReLU term in formula (2), the problem to be solved is still not a linear problem. However, this problem can be regarded as a linear assignment problem in the dual space for solution. Specifically, a Lagrangian multiplier matrix variable with the same dimension as the first node embedding vector and the second node embedding vector can be used to convert the first initial function into a dual form, obtaining the following second initial function:

[0157] (3)

[0158] Among them, is the Lagrange multiplier matrix variable. Through the above formula, the original minimization problem can be transformed into a linear assignment problem in the dual space. For example, by fixing C to calculate the optimal S, and then updating C through S, an alternating optimization process is formed to avoid the bottleneck of the non-differentiability of the traditional hard permutation matrix.

[0159] Furthermore, an approximate solution of the Lagrange multiplier matrix variable C can be calculated to obtain the Lagrange multiplier matrix . Based on this, the second initial function can be adjusted based on the Lagrange multiplier matrix to obtain the objective function:

[0160] (4)

[0161] where is the Lagrange multiplier matrix. Thus, the optimization problem corresponding to the first initial function can be transformed into a linear assignment problem with a cost matrix of . This linear assignment problem can be processed by differentiable continuous relaxation through the target model, so that the target model approaches the permutation matrix variable of the objective function through the soft permutation matrix , and the soft permutation matrix is used as the target permutation matrix corresponding to the permutation matrix variable.

[0162] In some embodiments, the first node embedding matrix and the second node embedding matrix can be input into the bilinear activation network of the target model to extract high-order features, calculate the node similarity matrix through inner product, and then output a continuous permutation matrix through the operator network to approximate the optimal solution of the permutation matrix variable to obtain the target permutation matrix.

[0163] Through the above method, the best target permutation matrix between the first node embedding matrix and the second node embedding matrix can be accurately obtained, which intuitively and effectively characterizes the structural similarity between the traceability graph and the query graph under different permutations, improves the accuracy and efficiency of identifying and locating network threat behaviors, and at the same time enhances the model's ability to process complex graph structures and robustness.

[0164] In some embodiments, in order to achieve the best matching of the nodes in the traceability graph and the query graph, by constructing the first initial function and introducing methods such as Lagrange multiplier matrix variables and dual transformation, the first initial function can be transformed into a form that enables the model to perform differentiable solution quickly, so as to overcome the optimization deviation caused by the non-differentiability of the hard permutation matrix and significantly improve the accuracy of the graph isomorphism correlation metric. For example, before step 103, it may further include:

[0165] (B.1) Obtain a preset permutation matrix variable, and perform a multiplication operation on the permutation matrix variable and the second node embedding matrix to construct a permutation adjustment term;

[0166] (B.2) With the constraint objective of minimizing the difference between the first node embedding matrix and the permutation adjustment term, establish a first initial function based on the difference between the first node embedding matrix and the permutation adjustment term;

[0167] (B.3) Obtain the Lagrange multiplier matrix variable, and perform a dual transformation on the first initial function through the Lagrange multiplier matrix variable to obtain a second initial function; wherein, the dimensions of the Lagrange multiplier matrix variable, the first node embedding matrix, and the second node embedding matrix are the same;

[0168] (B.4) Calculate the Lagrange multiplier matrix through the Lagrange multiplier matrix variable, and adjust the second initial function based on the Lagrange multiplier matrix to obtain the objective function.

[0169] Among them, the permutation matrix variable can be used to rearrange the traceable nodes in the traceability graph to match the query nodes in the query graph, so as to achieve the rapid detection of network threat data.

[0170] Among them, the permutation adjustment term can be an expression constructed by multiplying a preset permutation matrix variable by the second node embedding matrix.

[0171] Among them, the constraint objective can be to minimize the difference between the first node embedding matrix and the permutation adjustment term. The constraint objective can be used to ensure that the adjusted second node embedding matrix is as close as possible to the first node embedding matrix, thereby improving the matching degree between the graph structures of the two traceability graphs and the query graph.

[0172] Among them, the first initial function can be a function established based on the difference between the first node embedding matrix and the permutation adjustment term.

[0173] Among them, the Lagrange multiplier matrix variable can be a matrix variable introduced in the optimization process, used to convert the first initial function into a dual form and handle the constraint conditions.

[0174] Among them, the second initial function can be a new function obtained by performing a dual transformation on the first initial function. This transformation is achieved by introducing the Lagrange multiplier matrix variable, aiming to simplify the optimization problem and make it easier to solve.

[0175] Among them, the Lagrange multiplier matrix can be an actual matrix obtained by calculating the Lagrange multiplier matrix variable. Each element in the Lagrange multiplier matrix is used to represent the node matching cost, and can be used to represent the weighted constraint strength of the node alignment difference in the process of solving the target permutation matrix. For example, the larger the element, the more significant the impact of the embedding difference between the query node i and the corresponding traceable node on the overall optimization.

[0176] In some embodiments, the dimensions of the Lagrange multiplier matrix variable, the first node embedding matrix, and the second node embedding matrix can be set to be the same. When the dimensions of the three are the same, their trace can characterize the global node alignment difference, so that the Lagrange multiplier can accurately characterize the node alignment cost, which helps to implement end-to-end graph matching subsequently.

[0177] In some embodiments, the objective function can be constructed through the following process:

[0178] First, construct a first initial function (i.e., the distance function) :

[0179] ;

[0180] Among them, is the permutation matrix variable, is the permutation adjustment term, is the first node embedding matrix, is the second node embedding matrix. The smaller the value of , the higher the correlation between the and the graph isomorphism relationship of . Therefore, it is necessary to solve the permutation matrix variable

[0181] to maximize the similarity degree of the structures between the traceability graph and the query graph, which is convenient for matching and positioning. Furthermore, since is a hard permutation matrix (only allowing 0 / 1 elements), the calculation is relatively complex, and it can only represent the situation of complete or incomplete matching between nodes, resulting in the inability to fully utilize the potential information in the image. Therefore, can be relaxed to a doubly stochastic or "soft" permutation matrix (allowing continuous values) for subsequent calculation. In some embodiments, through the target model, two neural networks can be used to calculate to approximate , where the neural network with parameter can provide a differentiable solution for the optimization problem in the node alignment representation

[0182] That is, the "soft" permutation matrix, so as to adapt to the node arrangement alignment requirements of different graph structures. The neural network with parameters can be used to learn the node embedding vector matrices of the traceability graph and the query graph. In some embodiments, a neural network with parameter

[0183] ;

[0184] Furthermore, since there is a ReLU term in the above formula, the problem to be solved is still not a linear problem. However, this problem can be regarded as a linear assignment problem in the dual space for solution. Specifically, a Lagrange multiplier matrix variable of the same dimension as the first node embedding vector and the second node embedding vector can be used to transform the first initial function into a dual form, obtaining the following second initial function:

[0185] ;

[0186] where, is the Lagrange multiplier matrix variable. Through the above formula, the original minimization problem can be transformed into a linear assignment problem in the dual space. For example, by fixing C to calculate the optimal S, and then updating C through S, an alternating optimization process is formed to avoid the bottleneck of the non-differentiability of the traditional hard permutation matrix.

[0187] Furthermore, an approximate solution of the Lagrange multiplier matrix variable C can be calculated to obtain the Lagrange multiplier matrix . Based on this, the second initial function can be adjusted to obtain the objective function:

[0188] ;

[0189] where, is the Lagrange multiplier matrix. Thus, the optimization problem corresponding to the first initial function can be transformed into a linear assignment problem with a cost matrix of , so as to directly perform differentiable continuous relaxation processing through the target model, thereby improving the efficiency and accuracy in the process of matching network threat data.

[0190] By constructing a function based on the permutation matrix, the structural similarity between the traceability graph and the query graph can be quantified, and by solving the permutation matrix variable, this similarity can be maximized, so as to achieve the fast and accurate matching of the traceability graph and the query graph. Moreover, to overcome the problems of complex calculation of the hard permutation matrix and only being able to represent complete or incomplete matches, the permutation matrix is relaxed to a "soft" permutation matrix allowing continuous values through an adjustment function, which is more in line with the actual topological relationship than hard alignment, and significantly improves the efficiency and accuracy in processing network threat data.

[0191] In some embodiments, to improve the accuracy and efficiency of node matching in the traceability graph and the query graph, the similarity matrix of the first node embedding matrix and the second node embedding matrix can be processed by a differentiable continuous relaxation through a target model to obtain a target permutation matrix that infinitely approximates the optimal solution of the permutation matrix variable. At the same time, the target permutation matrix is a soft matrix, which solves the problem of non-differentiability of the hard permutation matrix. In this way, the solution process can be made more efficient, and the performance of the system can also be improved. Exemplarily, step 103 may include:

[0192] (103.1) Input the first node embedding matrix and the second node embedding matrix into the bilinear activation network of the target model respectively to obtain corresponding first low-dimensional embedding representations and second low-dimensional embedding representations;

[0193] (103.2) Perform an inner product operation on the first low-dimensional embedding representation and the second low-dimensional embedding representation to generate a corresponding node similarity matrix;

[0194] (103.3) Through the operator network, perform a differentiable continuous relaxation on the node similarity matrix to approximate the permutation matrix variable of the objective function through a soft permutation matrix, and use the soft permutation matrix as the target permutation matrix corresponding to the permutation matrix variable.

[0195] Among them, the bilinear activation network can be a neural network structure for converting a high-dimensional node embedding matrix into a low-dimensional representation. The bilinear activation network can capture the complex relationships between nodes through non-linear transformations (such as the ReLU activation function) to generate more compact and expressive low-dimensional embedding representations.

[0196] Among them, the first low-dimensional embedding representation can be a low-dimensional representation obtained by processing the first node embedding matrix through the bilinear activation network. It retains the key information of the first node embedding matrix, but the dimension is significantly reduced, which is convenient for subsequent similarity calculation.

[0197] Among them, the second low-dimensional embedding representation can be a low-dimensional representation obtained by processing the second node embedding matrix through the bilinear activation network. Similar to the first low-dimensional embedding representation, the second low-dimensional embedding representation retains the key information of the second node embedding matrix, but the dimension is significantly reduced, which is convenient for subsequent similarity calculation.

[0198] Among them, the node similarity matrix can be a matrix generated by performing an inner product operation on the first low-dimensional embedding representation and the second low-dimensional embedding representation. Each element in the node similarity matrix represents the similarity between the query node and the traceability node at the corresponding position. The larger the value, the higher the similarity between the two.

[0199] Among them, the operator network can be the Gumbel-Sinkhorn operator, abbreviated as the GS operator, which is a network structure for processing the node similarity matrix and can convert a hard permutation matrix (i.e., a matrix of 0 or 1) into a soft permutation matrix (i.e., a probability distribution), thereby realizing the optimization solution of the objective function.

[0200] In some embodiments, the first node embedding matrix and the second node embedding matrix can be respectively input into a bilinear activation network (Linear-ReLU-Linear Network, LRL) with parameters Through a series of linear transformations and ReLU activation functions of the LRL, the high-dimensional node embedding matrix is converted into a low-dimensional embedding representation. Among them, the specific acquisition formula of the first low-dimensional embedding representation is as follows:

[0201] ;

[0202] Among them, is the dimensionality reduction matrix, is the feature interaction matrix, is the first node embedding matrix.

[0203] Furthermore, the specific acquisition formula of the second low-dimensional embedding representation is as follows:

[0204] ;

[0205] Among them, is the second node embedding matrix.

[0206] By converting the node embedding matrix into the corresponding low-dimensional embedding representation, the noise in the node embedding matrix can be filtered, the core features of cross-graph topological alignment can be retained, and the subsequent computational complexity can be reduced.

[0207] Furthermore, the inner product calculation can be performed on the first low-dimensional embedding representation and the second low-dimensional embedding representation to generate the node similarity matrix :

[0208] ;

[0209] The larger the element value in

[0210] indicates that the projection directions of the corresponding source node and query node in the dimensionality reduction space are more consistent, and the possibility of topological alignment is higher.

[0211] ;

[0212] Therefore, it can be approximated to , that is, the target permutation matrix for the minimization problem in the correlation metric of graph isomorphism relationship.

[0213] In some embodiments, through the above derivation, it is proved that a soft permutation matrix can be obtained through the target model to be used as the target permutation matrix to approximate the optimal solution and solve the problems of difficult matching and inaccurate matching in the hard permutation matrix (only containing elements of 0 or 1, that is, only matching and non-matching). Therefore, in practical applications, the first node embedding matrix and the second node embedding matrix can be directly input into the target model to obtain the corresponding target permutation matrix, or the feasibility of the calculation method can be verified through the above derivation method and then input into the target model to obtain the optimal solution of the objective function.

[0214] Through the above method, the complex hard permutation matrix problem can be transformed into a differentiable continuous optimization problem, avoiding the non-differentiable bottleneck of the hard permutation matrix, improving the optimization efficiency and flexibility, realizing the best matching of nodes in the traceability graph and the query graph, helping to quickly locate and identify potential network threats, and improving the efficiency and accuracy of network threat detection.

[0215] In some embodiments, in order to enhance the model's understanding and matching ability of complex network structures, the preset model can be trained to improve the accuracy and efficiency of network threat detection. Exemplarily, the target model is trained in the following way:

[0216] (C.1) Obtain the preset positive sample pairs and negative sample pairs, where each positive sample pair includes a sample traceability graph and a sample first query graph whose isomorphism distance from the sample traceability graph is less than the training warning distance threshold, and each negative sample pair includes a sample traceability graph and a sample second query graph whose isomorphism distance from the sample traceability graph is greater than the training warning distance threshold;

[0217] (C.2) Through the preset model, obtain the first isomorphism distance between positive sample pairs and the second isomorphism distance between negative sample pairs;

[0218] (C.3) Based on the difference between the first isomorphism distance and the second isomorphism distance, construct a hinge loss;

[0219] (C.4) Based on the hinge loss, adjust the network parameters of the preset model to obtain the target model.

[0220] Among them, a positive sample pair can be a sample pair composed of a sample traceability graph and a first sample query graph that has a relatively small isomorphism distance (less than the training warning distance threshold) from it. The positive sample pair represents an actual threatening situation and is used to train a preset model to recognize similar graph structures.

[0221] Among them, a negative sample pair can be a sample pair composed of the same sample traceability graph as the positive sample pair and a second sample query graph that has a relatively large isomorphism distance (greater than the training warning distance threshold) from it. The negative sample pair represents a non-threatening situation and is used to train the model to recognize dissimilar graph structures.

[0222] Among them, the sample traceability graph can be, in the positive sample pair, an actual system state graph generated from the system kernel audit log.

[0223] Among them, the training warning distance threshold can be an isomorphism distance threshold used to distinguish between positive and negative sample pairs. If the isomorphism distance between the query graph and the traceability graph is less than this training warning distance threshold, then this pair of graphs is considered a positive sample pair; otherwise, it is considered a negative sample pair.

[0224] Among them, the first sample query graph can be, in the positive sample pair, a query graph constructed based on network threat intelligence.

[0225] Among them, the second sample query graph can be, in the negative sample pair, a query graph constructed based on network threat intelligence.

[0226] Among them, the first isomorphism distance can be, in the positive sample pair, the isomorphism distance between the sample traceability graph and the first sample query graph.

[0227] Among them, the second isomorphism distance can be, in the negative sample pair, the isomorphism distance between the sample traceability graph and the second sample query graph.

[0228] Among them, the hinge loss can be a loss function for classification problems, which can optimize the model parameters by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs, prompting the preset model to learn the correct classification boundary.

[0229] Among them, the preset model can be an initial, untrained, or partially trained model based on an interpretable subgraph matching network.

[0230] Exemplarily, a positive sample pair can be composed of a sample traceability graph and a first sample query graph that satisfy the first isomorphism distance , indicating that the corresponding attack model is highly correlated with the real threat behavior (such as the C2 communication mode in an APT attack and the logs of the infected host). A negative sample pair can be composed of a sample traceability graph and a second sample query graph Composition, satisfying the second isomorphism distance , representing irrelevant or low-correlation behaviors (such as normal user operation logs). Among them, is the training warning distance threshold (such as = 0.3, which can be specifically set according to the actual situation).

[0231] ;

[0232] Among them, the subscripts and are the parameters of the bilinear activation network, which are optimized and adjusted during the training of the preset model. Finally, through the hinge loss of the hyperparameter , the overall training of the preset model is achieved:

[0233] ;

[0234] Based on the hinge loss, update the network parameters and of the model through the gradient descent method, and minimize the hinge loss, which can maximize the distance between positive and negative sample pairs and enhance the discriminative ability of the model for threat behaviors.

[0235] The obtained by learning and calculating through the preset model is the isomorphism distance of the graph isomorphism relationship, is the optimal alignment scheme. After training the target model, that is, without adjusting the network parameters of the preset model, when the isomorphism distance is lower than the set training warning threshold, the model will trigger an alarm and output the traced nodes and paths matched by the corresponding query graph in the traced graph.

[0236] By training the preset model in the above way, the model can more accurately identify and locate potential threat behaviors when processing network threat data. When the isomorphism distance is lower than the set threshold, the model will trigger an alarm and output the traced nodes and paths matched by the query graph in the traced graph, significantly improving the accuracy and efficiency of network threat detection.

[0237] Step 104, through the target permutation matrix, permute the second node embedding matrix into a target node embedding matrix aligned with the element positions in the first node embedding matrix, and determine the target isomorphism distance between the target node embedding matrix and the first node embedding matrix.

[0238] In some embodiments, to precisely match the nodes in the traceability graph and the query graph, a target permutation matrix can be applied to rearrange the second node embedding matrix so that the positions of its elements are aligned with those in the first node embedding matrix, thereby generating a target node embedding matrix. Then, the target isomorphism distance between the target node embedding matrix and the first node embedding matrix is calculated to quantify the difference between the structures of the two graphs and improve the accuracy and reliability of network threat detection.

[0239] Among them, the target node embedding matrix can be a new matrix obtained by applying the target permutation matrix to the second node embedding matrix. The target node embedding matrix adjusts the positions of the traceability nodes in the second node embedding matrix to align them as closely as possible with the node positions in the first node embedding matrix (i.e., the node embedding matrix of the query graph).

[0240] Among them, the target isomorphism distance can be used to quantify the difference between the target node embedding matrix and the first node embedding matrix. Specifically, the target isomorphism distance measures the structural similarity between the traceability graph and the query graph after the best node alignment.

[0241] Exemplarily, the target embedding matrix can be calculated through the following process:

[0242] ;

[0243] Among them, is the target permutation matrix, is the second node embedding matrix.

[0244] By adjusting the node arrangement order of the second node embedding matrix with the target permutation matrix, the topological structure of the traceability graph can be aligned with the query graph, facilitating subsequent calculation of the target isomorphism distance between the target node embedding matrix and the first node embedding matrix. For example, if the second node embedding matrix is:

[0245] ;

[0246] The target permutation matrix is:

[0247] ;

[0248] Then the target embedding matrix can be:

[0249] ;

[0250] Furthermore, the target isomorphism distance can be obtained through the difference between the first node embedding matrix and the target embedding matrix. Specifically, the target isomorphism distance The calculation formula is as follows:

[0251] ;

[0252] Among them, is the first node embedding matrix, is the target permutation matrix, is the second node embedding matrix.

[0253] Through the above method, not only the accuracy of node matching is improved, but also the ability of the model to process complex graph structures is enhanced, thereby improving the reliability, response speed and efficiency of the overall threat data hunting, enabling the target model to more accurately and quickly identify potential threat behaviors in network threat detection.

[0254] Step 105, when the target isomorphism distance is less than the preset warning distance threshold, based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix, output the threat event detection result of the query graph.

[0255] In some embodiments, in order to accurately identify and locate threat behaviors, it is possible to compare the calculated target isomorphism distance with the preset warning distance threshold to determine whether to trigger an alarm, so as to improve the accuracy and efficiency of threat detection and achieve precise positioning and rapid response to potential threat behaviors.

[0256] Among them, the preset warning distance threshold can be a pre-set distance threshold used to determine whether the similarity between the query graph and the traceability graph is sufficient to indicate the existence of potential security threats. When the calculated target isomorphism distance is lower than the preset warning distance threshold, it indicates a high similarity between the query graph and the traceability graph, suggesting that there may be network threat behaviors that need further analysis and response.

[0257] Among them, the threat event detection result can be the output information generated by the system based on the alignment relationship between the target node embedding matrix and the second node embedding matrix when the target isomorphism distance is less than the preset warning distance threshold.

[0258] In some embodiments, since when training the preset model, the can provide a differentiable solution for the optimization problem in node alignment representation, that is, the target permutation matrix approaching the hard permutation matrix. Therefore, the target permutation matrix essentially reflects the best alignment relationship between the query nodes in the query graph and the traceability nodes in the traceability graph. Through this alignment relationship, the mapping positions and connection paths of each query node in the query graph in the traceability graph can be directly output, thereby realizing the positioning of the specific location of the threat behavior in the system log.

[0259] Exemplarily, if the calculated target isomorphism distance is 0.12 and the preset warning distance threshold is 0.3, then the target isomorphism distance is less than the preset warning distance threshold, confirming that there may be a threat, and outputting the threat event detection result of the query graph in the traceability graph. Further, the preset warning distance threshold can be set according to the actual situation, for example, set to 0.1, 0.5, etc., and the embodiments of the present application do not make specific limitations on this.

[0260] In some embodiments, the threat event detection result may include the locations of the query node and the corresponding matching traceability node, the attack path (such as [malicious file] → [create] → [malicious process] → [connect] → [intranet server] → [steal] → [database file]), and the key evidence chain (such as timeline, behavior association, etc.). For example, the timeline can be 2025-10-01, 02:15:00 file download → 02:16:30 process creation → 02:20:00 lateral movement → 02:25:00 data exfiltration, and the behavior association can be that the malicious process executes the lateral movement command through the PowerShell command-line tool and establishes an encrypted tunnel with the server (185.xxx.xxx.xxx), etc.

[0261] In the embodiments of the present application, a query graph constructed based on cyber threat intelligence and a traceability graph generated based on preset kernel audit logs are obtained; the query graph and the traceability graph are input into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph; through the target model, an isomorphic distance minimization calculation is performed on a preset target function to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, where the target function is used to represent the isomorphic distance relationship between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph under different permutation matrices; through the target permutation matrix, the second node embedding matrix is permuted into a target node embedding matrix whose element positions are aligned with those in the first node embedding matrix, and the target isomorphic distance between the target node embedding matrix and the first node embedding matrix is determined; when the target isomorphic distance is less than a preset warning distance threshold, based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix, a threat event detection result of the query graph is output. In this way, by generating a traceability graph from kernel audit logs, the dynamic changes during the system operation can be fully presented, newly emerging attack behaviors can be comprehensively covered, the limitations of the static feature library are broken through, and the detection accuracy is improved. At the same time, by introducing isomorphic distance minimization calculation and permutation matrix optimization, node-level alignment mapping is synchronously achieved when measuring the graph structure similarity. It can not only capture the topological pattern similarity of unknown attacks in the traceability graph to improve the accuracy of threat event detection, but also accurately trace the specific occurrence nodes of the attack in the logs through the node mapping relationship, realizing the accurate positioning of cyber threat data. In summary, the present application can improve the accuracy of threat event detection while achieving the accurate positioning of cyber threat data.

[0262] In some embodiments, the general embodiments of the present application will be introduced in conjunction with Figure 3 Exemplarily, first, for the convenience of subsequent threat event detection, a query graph and a traceability graph can be constructed according to network logs and threat intelligence data. The query graph represents known threat patterns or attack behaviors, while the traceability graph is generated from system kernel audit logs and records the causal relationships and information flows among all entities in the system. It is necessary to perform threat event detection on the traceability graph according to the query graph to facilitate the identification of potential threats.

[0263] Furthermore, the traceability graph can be further filtered according to the query graph, and the nodes and edges in the traceability graph that are irrelevant to the query graph are pruned to obtain a traceability subgraph, so as to eliminate the interference of irrelevant data and improve the efficiency and accuracy of graph structure matching.

[0264] Furthermore, through graph representation learning techniques, the nodes of the query graph and the traceability graph (if pruned, the corresponding traceability subgraph here) can be transformed into a first node embedding matrix and a second node embedding matrix to capture the neighborhood structure and feature information of the nodes, facilitating subsequent efficient and accurate matching.

[0265] Furthermore, through the target model, a target permutation matrix can be calculated to facilitate the permutation of the second node embedding matrix so that its element positions are aligned with those in the first node embedding matrix, forming a target node embedding matrix corresponding to the second node embedding matrix. Then, the target isomorphism distance between the first node embedding matrix and the target node embedding matrix is calculated to evaluate the structural similarity between the two.

[0266] Furthermore, by calculating the target isomorphism distance, the matching degree of the query graph in the traceability graph can be quantified. When the target isomorphism distance is lower than the preset warning distance threshold, the target model triggers an alarm and outputs the nodes and paths matched by the query graph in the traceability graph (i.e., the threat event detection result), achieving precise positioning of network threats.

[0267] Please refer to Figure 4 , in some embodiments, the embodiments of the present application further provide a threat event detection device based on traceability graph matching, which can implement the above threat event detection method based on traceability graph matching. The threat event detection device based on traceability graph matching includes:

[0268] An acquisition module 41, configured to acquire a query graph constructed according to network threat intelligence and a traceability graph generated according to preset kernel audit logs;

[0269] An input module 42, configured to input the query graph and the traceability graph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph;

[0270] A calculation module 43, configured to perform an isomorphism distance minimization calculation on a preset objective function through the target model to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, where the objective function is used to represent the isomorphism distance relationship between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph under different permutation matrices;

[0271] A permutation module 44, configured to permute the second node embedding matrix into a target node embedding matrix with element positions aligned with those in the first node embedding matrix through the target permutation matrix, and determine the target isomorphism distance between the target node embedding matrix and the first node embedding matrix;

[0272] An output module 45, configured to output a threat event detection result of the query graph based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix when the target isomorphic distance is less than a preset warning distance threshold.

[0273] The specific implementation manner of the threat event detection device based on traceability graph matching is basically the same as the specific embodiments of the above-mentioned threat event detection method based on traceability graph matching, and will not be elaborated here. On the premise of meeting the requirements of the embodiments of the present application, other functional modules can be set in the threat event detection device based on traceability graph matching to implement the threat event detection method based on traceability graph matching in the above embodiments.

[0274] The embodiments of the present application further provide a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned threat event detection method based on traceability graph matching. The computer device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0275] Please refer to Figure 5 , Figure 5 which schematically shows the hardware structure of a computer device in another embodiment. The computer device includes:

[0276] A processor 51, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0277] A memory 52, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 52 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 52, and the processor 51 is used to call and execute the threat event detection method based on traceability graph matching in the embodiments of the present application;

[0278] An input / output interface 53, configured to implement information input and output;

[0279] A communication interface 54, configured to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.);

[0280] A bus 55 transmits information between various components of the device (such as a processor 51, a memory 52, an input / output interface 53, and a communication interface 54).

[0281] Among them, the processor 51, the memory 52, the input / output interface 53, and the communication interface 54 achieve communication connections with each other inside the device through the bus 55.

[0282] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned threat event detection method based on traceability graph matching.

[0283] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0284] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0285] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0286] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0287] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0288] In the description of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0289] It should be understood that in this application, "at least one (item)" and "several" mean one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0290] In several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces, devices, or units, and can be in an electrical, mechanical, or other form.

[0291] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0292] In addition, in each embodiment of the present application, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0293] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0294] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. This does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A threat event detection method based on provenance graph matching, characterized in that: The method comprises: Obtain the query graph built based on network threat intelligence and the traceability graph generated based on the preset kernel audit log; Inputting the query graph and the tracing graph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the tracing graph; Through the target model, the preset target function is subjected to isomorphism distance minimization calculation to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, wherein the target function is used to characterize the isomorphism distance relationship between the second node embedding matrix of the source graph and the first node embedding matrix of the query graph under different permutation matrices; By using the target permutation matrix, the second node embedding matrix is ​​permuted into a target node embedding matrix aligned with the element positions in the first node embedding matrix, and a target isomorphism distance between the target node embedding matrix and the first node embedding matrix is ​​determined; When the target isomorphism distance is less than a preset warning distance threshold, based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix, the threat event detection result of the query graph is output.

2. The threat event detection method based on traceability graph matching according to claim 1 is characterized in that: Before inputting the query graph and the traceability graph into the pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph, the method further includes: For each query node in the query graph, obtaining a corresponding query neighborhood sequence; Perform sequence matching in the traceability graph based on the query neighborhood sequence, and when the query neighborhood sequence is a subsequence of any traceability node in the traceability graph, add the any traceability node as a candidate node to the candidate node set of the corresponding query node; According to the correlation between each candidate node in the candidate node set and the corresponding query node, the candidate node set is globally optimized to obtain a corresponding target node set; Based on multiple target nodes included in the multiple target node sets and the connection relationship between the multiple target nodes in the traceability graph, construct a corresponding traceability subgraph; Then, the query graph and the traceability graph are input into the pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph, including: The query graph and the tracing subgraph are input into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the tracing subgraph.

3. The threat event detection method based on traceability graph matching according to claim 2 is characterized in that: The globally optimizing the candidate node set according to the correlation between each candidate node in the candidate node set and the corresponding query node to obtain the corresponding target node set includes: Obtaining candidate adjacent nodes of each candidate node in the candidate node set and query adjacent nodes of the corresponding query node; Based on the candidate adjacent nodes and the query adjacent nodes, construct a bipartite graph of each candidate node and the corresponding query node; When the subsequence of the query adjacent node is a subsequence of the candidate adjacent node, adding a connecting edge between the query adjacent node and the candidate adjacent node in the bipartite graph to obtain a target bipartite graph; According to the number of connection edges of each query adjacent node corresponding to the query node in the target bipartite graph, the candidate node set is globally optimized to obtain a corresponding target node set.

4. The threat event detection method based on traceability graph matching according to claim 1 is characterized in that: Before the target model is used to perform isomorphic distance minimization calculation on a preset target function to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, the method further includes: Obtaining a preset permutation matrix variable, and performing a product operation by the permutation matrix variable and the second node embedding matrix to construct an arrangement adjustment item; Taking minimizing the difference between the first node embedding matrix and the arrangement adjustment item as a constraint target, establishing a first initial function based on the difference between the first node embedding matrix and the arrangement adjustment item; Obtaining a Lagrange multiplier matrix variable, and performing a dual transformation on the first initial function through the Lagrange multiplier matrix variable to obtain a second initial function; wherein the Lagrange multiplier matrix variable, the first node embedding matrix, and the second node embedding matrix have the same dimension; The Lagrange multiplier matrix is ​​obtained by calculating the Lagrange multiplier matrix variables, and the second initial function is adjusted based on the Lagrange multiplier matrix to obtain the objective function.

5. The threat event detection method based on traceability graph matching according to claim 4 is characterized in that: The method of performing isomorphic distance minimization calculation on a preset objective function through the target model to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix includes: Inputting the first node embedding matrix and the second node embedding matrix into the bilinear activation network of the target model respectively to obtain a corresponding first low-dimensional embedding representation and a second low-dimensional embedding representation; Performing an inner product operation on the first low-dimensional embedding representation and the second low-dimensional embedding representation to generate a corresponding node similarity matrix; Through the operator network, the node similarity matrix is ​​subjected to differentiable continuous relaxation processing to approach the permutation matrix variable of the objective function through a soft permutation matrix, and the soft permutation matrix is ​​used as the target permutation matrix corresponding to the permutation matrix variable.

6. The threat event detection method based on traceability graph matching according to claim 1 is characterized in that: The target model is trained in the following way: Obtain preset positive sample pairs and corresponding negative sample pairs, wherein each positive sample pair includes a sample provenance graph and a sample first query graph whose isomorphic distance with the sample provenance graph is less than a training warning distance threshold, and each negative sample pair includes the sample provenance graph and a sample second query graph whose isomorphic distance with the sample provenance graph is greater than a training warning distance threshold; Obtaining a first isomorphic distance between the positive sample pairs and a second isomorphic distance between the negative sample pairs through a preset model; constructing a hinge loss based on the difference between the first isomorphism distance and the second isomorphism distance; Based on the hinge loss, the network parameters of the preset model are adjusted to obtain a target model.

7. The threat event detection method based on traceability graph matching according to claim 1 is characterized in that: The step of inputting the query graph and the traceability graph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph includes: Inputting the query graph and the traceability graph into a pre-trained target model; By each neural network layer in the target model, for each query node in the query graph, based on adjacent query node features updated by adjacent query nodes in a previous neural network layer, query node features of each query node are updated, and a first node embedding matrix is ​​generated based on a plurality of query node features corresponding to the updated query nodes of a last level; Through each neural network layer, for each tracing node in the tracing graph, the tracing node features of each tracing node are updated based on the adjacent tracing node features updated in the previous neural network layer, and based on the multiple tracing node features updated corresponding to the multiple tracing nodes of the last level, a second node embedding matrix is ​​generated.

8. A threat event detection device based on provenance graph matching, characterized in that: The device comprises: The acquisition module is used to obtain the query graph constructed based on network threat intelligence and the traceability graph generated based on the preset kernel audit log; An input module, used to input the query graph and the tracing graph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the tracing graph; A calculation module, used to perform isomorphic distance minimization calculation on a preset objective function through the target model to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, wherein the objective function is used to characterize the isomorphic distance relationship between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph under different permutation matrices; a replacement module, configured to replace the second node embedding matrix with a target node embedding matrix aligned with element positions in the first node embedding matrix through the target replacement matrix, and determine a target isomorphism distance between the target node embedding matrix and the first node embedding matrix; An output module is used to output the threat event detection result of the query graph based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix when the target isomorphism distance is less than a preset warning distance threshold.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the threat event detection method based on provenance graph matching as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the threat event detection method based on provenance graph matching according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Isomorphic subgraph query method and device, electronic device and storage medium

    CN110489607A

  • A line clearance system

    US20230351582A1