Threat event detection method and device based on traceability graph matching, equipment and medium

Through the threat event detection method based on traceability map matching, using isomorphic distance minimization calculation and permutation matrix optimization, the problem of not being able to identify unknown attack patterns in the prior art is solved, and accurate positioning and efficient detection of network threat data are achieved.

CN119945798AActive Publication Date: 2025-05-06PENG CHENG LAB

Patent Information

Application Number
CN202510398821.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-05-06
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The prior art cannot effectively identify and detect unknown attack patterns when detecting cyber threats, resulting in reduced detection accuracy.

Method used

Using the threat event detection method based on traceability map matching, the traceability map generated by network threat intelligence is obtained by obtaining the query map constructed by network threat intelligence and the traceability map generated by preset kernel audit logs, the pre-trained target model is used for isomorphic distance minimization calculation, and the target permutation matrix is ​​generated, and the traceability map nodes are embedded in the matrix permutation through this matrix to realize node-level alignment mapping.

Benefits of technology

It improves the accuracy of detection of network threat data, breaks through the limitations of the static feature library, can capture the topological pattern similarity of unknown attacks in the traceability map, and accurately traces the specific occurrence nodes of the attack in the log through node mapping relationship.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119945798A_ABST
    Figure CN119945798A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a threat event detection method and device based on traceability graph matching, equipment and a medium. The method comprises the steps of obtaining a query graph and a traceability graph; inputting the query graph and the traceability graph into a target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph; performing isomorphic distance minimization calculation on the target function through the target model to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix; replacing the second node embedding matrix with a target node embedding matrix through a target replacement matrix, and determining a target isomorphic distance between the target node embedding matrix and the first node embedding matrix; and when the target isomorphic distance is smaller than a preset early warning distance threshold, outputting a threat event detection result based on an element position alignment relationship between the target node embedding matrix and the second node embedding matrix. Therefore, the accuracy of threat event detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and in particular to a threat event detection method, device, equipment and medium based on provenance graph matching. Background Art

[0002] With the rapid development of information technology and the widespread popularity of the Internet, cyber attack methods are becoming increasingly diverse and complex. Cyber ​​attacks carry out malicious acts against computer systems, network infrastructure or data assets through illegal intrusion, data theft, service destruction and other means, posing a serious threat to network security, corporate operations and personal privacy. Therefore, it is necessary to detect network threat data to enhance network security protection capabilities.

[0003] In the related art, a signature database of known attack patterns is generally maintained. Each attack signature in the signature database contains the characteristics of a specific attack, such as a specific string, data packet structure or behavior pattern, and the data in the network traffic or system log is matched with these signatures to identify potential attack behaviors. However, attackers will continue to evolve their attack methods, resulting in new attack patterns that may not be covered by the signature database. At this time, these unknown attack patterns cannot be identified and detected by matching, which reduces the accuracy of detection. Summary of the invention

[0004] The main purpose of the embodiments of the present application is to propose a threat event detection method, device, equipment and medium based on traceability graph matching, which can improve the accuracy of detecting network threat data.

[0005] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application proposes a threat event detection method based on provenance graph matching, the method comprising: Obtain the query graph built based on network threat intelligence and the traceability graph generated based on the preset kernel audit log; Inputting the query graph and the tracing graph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the tracing graph; Through the target model, the preset target function is subjected to isomorphism distance minimization calculation to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, wherein the target function is used to characterize the isomorphism distance relationship between the second node embedding matrix of the source graph and the first node embedding matrix of the query graph under different permutation matrices; By using the target permutation matrix, the second node embedding matrix is ​​permuted into a target node embedding matrix aligned with the element positions in the first node embedding matrix, and a target isomorphism distance between the target node embedding matrix and the first node embedding matrix is ​​determined; When the target isomorphism distance is less than a preset warning distance threshold, based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix, the threat event detection result of the query graph is output.

[0006] Accordingly, a second aspect of an embodiment of the present application proposes a threat event detection device based on provenance graph matching, the device comprising: The acquisition module is used to obtain the query graph constructed based on network threat intelligence and the traceability graph generated based on the preset kernel audit log; An input module, used to input the query graph and the tracing graph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the tracing graph; A calculation module, used to perform isomorphic distance minimization calculation on a preset objective function through the target model to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, wherein the objective function is used to characterize the isomorphic distance relationship between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph under different permutation matrices; a replacement module, configured to replace the second node embedding matrix with a target node embedding matrix aligned with element positions in the first node embedding matrix through the target replacement matrix, and determine a target isomorphism distance between the target node embedding matrix and the first node embedding matrix; An output module is used to output the threat event detection result of the query graph based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix when the target isomorphism distance is less than a preset warning distance threshold.

[0007] In some implementations, the threat event detection device based on provenance graph matching further includes a construction module for: For each query node in the query graph, obtaining a corresponding query neighborhood sequence; Perform sequence matching in the traceability graph based on the query neighborhood sequence, and when the query neighborhood sequence is a subsequence of any traceability node in the traceability graph, add the any traceability node as a candidate node to the candidate node set of the corresponding query node; According to the correlation between each candidate node in the candidate node set and the corresponding query node, the candidate node set is globally optimized to obtain a corresponding target node set; Based on multiple target nodes included in the multiple target node sets and the connection relationship between the multiple target nodes in the traceability graph, construct a corresponding traceability subgraph; Then, the query graph and the traceability graph are input into the pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph, including: The query graph and the tracing subgraph are input into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the tracing subgraph.

[0008] In some embodiments, the building block is further used to: Obtaining candidate adjacent nodes of each candidate node in the candidate node set and query adjacent nodes of the corresponding query node; Based on the candidate adjacent nodes and the query adjacent nodes, construct a bipartite graph of each candidate node and the corresponding query node; When the subsequence of the query adjacent node is a subsequence of the candidate adjacent node, adding a connecting edge between the query adjacent node and the candidate adjacent node in the bipartite graph to obtain a target bipartite graph; According to the number of connection edges of each query adjacent node corresponding to the query node in the target bipartite graph, the candidate node set is globally optimized to obtain a corresponding target node set.

[0009] In some implementations, the threat event detection device based on provenance graph matching further includes a conversion module, which is used to: Obtaining a preset permutation matrix variable, and performing a product operation by the permutation matrix variable and the second node embedding matrix to construct an arrangement adjustment item; Taking minimizing the difference between the first node embedding matrix and the arrangement adjustment item as a constraint target, establishing a first initial function based on the difference between the first node embedding matrix and the arrangement adjustment item; Obtaining a Lagrange multiplier matrix variable, and performing a dual transformation on the first initial function through the Lagrange multiplier matrix variable to obtain a second initial function; wherein the Lagrange multiplier matrix variable, the first node embedding matrix, and the second node embedding matrix have the same dimension; The Lagrange multiplier matrix is ​​obtained by calculating the Lagrange multiplier matrix variables, and the second initial function is adjusted based on the Lagrange multiplier matrix to obtain the objective function.

[0010] In some embodiments, the computing module is further configured to: Inputting the first node embedding matrix and the second node embedding matrix into the bilinear activation network of the target model respectively to obtain a corresponding first low-dimensional embedding representation and a second low-dimensional embedding representation; Performing an inner product operation on the first low-dimensional embedding representation and the second low-dimensional embedding representation to generate a corresponding node similarity matrix; Through the operator network, the node similarity matrix is ​​subjected to differentiable continuous relaxation processing to approach the permutation matrix variable of the objective function through a soft permutation matrix, and the soft permutation matrix is ​​used as the target permutation matrix corresponding to the permutation matrix variable.

[0011] In some embodiments, the threat event detection device based on provenance graph matching further includes a training module for: Obtain preset positive sample pairs and negative sample pairs, wherein each positive sample pair includes a sample provenance graph and a sample first query graph whose isomorphic distance with the sample provenance graph is less than a training warning distance threshold, and each negative sample pair includes the sample provenance graph and a sample second query graph whose isomorphic distance with the sample provenance graph is greater than a training warning distance threshold; Obtaining a first isomorphic distance between the positive sample pairs and a second isomorphic distance between the negative sample pairs through a preset model; constructing a hinge loss based on the difference between the first isomorphism distance and the second isomorphism distance; Based on the hinge loss, the network parameters of the preset model are adjusted to obtain a target model.

[0012] In some embodiments, the input module is further used to: Inputting the query graph and the traceability graph into a pre-trained target model; By each neural network layer in the target model, for each query node in the query graph, based on adjacent query node features updated by adjacent query nodes in a previous neural network layer, query node features of each query node are updated, and a first node embedding matrix is ​​generated based on a plurality of query node features corresponding to the updated query nodes of a last level; Through each neural network layer, for each tracing node in the tracing graph, the tracing node features of each tracing node are updated based on the adjacent tracing node features updated in the previous neural network layer, and based on the multiple tracing node features updated corresponding to the multiple tracing nodes of the last level, a second node embedding matrix is ​​generated.

[0013] Correspondingly, the third aspect of the embodiments of the present application proposes a computer device, which includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the threat event detection method based on traceability graph matching of any one of the embodiments of the first aspect of the present application.

[0014] Correspondingly, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the threat event detection method based on traceability graph matching of any one of the embodiments of the first aspect of the present application.

[0015] The embodiment of the present application obtains a query graph constructed according to network threat intelligence, and a traceability graph generated according to a preset kernel audit log; the query graph and the traceability graph are input into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph, and a second node embedding matrix corresponding to the traceability graph; through the target model, the preset target function is minimized by isomorphic distance calculation to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, wherein the target function is used to characterize the isomorphic distance relationship between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph under different permutation matrices; through the target permutation matrix, the second node embedding matrix is ​​replaced with a target node embedding matrix aligned with the element position in the first node embedding matrix, and the target isomorphic distance between the target node embedding matrix and the first node embedding matrix is ​​determined; when the target isomorphic distance is less than the preset warning distance threshold, based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix, the threat event detection result of the query graph is output. In this way, by generating a traceability graph from the kernel audit log, the dynamic changes in the system operation process can be fully displayed, and the newly generated attack behaviors can be fully covered, breaking through the limitations of the static feature library, which is conducive to improving the accuracy of detection. At the same time, by introducing isomorphic distance minimization calculation and permutation matrix optimization, node-level alignment mapping is realized synchronously when measuring graph structure similarity, which can not only capture the topological pattern similarity of unknown attacks in the traceability graph to improve the accuracy of threat event detection, but also accurately trace the specific occurrence node of the attack in the log through the node mapping relationship, thus realizing the accurate positioning of network threat data. In summary, this application can achieve accurate positioning of network threat data while improving the accuracy of threat event detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a schematic diagram of the architecture of a threat event detection system based on provenance graph matching provided in an embodiment of the present application; Figure 2 It is a flowchart of a threat event detection method based on provenance graph matching provided in an embodiment of the present application; Figure 3 It is an overall flow chart of a threat event detection method based on provenance graph matching provided in an embodiment of the present application; Figure 4 It is a functional module diagram of a threat event detection device based on provenance graph matching provided in an embodiment of the present application; Figure 5 It is a schematic diagram of the hardware structure of the computer device provided in the embodiment of the present application. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0018] It should be noted that, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0020] With the rapid development of information technology and the widespread popularity of the Internet, cyber attack methods are becoming increasingly diverse and complex. Cyber ​​attacks carry out malicious acts against computer systems, network infrastructure or data assets through illegal intrusion, data theft, service destruction and other means, posing a serious threat to network security, corporate operations and personal privacy. Therefore, it is necessary to detect network threat data to enhance network security protection capabilities.

[0021] In the related art, a signature database of known attack patterns is generally maintained. Each attack signature in the signature database contains the characteristics of a specific attack, such as a specific string, data packet structure or behavior pattern, and the data in the network traffic or system log is matched with these signatures to identify potential attack behaviors. However, attackers will continue to evolve their attack methods, resulting in new attack patterns that may not be covered by the signature database. At this time, these unknown attack patterns cannot be identified and detected by matching, which reduces the accuracy of detection.

[0022] Based on this, the embodiments of the present application provide a threat event detection method, apparatus, device and medium based on provenance graph matching, which can improve the accuracy of detecting network threat data.

[0023] The threat event detection method, device, equipment and medium based on provenance graph matching provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the threat event detection system based on provenance graph matching in the embodiments of the present application is described.

[0024] Please refer to Figure 1 In some implementations, an embodiment of the present application provides a threat event detection system based on provenance graph matching, including a terminal 11 and a server 12.

[0025] Exemplarily, the terminal 11 may be a personal computer, a dedicated security device (such as a network security gateway, an intrusion detection system, etc.), a mobile device, etc. The terminal 11 may collect system logs, network traffic and other data in real time through a kernel audit module or a network interface, provide raw data for the construction of a traceability graph, and pre-process the data. At the same time, the technician may also construct a query graph based on threat intelligence on the terminal 11 to define the characteristics and patterns of threat behaviors. The terminal 11 may also provide a graphical interface for displaying the threat event detection results corresponding to the query graph, and allow the user to perform further analysis and operations.

[0026] Furthermore, the server 12 can be deployed on a computer device, for example, the server 12 can be a high-performance server, a distributed computing cluster, etc. The server 12 can generate training samples according to the traceability graph and train the parameters of the preset model. In addition, the server 12 can also receive the query graph input by the terminal 11, combine the traceability graph constructed in real time, perform reasoning and matching through the model, and return the threat event detection results obtained to the terminal 11.

[0027] The threat event detection method based on provenance graph matching in the embodiments of the present application can be illustrated by the following embodiments.

[0028] It should be noted that in each specific implementation of the present application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for enabling the normal operation of the embodiment of the present application will be obtained.

[0029] In the embodiment of the present application, the threat event detection device based on the traceability graph matching will be described from the perspective of the threat event detection device, which can be integrated into a computer device. Figure 2 , Figure 2 The flowchart of the steps of the threat event detection method based on traceability graph matching provided in the embodiment of the present application is as follows: Step 101, obtain a query graph constructed based on network threat intelligence and a traceability graph generated based on a preset kernel audit log.

[0030] In some embodiments, in order to identify possible security threats in the system, key data structures for network threat hunting, namely, query graphs and tracing graphs, can be obtained to facilitate subsequent network threat detection and location by comparing the graph structures of the query graphs and tracing graphs.

[0031] Cyber ​​threat intelligence can be information collected and analyzed from various sources about potential or current cyber attacks. Specifically, cyber threat intelligence includes but is not limited to attack methods, malware characteristics, attacker identities, etc., which helps predict cyber threats.

[0032] The query graph can be a graphical representation constructed based on network threat intelligence, which is used to describe the behavior patterns of potential or known threats and reflect the causal relationship and information flow between all entities within the system. The query graph is used to search for matching items in the traceability graph, and the corresponding threats existing in the system can be determined.

[0033] The kernel audit log may be an activity log recording the kernel level of the operating system, and may include real-time interaction information between all processes, files, network connections and other entities in the system.

[0034] The provenance graph can be a directed graph converted from the kernel audit log, where nodes represent entities in the system (such as files and processes) and edges represent causal relationships or information flows between entities. The provenance graph can be used to characterize the causal relationships and information flows between all entities in the system.

[0035] For example, network threat intelligence can be obtained from a threat intelligence platform and a query graph can be constructed. For example, if a threat intelligence platform reports an attack pattern of an Advanced Persistent Threat (APT) organization, which involves the following steps: creating a malicious process (entity A); downloading malicious files through the process (entity B); establishing a network connection (entity C) to transmit data back, then these entities and their relationships can be constructed into a query graph to search for matching patterns in the traceability graph.

[0036] Exemplarily, the query graph may also be constructed manually by a security analyst, or generated by an automated tool, and this application does not limit the specific method of generating the query graph.

[0037] In some implementations, the traceability graph can be obtained from kernel audit logs, system monitoring tools, network traffic analysis tools, and log management platforms. Taking the acquisition from kernel audit logs as an example, in a Linux system, the auditd (audit daemon) tool can be used to record kernel audit logs. The logs may include process creation events, file access events, network connection events, etc. This information can be parsed and constructed into traceability nodes and edges of the traceability graph, and finally a constructed traceability graph is obtained.

[0038] By generating query graphs and traceability graphs, it is possible to facilitate the subsequent effective detection and response to complex attack patterns that may exist in network threat data.

[0039] Step 102, input the query graph and the traceability graph into the pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph.

[0040] In some embodiments, in order to provide basic data for subsequent graph matching and threat location, a node embedding matrix corresponding to the query graph and the traceability graph can be generated to capture the structure and feature information of each node in the graph, and efficiently and accurately calculate the similarity between the query graph and the traceability graph, thereby identifying potential security threats.

[0041] Among them, the target model can be a pre-trained model based on the Interpretable Node-Preserving Graph Submatching Network (INPGS), which combines GraphQL (graph filtering algorithm) and graph messaging framework, and can efficiently learn and calculate the distance value and the best matching solution of the graph isomorphism relationship between the query graph and the traceability graph, so as to locate the query graph in the traceability graph.

[0042] The first node embedding matrix may be a node embedding matrix generated by the target model according to the query graph, which includes the feature vectors of each query node in the query graph, and these feature vectors include the contextual semantic information of the node, such as the neighborhood structure and features of the node.

[0043] The second node embedding matrix may be a node embedding matrix generated by the target model according to the source graph, which includes the feature vector of each source node in the source graph and also includes the contextual semantic information of the node.

[0044] In some embodiments, the query graph and the traceability graph can be processed using a pre-trained target model to generate a corresponding node embedding matrix. Taking the query graph as an example, the target model can first initialize the embedding vector of each node, and then in each round of iteration, in each corresponding neural network layer, for each query node, collect the adjacent query node features of the adjacent query nodes in the previous round of iteration. After the query node receives the adjacent query node features, it combines its own query node features in the previous round of iteration, updates its own features through a state update function, and after the last round of update is completed, outputs the first node embedding matrix corresponding to the query graph through the last neural network layer.

[0045] Similarly, the method for generating the second node embedding matrix of the traceability graph is the same as the method for generating the first node embedding matrix. You can refer to the process of generating the first node embedding matrix according to the query graph, which will not be described in detail here.

[0046] Through the above methods, the target model can effectively extract key information from the query graph and traceability graph, and convert it into a form that is easy to calculate and analyze, laying the foundation for further threat detection.

[0047] In some embodiments, in order to provide high-quality input data for subsequent graph matching, the neural network layer in the pre-trained target model can be used to update the features of each node in the query graph and the traceability graph, respectively, to generate the respective first node embedding matrix and second node embedding matrix. This process utilizes the neighborhood information of the node to obtain rich contextual information and improve the accuracy of threat event detection. Exemplarily, step 102 may include: (102.1) Input the query graph and the traceability graph into the pre-trained target model; (102.2) for each query node in the query graph, updating query node features of each query node based on adjacent query node features updated in a previous neural network layer, and generating a first node embedding matrix based on a plurality of query node features updated corresponding to a plurality of query nodes in a last level, through each neural network layer in the target model; (102.3) Through each neural network layer, for each traceability node in the traceability graph, the traceability node features of each traceability node are updated based on the adjacent traceability node features updated in the previous neural network layer, and based on the multiple traceability node features updated corresponding to the multiple traceability nodes of the last level, a second node embedding matrix is ​​generated.

[0048] The neural network layer may be a hierarchical structure used to process and update node features in the target model. Each neural network layer may receive information from adjacent nodes through a message passing mechanism and update the feature vector of the current node.

[0049] Among them, the query node can be a node in the query graph that represents a specific entity or behavior pattern, which is used to characterize key elements in network threat intelligence, such as the behavior characteristics of malware, the operation steps of attackers, etc.

[0050] The adjacent query nodes may be other query nodes in the query graph that are directly connected to the query node that currently needs to perform feature update.

[0051] The adjacent query node features may be feature vectors of the adjacent query nodes that have been updated in the previous neural network layer (i.e., the previous iteration).

[0052] The query node feature may be a feature vector of the query node itself that currently needs to be updated in the query graph. The query node feature may be continuously updated during the message transmission process to eventually form a corresponding node embedding.

[0053] Among them, the traceability node can be a node in the traceability graph that represents a specific entity or behavior pattern, such as a file, process, etc.

[0054] Among them, the adjacent traceability nodes can be other traceability nodes in the traceability graph that are directly connected to the traceability node that currently needs to perform feature update.

[0055] Among them, the adjacent tracing node features can be the feature vectors of the adjacent tracing nodes after being updated in the previous neural network layer (that is, the previous iteration).

[0056] Among them, the traceability node feature can be the feature vector of the traceability node itself that currently needs to update its features in the traceability graph. The traceability node feature can be continuously updated with the message transmission process to finally form the corresponding node embedding.

[0057] In some embodiments, the process of inputting the query graph into the target model, performing message transmission and state update, and finally obtaining the first node embedding matrix is ​​introduced. Exemplarily, after the query graph is input into the target model, each query node Corresponding to an initial embedding .

[0058] For each query node, the messages from its neighboring query nodes can be calculated : ; Among them, Is a query node Historical query node features in the previous iteration, is the adjacent query node The features of the adjacent query nodes in the previous iteration, It is the edge (i.e. and The feature vectors of the edges between In the tth iteration, query node Messages received from neighboring query nodes, is the message function used to calculate the information from the adjacent query nodes To the query node news.

[0059] Furthermore, the target model can update the query node according to the received message corresponding to the feature of the adjacent query node. The specific process can be described by the following formula: ; in, is the query node feature that needs to be updated currently (that is, the node embedding corresponding to the query node), is a state update function that can be implemented by a neural network, such as a Multi-Layer Perceptron (MLP).

[0060] Furthermore, after the last neural network layer is updated, the query node features of all query nodes in the query graph can be Get the final first node embedding matrix , .

[0061] So far, the process of generating the first node embedding matrix of the query graph through the target model has been described. The process of generating the second node embedding matrix of the traceability graph through the target model is the same as the process of the query graph, and only differs in the graph of the input model. The embodiment of the present application does not elaborate on the process of generating the second node embedding matrix. For details, please refer to the aforementioned process of generating the first node embedding matrix.

[0062] In some embodiments, when the network threat intelligence characterizes that there are multiple edges between any two query nodes (each edge corresponds to a different connection relationship), multiple edges can be added between the corresponding two query nodes in the query graph (for example, one edge indicates that query node A directly calls the kinetic energy of query node B, and the other edge indicates that query node B reversely calls part of the interface of query node A). Afterwards, in the process of generating the first node embedding, in each round of iteration, for each query node, the adjacent query node features of the adjacent query nodes can be respectively transmitted through each edge to obtain multiple messages, and different weights are set based on the messages transmitted by each edge. After adjusting and weighting the multiple messages transmitted according to the weights, the query node features that need to be updated are updated based on the final message, and finally the first node embedding matrix corresponding to multiple query nodes is obtained. In this way, the structure and behavior of complex systems can be described more accurately. In practical applications, the target model is allowed to capture more dimensional information, thereby improving the accuracy of analysis and prediction. Similarly, the above update method is also applicable to the traceability graph. The specific process can refer to the query graph, which will not be repeated here.

[0063] By generating high-quality node embedding matrices for the query graph and the traceability graph respectively, not only the local and global contextual information of the nodes in the graph is fully captured, but also the semantic richness and discrimination of the node representation are enhanced through the deep feature extraction of the multi-layer neural network. The final node embedding matrix can more accurately reflect the relationship and characteristics of the nodes in the graph structure, thereby providing more reliable data support in subsequent tasks such as graph matching, anomaly detection or classification, and significantly improving the accuracy and efficiency of network threat detection.

[0064] In some implementations, in order to improve the efficiency and accuracy of threat event detection, the nodes and edges in the provenance graph that are not related to the query may be filtered and pruned, thereby constructing a streamlined provenance subgraph that is closely related to the query graph to reduce the interference of these nodes and edges on the detection results. Exemplarily, before step 102, the following may also be included: (A.1) For each query node in the query graph, obtain the corresponding query neighborhood sequence; (A.2) Perform sequence matching in the traceability graph based on the query neighborhood sequence. When the query neighborhood sequence is a subsequence of any traceability node in the traceability graph, add any traceability node as a candidate node to the candidate node set of the corresponding query node; (A.3) According to the correlation between each candidate node in the candidate node set and the corresponding query node, the candidate node set is globally optimized to obtain the corresponding target node set; (A.4) Based on the multiple target nodes contained in the multiple target node sets and the connection relationship between the multiple target nodes in the traceability graph, construct a corresponding traceability subgraph; The query graph and the traceability graph are input into the pre-trained target model to obtain the first node embedding matrix corresponding to the query graph and the second node embedding matrix corresponding to the traceability graph, including: The query graph and the tracing subgraph are input into the pre-trained target model to obtain the first node embedding matrix corresponding to the query graph and the second node embedding matrix corresponding to the tracing subgraph.

[0065] The query neighborhood sequence can be a label sequence arranged in lexicographic order for the current query node itself and its hop neighborhood (i.e., adjacent nodes within a certain distance, the specific distance can be set according to actual conditions) in the query graph. The query neighborhood sequence can be used for matching in the provenance graph to identify potential provenance nodes that may be related to the query node.

[0066] Among them, a subsequence can be a partial sequence that appears continuously in a sequence. For example, if a query neighborhood sequence of a query node in the query graph is a subsequence of a traceability neighborhood sequence of a traceability node in the traceability graph, then it is considered that the corresponding query node and the traceability node have similarity or correlation.

[0067] Among them, the candidate node set can be for each query node in the query graph. If the query neighborhood sequence of the query node can be used as a subsequence of the tracing neighborhood sequence of the tracing node, then the set composed of these tracing nodes is the candidate node set corresponding to the query node.

[0068] The candidate node may be a node in the traceability graph that satisfies the condition that the query neighborhood sequence is a subsequence of its traceability neighborhood sequence. Each candidate node is a potential matching object of the corresponding query node.

[0069] The target node set may be a set of traceability nodes that are confirmed to be related to the query nodes in the query graph after global optimization.

[0070] The target node may be a single node in the target node set.

[0071] Among them, the provenance subgraph can be a subgraph constructed based on multiple target nodes and their connection relationships in the provenance graph. The provenance subgraph can only contain nodes and edges related to the query graph, removing the interference of irrelevant information, making subsequent subgraph matching and threat positioning more efficient and accurate.

[0072] Exemplarily, for each query node in the query graph , you can get its query neighborhood sequence, for example, query node The query neighborhood sequence can be ,express and After that, for each query node in the traceability graph, check whether it is a subsequence of any traceability node in the traceability graph. If so, add the traceability node as a candidate node to the candidate node set of the corresponding query node. , assuming The traceability neighborhood sequence of . The query neighborhood sequence in the query graph yes Therefore, we can Add to in the candidate node set.

[0073] Furthermore, in order to globally optimize the candidate node set, the candidate nodes can be screened by constructing a bipartite graph between each query node and the corresponding candidate node to obtain multiple target nodes. Specifically, a bipartite graph can be constructed, one side of which contains the query adjacent nodes of the query node in the query graph, and the other side contains the candidate adjacent nodes of each candidate node in the candidate node set. Next, it is detected whether the subsequence of the query adjacent nodes in the bipartite graph is a subsequence of the candidate adjacent nodes. If so, an edge is added between the query adjacent nodes and the candidate adjacent nodes corresponding to the bipartite graph. Further, when all query adjacent nodes corresponding to the query node have at least one edge in the bipartite graph composed of all query nodes and candidate nodes, that is, when there is a semi-perfect match in the bipartite graph in which each query node adjacent node matches at least one candidate node adjacent node, the candidate node can be determined as the target node. Otherwise, the candidate node corresponding to the bipartite graph is removed from the candidate node set, and this process can be repeated to ensure that the candidate node set is highly correlated with the query node. Ultimately, through this optimization process, the obtained target node set can more accurately reflect the corresponding relationship of the query nodes in the traceability graph, thus providing a solid foundation for constructing the traceability subgraph.

[0074] Furthermore, the target nodes in the set of all target nodes and their connection relationships in the traceability graph can be extracted to construct a traceability subgraph. For example, suppose there are two query nodes in the query graph and , the corresponding target nodes are and , in the traceability diagram, and If there is an edge between them, then the constructed traceability subgraph contains the traceability node and , and the edges between them.

[0075] Through the above steps, the traceability graph can be effectively simplified, reducing the computational complexity while improving the performance of the subgraph matching network, thereby achieving more accurate detection of network threat data.

[0076] In some implementations, in order to further improve the matching accuracy and efficiency of the query graph, a bipartite graph may be constructed and optimized to filter out target nodes that are highly relevant to the query nodes in the query graph from the candidate node set, remove candidate nodes that do not completely match the query nodes in the candidate node set, and finally generate a target node set to improve the quality and reliability of graph matching and analysis. Exemplarily, (A.3) may include: (A.3.1) Obtain the candidate adjacent nodes of each candidate node in the candidate node set, and the query adjacent nodes of the corresponding query node; (A.3.2) Based on the candidate adjacent nodes and the query adjacent nodes, construct a bipartite graph of each candidate node and the corresponding query node; (A.3.3) When the subsequence of the query adjacent node is a subsequence of the candidate adjacent node, add the edge connecting the query adjacent node and the candidate adjacent node to the bipartite graph to obtain the target bipartite graph; (A.3.4) Based on the number of connecting edges of each query adjacent node corresponding to the query node in the target bipartite graph, the candidate node set is globally optimized to obtain the corresponding target node set.

[0077] The candidate adjacent nodes may be other traceability nodes directly connected to the corresponding candidate nodes in the traceability graph.

[0078] The query adjacent nodes may be other query nodes in the query graph that are directly connected to the query node corresponding to the candidate node.

[0079] The bipartite graph may be a graph structure constructed based on the candidate adjacent nodes of each candidate node and the query adjacent nodes of the corresponding query node in the candidate node set, and each candidate node and the corresponding query node correspond to a bipartite graph.

[0080] The target bipartite graph may be an adjusted bipartite graph, in which connecting edges are added only between node pairs whose subsequences of query adjacent nodes match the subsequences of candidate adjacent nodes.

[0081] Exemplarily, for each query node u in the query graph, each candidate node v in its candidate node set CS(u) can be obtained, and the adjacent nodes of the query node u and each candidate node v can be obtained respectively. Taking the query node u1 and the corresponding candidate node set CS(u) as an example, for the query node u1 and the candidate node v1 in the candidate node set, a corresponding bipartite graph can be constructed. Specifically, the set N(u) consisting of multiple query adjacent nodes corresponding to the query node u1 and the set N(v) consisting of multiple candidate adjacent nodes of the candidate node v1 can be obtained.

[0082] Furthermore, a bipartite graph Bvu can be constructed based on N(u) and N(v). In this bipartite graph, one side is all query adjacent nodes in N(u), and the other side is all candidate adjacent nodes in N(v). And check whether each query adjacent node in N(u) is a subsequence of each candidate adjacent node in N(v). If so, add corresponding edges between the corresponding query adjacent nodes and candidate adjacent nodes in the bipartite graph Bvu. After all nodes in N(u) and N(v) are traversed, the target bipartite graph corresponding to query node u1 and candidate node v1 can be generated.

[0083] Furthermore, it is possible to check whether there is a semi-perfect match in the target bipartite graph, that is, whether all query adjacent nodes in N(u) have at least one edge in the target bipartite graph. If not, remove the corresponding candidate node from the candidate node set corresponding to the query node u1 and the candidate node v1; otherwise, it is not necessary to remove it. Through multiple optimizations (the number of optimizations can be set according to actual conditions, such as 3 times, 5 times, 10 times, etc.), the optimal target node set can be finally obtained.

[0084] It can be understood that the above examples are merely embodiments. In fact, the above optimization method can be used for any candidate node set of a query node, which will not be described in detail here.

[0085] Through the above method, the target node set in the traceability graph that is most relevant to the query graph can be screened more accurately. In this way, not only the efficiency of graph matching is improved, but also the quality of matching results is significantly improved, so that threat event detection can be performed more accurately in practical applications.

[0086] Step 103, through the target model, the preset target function is calculated to minimize the isomorphism distance, and the target permutation matrix between the first node embedding matrix and the second node embedding matrix is ​​obtained, wherein the target function is used to characterize the isomorphism distance relationship between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph under different permutation matrices.

[0087] In some embodiments, in order to accurately match the nodes in the two graph structures and thus locate potential threat behaviors, the preset objective function can be optimized and calculated through the target model to find the target permutation matrix that minimizes the isomorphic distance between the second node embedding matrix of the tracing graph and the first node embedding matrix of the query graph, so as to improve the accuracy and efficiency of network threat hunting.

[0088] The objective function may be a mathematical expression for measuring the difference between the second node embedding matrix of the source graph and the first node embedding matrix of the query graph.

[0089] Among them, the isomorphism distance can be a metric for quantifying the structural similarity between the traceability graph and the query graph. The isomorphism distance can be calculated by comparing the first node embedding matrix and the second node embedding matrix. A smaller isomorphism distance indicates that the traceability graph and the query graph are more similar in structure, while a larger isomorphism distance indicates that the traceability graph and the query graph are more different.

[0090] Among them, the target permutation matrix can be used to rearrange the nodes in the provenance graph to match the query graph to the greatest extent. Specifically, the target permutation matrix is ​​a double random matrix that can rearrange the second node embedding matrix of the provenance graph so as to minimize the isomorphic distance between the node embedding matrix of the query graph and the query graph. In this way, the existence of the query node in the query graph in the provenance graph can be determined, and if the existence is, the corresponding position of the query node in the provenance node can be determined, thereby achieving the specific positioning of the network threat behavior.

[0091] In some embodiments, if the traceability graph and query graph The node embedding matrices are as well as If there is a permutation matrix , so that , it indicates that there is a similarity between the provenance graph and the query graph. It is understandable that since the query graph is usually constructed based on known threat intelligence and represents a specific attack mode or malicious behavior, if there is a subgraph similar to the query graph in the provenance graph, it can indicate that similar attack activities have occurred in the system or there are potential security threats. Therefore, in the process of network threat hunting, the permutation matrix that minimizes the distance between the provenance graph and the query graph should be found as much as possible, so as to conduct threat hunting in the provenance graph as accurately as possible without missing any abnormal behavior.

[0092] In some implementations, the objective function may be constructed by the following process: First, construct a first initial function (also known as the distance function) : (1) in, is the permutation matrix variable, To arrange the adjustments, is the first node embedding matrix, Embedding matrix for the second node. The smaller the value of and The higher the correlation of the graph isomorphism relationship, the more necessary it is to permute the matrix variables The solution is performed to maximize the structural similarity between the source graph and the query graph, which facilitates matching and positioning.

[0093] Furthermore, since S is a hard permutation matrix (only 0 / 1 elements are allowed), the calculation is relatively complex and can only represent the situation where the nodes are completely matched or not completely matched, resulting in the inability to fully utilize the potential information in the image. Therefore, it is necessary to relax S into a double random or "soft" permutation matrix (allowing continuous values). In some embodiments, the target model can be used to calculate the two neural networks To approximate , where the parameter The neural network can provide a differentiable solution to the optimization problem in node alignment representation. , which is a "soft" permutation matrix, so that it can adapt to the node arrangement and alignment requirements of different graph structures, with The neural network with parameters of can be used to learn the node embedding vector matrices of the tracing graph and the query graph.

[0094] In some implementations, a Neural Network To learn a "soft" permutation matrix, we get the following distance function: (2) Furthermore, due to the presence of the ReLU term in equation (2), the problem to be solved is still not a linear problem, but it can be considered as a linear assignment problem in the dual space for solution. Specifically, the first initial function can be converted to a dual form by using the Lagrange multiplier matrix variable with the same dimension as the first node embedding vector and the second node embedding vector, and the following second initial function is obtained: (3) in, is the Lagrange multiplier matrix variable. Through the above formula, the original minimization problem can be converted into a linear assignment problem in the dual space. For example, by fixing C to calculate the optimal S, and then updating C through S, an alternating optimization process is formed to avoid the bottleneck of the non-differentiable traditional hard permutation matrix.

[0095] Furthermore, the approximate solution of the Lagrange multiplier matrix variable C can be calculated to obtain the Lagrange multiplier matrix Therefore, the second initial function can be adjusted based on the Lagrange multiplier matrix to obtain the objective function: (4) in, is the Lagrange multiplier matrix. Thus, the optimization problem corresponding to the first initial function can be converted into a cost matrix as The linear assignment problem can be processed by differentiable continuous relaxation of the target model so that the target model can be solved by soft permutation matrix The permutation matrix variable of the objective function is approached, and the soft permutation matrix is ​​used as the target permutation matrix corresponding to the permutation matrix variable.

[0096] In some embodiments, the first node embedding matrix and the second node embedding matrix can be input into the bilinear activation network of the target model to extract high-order features, the node similarity matrix is ​​calculated by inner product, and then the continuous permutation matrix is ​​output through the operator network to approximate the optimal solution of the permutation matrix variables to obtain the target permutation matrix.

[0097] Through the above method, the optimal target permutation matrix between the first node embedding matrix and the second node embedding matrix can be accurately obtained, thereby intuitively and effectively representing the structural similarity between the traceability graph and the query graph under different arrangements, improving the accuracy and efficiency of identifying and locating network threat behaviors, and enhancing the model's ability to handle complex graph structures and robustness.

[0098] In some implementations, in order to achieve the best match between the nodes in the source graph and the query graph, the first initial function can be constructed and the Lagrange multiplier matrix variables and dual transformation methods can be introduced to convert the first initial function into a form that enables the model to be quickly differentiable, so as to overcome the optimization deviation caused by the non-differentiable hard permutation matrix and significantly improve the accuracy of the graph isomorphism correlation measurement. For example, before step 103, it can also include: (B.1) Obtain a preset permutation matrix variable, and perform a product operation by the permutation matrix variable and the second node embedding matrix to construct a permutation adjustment term; (B.2) taking minimizing the difference between the first node embedding matrix and the permutation adjustment item as a constraint target, establishing a first initial function based on the difference between the first node embedding matrix and the permutation adjustment item; (B.3) obtaining a Lagrange multiplier matrix variable, and performing a dual transformation on the first initial function through the Lagrange multiplier matrix variable to obtain a second initial function; wherein the Lagrange multiplier matrix variable, the first node embedding matrix, and the second node embedding matrix have the same dimension; (B.4) The Lagrange multiplier matrix is ​​obtained by calculating the Lagrange multiplier matrix variables, and the second initial function is adjusted based on the Lagrange multiplier matrix to obtain the objective function.

[0099] Among them, the permutation matrix variables can be used to rearrange the traceability nodes in the traceability graph to match the query nodes in the query graph, thereby achieving rapid detection of network threat data.

[0100] The arrangement adjustment term may be an expression constructed by performing a product operation on a preset permutation matrix variable and a second node embedding matrix.

[0101] The constraint target may be to minimize the difference between the first node embedding matrix and the permutation adjustment item. The constraint target may be used to ensure that the adjusted second node embedding matrix is ​​as close as possible to the first node embedding matrix, thereby improving the matching degree between the graph structures of the two traceability graphs and the query graph.

[0102] The first initial function may be a function established based on the difference between the first node embedding matrix and the arrangement adjustment item.

[0103] The Lagrange multiplier matrix variable may be a matrix variable introduced in the optimization process, which is used to transform the first initial function into a dual form and process constraints.

[0104] The second initial function may be a new function obtained by performing a dual transformation on the first initial function. This transformation is achieved by introducing Lagrange multiplier matrix variables, which aims to simplify the optimization problem and make it easier to solve.

[0105] The Lagrange multiplier matrix can be an actual matrix obtained by calculating the Lagrange multiplier matrix variables. Each element in the Lagrange multiplier matrix is ​​used to characterize the node matching cost, and can be used to characterize the weighted constraint strength on the node alignment difference in the process of solving the target permutation matrix. For example, the larger the element, the more significant the impact of the embedding difference between the query node i and the corresponding traceability node on the overall optimization.

[0106] In some embodiments, the dimensions of the Lagrange multiplier matrix variable, the first node embedding matrix, and the second node embedding matrix can be set to be the same. When the dimensions of the three are the same, their traces can characterize the global node alignment differences, so that the Lagrange multipliers can accurately characterize the node alignment costs, to help achieve end-to-end graph matching later.

[0107] In some implementations, the objective function may be constructed by the following process: First, construct a first initial function (also known as the distance function) : ; in, is the permutation matrix variable, To arrange the adjustments, is the first node embedding matrix, Embedding matrix for the second node. The smaller the value of and The higher the correlation of the graph isomorphism relationship, the more necessary it is to permute the matrix variables The solution is performed to maximize the structural similarity between the source graph and the query graph, which facilitates matching and positioning.

[0108] Furthermore, due to is a hard permutation matrix (only 0 / 1 elements are allowed), which is more complicated to calculate and can only represent the situation where the nodes are completely matched or incompletely matched, resulting in the inability to fully utilize the potential information in the image. Therefore, S can be relaxed to a double random or "soft" permutation matrix (allowing continuous values) to facilitate subsequent calculations. In some embodiments, the target model can be used to calculate using two neural networks To approximate , where the parameter The neural network can provide a differentiable solution to the optimization problem in node alignment representation. , which is a "soft" permutation matrix, so that it can adapt to the node arrangement and alignment requirements of different graph structures, with The neural network with parameters of can be used to learn the node embedding vector matrices of the tracing graph and the query graph.

[0109] In some implementations, a Neural Network To learn a "soft" permutation matrix, we get the following distance function: ; Furthermore, since there is a ReLU term in the above formula, the problem to be solved is still not a linear problem, but it can be regarded as a linear assignment problem in the dual space for solution. Specifically, the first initial function can be converted into a dual form by using a Lagrange multiplier matrix variable of the same dimension as the first node embedding vector and the second node embedding vector, and the following second initial function is obtained: ; in, is the Lagrange multiplier matrix variable. Through the above formula, the original minimization problem can be converted into a linear assignment problem in the dual space. For example, by fixing C to calculate the optimal S, and then updating C through S, an alternating optimization process is formed to avoid the bottleneck of the non-differentiable traditional hard permutation matrix.

[0110] Furthermore, the approximate solution of the Lagrange multiplier matrix variable C can be calculated to obtain the Lagrange multiplier matrix Therefore, the second initial function can be adjusted based on the Lagrange multiplier matrix to obtain the objective function: ; in, is the Lagrange multiplier matrix. Thus, the optimization problem corresponding to the first initial function can be converted into a cost matrix as The linear assignment problem is solved so that differentiable continuous relaxation can be directly performed through the target model, thereby improving the efficiency and accuracy of the network threat data matching process.

[0111] By constructing a function based on the permutation matrix, the structural similarity between the traceability graph and the query graph can be quantified, and this similarity can be maximized by solving the permutation matrix variables, so as to achieve fast and accurate matching of the traceability graph and the query graph. In addition, in order to overcome the problem that the hard permutation matrix is ​​complex to calculate and can only represent complete or incomplete matching, the permutation matrix is ​​relaxed into a "soft" permutation matrix that allows continuous values ​​by adjusting the function. Compared with hard alignment, it is more in line with the actual topological relationship and significantly improves the efficiency and accuracy in processing network threat data.

[0112] In some embodiments, in order to improve the accuracy and efficiency of node matching in the traceability graph and the query graph, the similarity matrix of the first node embedding matrix and the second node embedding matrix can be subjected to differentiable continuous relaxation processing through the target model to obtain a target permutation matrix that is infinitely close to the optimal solution of the permutation matrix variable. At the same time, the target permutation matrix is ​​a soft matrix, which solves the problem that the hard permutation matrix is ​​not differentiable, thereby making the solution process more efficient and improving the performance of the system. Exemplarily, step 103 may include: (103.1) inputting the first node embedding matrix and the second node embedding matrix into the bilinear activation network of the target model respectively to obtain the corresponding first low-dimensional embedding representation and second low-dimensional embedding representation; (103.2) performing an inner product operation on the first low-dimensional embedding representation and the second low-dimensional embedding representation to generate a corresponding node similarity matrix; (103.3) Through the operator network, the node similarity matrix is ​​subjected to differentiable continuous relaxation so as to approach the permutation matrix variable of the objective function through the soft permutation matrix, and the soft permutation matrix is ​​used as the target permutation matrix corresponding to the permutation matrix variable.

[0113] Among them, the bilinear activation network can be a neural network structure used to convert a high-dimensional node embedding matrix into a low-dimensional representation. The bilinear activation network can capture the complex relationship between nodes through nonlinear transformations (such as the ReLU activation function) to generate a more compact and expressive low-dimensional embedding representation.

[0114] Among them, the first low-dimensional embedding representation can be a low-dimensional representation obtained by processing the first node embedding matrix through a bilinear activation network, which retains the key information of the first node embedding matrix, but the dimension is significantly reduced, which is convenient for subsequent similarity calculation.

[0115] The second low-dimensional embedding representation may be a low-dimensional representation obtained by processing the second node embedding matrix through a bilinear activation network. Similar to the first low-dimensional embedding representation, the second low-dimensional embedding representation retains the key information of the second node embedding matrix, but the dimension is significantly reduced, which is convenient for subsequent similarity calculations.

[0116] The node similarity matrix may be a matrix generated by performing an inner product operation on the first low-dimensional embedding representation and the second low-dimensional embedding representation. Each element in the node similarity matrix represents the similarity between the query node and the traceability node at the corresponding position, and the larger the value, the higher the similarity between the two.

[0117] The operator network may be a Gumbel-Sinkhorn operator, abbreviated as GS operator, which is used to process the network structure of the node similarity matrix and can convert a hard permutation matrix (i.e., a matrix of 0 or 1) into a soft permutation matrix (i.e., a probability distribution), thereby achieving an optimization solution to the objective function.

[0118] In some embodiments, the first node embedding matrix and the second node embedding matrix can be input into a matrix with parameters In the bilinear activation network (Linear-ReLU-Linear Network, LRL), the high-dimensional node embedding matrix is ​​converted to a low-dimensional embedding representation by performing a series of linear transformations and ReLU activation functions. Among them, the first low-dimensional embedding representation The specific formula for obtaining is: ; in, is the dimension reduction matrix, is the feature interaction matrix, Embedding matrix for the first node.

[0119] Furthermore, the second low-dimensional embedding representation The specific formula for obtaining is: ; in, Embedding matrix for the second node.

[0120] By converting the node embedding matrix into the corresponding low-dimensional embedding representation, the noise in the node embedding matrix can be filtered out, retaining the core feature of cross-graph topological alignment while reducing the subsequent computational complexity.

[0121] Furthermore, the inner product calculation can be performed on the first low-dimensional embedding representation and the second low-dimensional embedding representation to generate a node similarity matrix : ; The larger the element value in , the more consistent the projection directions of the corresponding traceability nodes and query nodes in the dimensionality reduction space, and the higher the possibility of topological alignment.

[0122] Furthermore, the node similarity matrix can be subjected to differentiable continuous relaxation through the operator network to approximate the optimal solution of the permutation matrix variables and obtain the corresponding target permutation matrix. The specific process is as follows: ; therefore, Can be approximated to , which is the target permutation matrix of the minimization problem in the correlation measurement of graph isomorphism relations.

[0123] In some embodiments, through the above derivation, it is proved that a soft permutation matrix can be obtained through the target model As the target permutation matrix, in order to approximate the optimal solution, solve the matching difficulties and inaccurate matching problems in the hard permutation matrix (containing only elements of 0 or 1, that is, only matching and non-matching). Therefore, in practical applications, the first node embedding matrix and the second node embedding matrix can be directly input into the target model to obtain the corresponding target permutation matrix, or the feasibility of the calculation method can be confirmed by the above derivation method, and then input into the target model to obtain the optimal solution of the objective function.

[0124] Through the above method, the complex hard permutation matrix problem can be transformed into a differentiable continuous optimization problem, avoiding the bottleneck of the non-differentiable hard permutation matrix, improving the optimization efficiency and flexibility, and achieving the best matching of nodes in the traceability graph and the query graph, which helps to quickly locate and identify potential network threats and improve the efficiency and accuracy of network threat detection.

[0125] In some implementations, in order to enhance the model's ability to understand and match complex network structures, the preset model can be trained to improve the accuracy and efficiency of network threat detection. Exemplarily, the target model is trained in the following ways: (C.1) Obtaining preset positive sample pairs and negative sample pairs, wherein each positive sample pair includes a sample provenance graph and a sample first query graph whose isomorphic distance with the sample provenance graph is less than a training warning distance threshold, and each negative sample pair includes a sample provenance graph and a sample second query graph whose isomorphic distance with the sample provenance graph is greater than a training warning distance threshold; (C.2) Obtaining the first isomorphic distance between positive sample pairs and the second isomorphic distance between negative sample pairs through the preset model; (C.3) constructing a hinge loss based on the difference between the first isomorphism distance and the second isomorphism distance; (C.4) Based on the hinge loss, the network parameters of the preset model are adjusted to obtain the target model.

[0126] Among them, the positive sample pair can be a sample pair consisting of a sample traceability graph and a sample first query graph with a small isomorphic distance (less than the training warning distance threshold). The positive sample pair represents the actual threat situation and is used to train the preset model to identify similar graph structures.

[0127] Among them, the negative sample pair can be a sample pair consisting of the same sample traceability graph as the positive sample pair and a sample second query graph with a large isomorphic distance (greater than the training warning distance threshold). Negative sample pairs represent situations where there is no threat and are used to train the model to identify dissimilar graph structures.

[0128] The sample traceability diagram may be an actual system status diagram generated by a system kernel audit log in a positive sample pair.

[0129] The training warning distance threshold may be an isomorphic distance threshold used to distinguish positive sample pairs from negative sample pairs. If the isomorphic distance between the query graph and the traceability graph is less than the training warning distance threshold, the graph pair is considered to be a positive sample pair; otherwise, it is considered to be a negative sample pair.

[0130] Among them, the sample first query graph can be a query graph constructed based on network threat intelligence in the positive sample pair.

[0131] Among them, the sample second query graph can be a query graph constructed based on network threat intelligence in the negative sample pair.

[0132] The first isomorphic distance may be the isomorphic distance between the sample source graph and the sample first query graph in the positive sample pair.

[0133] The second isomorphic distance may be the isomorphic distance between the sample source graph and the sample second query graph in the negative sample pair.

[0134] Among them, hinge loss can be a loss function for classification problems, which can optimize model parameters by maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs, prompting the preset model to learn the correct classification boundaries.

[0135] The preset model may be an initial, untrained or partially trained model based on an interpretable subgraph matching network.

[0136] For example, the positive sample pair can be traced back to the sample source graph and the sample first query graph Constitute, satisfying the first isomorphism distance , indicating that the corresponding attack model is highly correlated with the real threat behavior (such as the C2 communication mode and the log of the infected host in APT attacks). The negative sample pairs can be traced by the sample source graph and the sample second query graph Constitute, satisfying the second isomorphism distance , represents irrelevant or low-relevance behavior (such as normal user operation logs). To train the warning distance threshold (such as =0.3, which can be set according to actual conditions).

[0137] ; The subscript and The parameters of the bilinear activation network are optimized and adjusted during the training of the preset model. Finally, the hyperparameters The hinge loss is used to implement the overall training of the preset model: ; Based on hinge loss, the network parameters of the model are updated by gradient descent method. and , and minimizing the hinge loss can maximize the distance between positive and negative sample pairs and enhance the model's ability to discriminate threatening behaviors.

[0138] Calculated by learning from the preset model is the isomorphism distance of graph isomorphism relations, After the target model is trained, that is, there is no need to adjust the network parameters of the preset model, when the isomorphism distance is lower than the set training warning threshold, the model will trigger an alarm and output the traceability nodes and paths matched by the corresponding query graph in the traceability graph.

[0139] By training the preset model in the above way, the model can more accurately identify and locate potential threat behaviors when processing network threat data. When the isomorphism distance is lower than the set threshold, the model will trigger an alarm and output the traceability nodes and paths matched by the query graph in the traceability graph, significantly improving the accuracy and efficiency of network threat detection.

[0140] Step 104, using a target permutation matrix, permuting the second node embedding matrix into a target node embedding matrix aligned with the element positions in the first node embedding matrix, and determining a target isomorphism distance between the target node embedding matrix and the first node embedding matrix.

[0141] In some embodiments, in order to accurately match the nodes in the traceability graph and the query graph, the target permutation matrix can be applied to rearrange the second node embedding matrix so that its element positions are aligned with the element positions in the first node embedding matrix, thereby generating a target node embedding matrix. Then, the target isomorphism distance between the target node embedding matrix and the first node embedding matrix is ​​calculated to quantify the difference between the two graph structures and improve the accuracy and reliability of network threat detection.

[0142] The target node embedding matrix can be a new matrix obtained by applying the target permutation matrix to the second node embedding matrix. The target node embedding matrix adjusts the positions of the traceable nodes in the second node embedding matrix so that they are aligned with the node positions in the first node embedding matrix (i.e., the node embedding matrix of the query graph) as much as possible.

[0143] The target isomorphism distance can be used to quantify the difference between the target node embedding matrix and the first node embedding matrix. Specifically, the target isomorphism distance measures the structural similarity between the source graph and the query graph after the optimal node alignment.

[0144] For example, the target embedding matrix It can be calculated through the following process: ; in, is the target permutation matrix, Embedding matrix for the second node.

[0145] By adjusting the node arrangement order of the second node embedding matrix through the target permutation matrix, the topological structure of the traceability graph can be aligned with the query graph, so as to facilitate the subsequent calculation of the target isomorphism distance between the target node embedding matrix and the first node embedding matrix. For example, if the second node embedding matrix for: ; Target permutation matrix for: ; Then the target embedding matrix Can be: ; Furthermore, the target isomorphism distance can be obtained by the difference between the first node embedding matrix and the target embedding matrix. Specifically, the target isomorphism distance The calculation formula is as follows: ; in, is the first node embedding matrix, is the target permutation matrix, Embedding matrix for the second node.

[0146] Through the above methods, not only the accuracy of node matching is improved, but also the model's ability to process complex graph structures is enhanced, thereby improving the reliability, response speed and efficiency of overall threat data hunting, enabling the target model to more accurately and quickly identify potential threat behaviors in network threat detection.

[0147] Step 105: When the target isomorphism distance is less than a preset warning distance threshold, based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix, the threat event detection result of the query graph is output.

[0148] In some embodiments, in order to accurately identify and locate threatening behaviors, the calculated target isomorphic distance can be compared with a preset warning distance threshold to determine whether to trigger an alarm, thereby improving the accuracy and efficiency of threat detection and achieving accurate positioning and rapid response to potential threatening behaviors.

[0149] The preset warning distance threshold may be a pre-set distance threshold used to determine whether the similarity between the query graph and the traceability graph is sufficient to indicate the presence of a potential security threat. When the calculated target isomorphism distance is lower than the preset warning distance threshold, it indicates that the query graph and the traceability graph have a high degree of similarity, indicating that there may be network threat behavior, which requires further analysis and response.

[0150] Among them, the threat event detection result can be the output information generated by the system based on the alignment relationship between the target node embedding matrix and the second node embedding matrix when the target isomorphism distance is less than a preset warning distance threshold.

[0151] In some embodiments, when the preset model is trained, the model It can provide a differentiable solution to the optimization problem in node alignment representation, that is, a target permutation matrix that is close to a hard permutation matrix. Therefore, the target permutation matrix essentially reflects the optimal alignment relationship between the query nodes in the query graph and the traceability nodes in the traceability graph. Through this alignment relationship, the mapping position and connection path of each query node in the query graph can be directly output, thereby locating the specific position of the threat behavior in the system log.

[0152] For example, if the calculated target isomorphism distance is 0.12, and the preset warning distance threshold is 0.3, then the target isomorphism distance is less than the preset warning distance threshold, confirming that there may be a threat, and outputting the threat event detection result of the query graph in the traceability graph. Furthermore, the preset warning distance threshold can be set according to actual conditions, for example, set to 0.1, 0.5, etc., and the embodiments of the present application do not impose specific restrictions on this.

[0153] In some embodiments, the threat event detection result may include the location of the query node and the corresponding matching tracing node, the attack path (such as [malicious file] → [create] → [malicious process] → [connect] → [intranet server] → [steal] → [database file]), key evidence chain (such as timeline, behavior association, etc.). For example, the timeline can be 2025-10-01, 02:15:00 file download → 02:16:30 process creation → 02:20:00 lateral movement → 02:25:00 data leakage, and the behavior association can be that the malicious process executes the lateral movement command through the PowerShell command line tool to establish an encrypted tunnel with the server (185.xxx.xxx.xxx), etc.

[0154] The embodiment of the present application obtains a query graph constructed according to network threat intelligence, and a traceability graph generated according to a preset kernel audit log; the query graph and the traceability graph are input into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph, and a second node embedding matrix corresponding to the traceability graph; through the target model, the preset target function is minimized by isomorphic distance calculation to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, wherein the target function is used to characterize the isomorphic distance relationship between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph under different permutation matrices; through the target permutation matrix, the second node embedding matrix is ​​replaced with a target node embedding matrix aligned with the element position in the first node embedding matrix, and the target isomorphic distance between the target node embedding matrix and the first node embedding matrix is ​​determined; when the target isomorphic distance is less than the preset warning distance threshold, based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix, the threat event detection result of the query graph is output. In this way, by generating a traceability graph from the kernel audit log, the dynamic changes in the system operation process can be fully displayed, and the newly generated attack behaviors can be fully covered, breaking through the limitations of the static feature library, which is conducive to improving the accuracy of detection. At the same time, by introducing isomorphic distance minimization calculation and permutation matrix optimization, node-level alignment mapping is realized synchronously when measuring graph structure similarity, which can not only capture the topological pattern similarity of unknown attacks in the traceability graph to improve the accuracy of threat event detection, but also accurately trace the specific occurrence node of the attack in the log through the node mapping relationship, thus realizing the accurate positioning of network threat data. In summary, this application can achieve accurate positioning of network threat data while improving the accuracy of threat event detection.

[0155] In some embodiments, the combination Figure 3 The overall embodiment of the present application is introduced. Exemplarily, first, in order to facilitate subsequent threat event detection, a query graph and a traceability graph can be constructed based on network logs and threat intelligence data. The query graph represents a known threat pattern or attack behavior, while the traceability graph is generated by the system kernel audit log, which records the causal relationship and information flow between all entities in the system. It is necessary to perform threat event detection on the traceability graph based on the query graph to identify potential threats.

[0156] Furthermore, the provenance graph can be further screened according to the query graph, and the nodes and edges in the provenance graph that are not related to the query graph can be pruned to obtain a provenance subgraph, so as to eliminate the interference of irrelevant data and improve the efficiency and accuracy of graph structure matching.

[0157] Furthermore, through graph representation learning technology, the nodes of the query graph and the traceability graph (if pruned, it corresponds to the traceability subgraph here) can be converted into a first node embedding matrix and a second node embedding matrix to capture the neighborhood structure and feature information of the nodes, which facilitates subsequent efficient and accurate matching.

[0158] Furthermore, the target permutation matrix can be calculated through the target model to permute the second node embedding matrix so as to align the element positions with those in the first node embedding matrix, thereby forming a target node embedding matrix corresponding to the second node embedding matrix. Afterwards, the target isomorphism distance between the first node embedding matrix and the target node embedding matrix is ​​calculated to evaluate the structural similarity between the two.

[0159] Furthermore, by calculating the target isomorphism distance, the matching degree of the query graph in the traceability graph can be quantified. When the target isomorphism distance is lower than the preset warning distance threshold, the target model triggers an alarm and outputs the nodes and paths matched by the query graph in the traceability graph (i.e., the threat event detection results), thus achieving accurate positioning of network threats.

[0160] See also Figure 4 In some implementations, the present application also provides a threat event detection device based on provenance graph matching, which can implement the above-mentioned threat event detection method based on provenance graph matching. The threat event detection device based on provenance graph matching includes: An acquisition module 41 is used to acquire a query graph constructed according to network threat intelligence and a traceability graph generated according to a preset kernel audit log; An input module 42 is used to input the query graph and the traceability graph into the pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph; A calculation module 43 is used to perform isomorphism distance minimization calculation on a preset objective function through a target model to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, wherein the objective function is used to characterize the isomorphism distance relationship between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph under different permutation matrices; a replacement module 44, configured to replace the second node embedding matrix with a target node embedding matrix aligned with the element positions in the first node embedding matrix through a target replacement matrix, and determine a target isomorphism distance between the target node embedding matrix and the first node embedding matrix; The output module 45 is used to output the threat event detection result of the query graph based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix when the target isomorphism distance is less than a preset warning distance threshold.

[0161] The specific implementation of the threat event detection device based on provenance graph matching is basically the same as the specific implementation of the threat event detection method based on provenance graph matching described above, and will not be repeated here. On the premise of meeting the requirements of the embodiments of this application, the threat event detection device based on provenance graph matching can also be provided with other functional modules to implement the threat event detection method based on provenance graph matching in the above embodiments.

[0162] The embodiment of the present application also provides a computer device, the computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned threat event detection method based on traceability graph matching when executing the computer program. The computer device can be any intelligent terminal including a tablet computer, a car computer, etc.

[0163] See also Figure 5 , Figure 5 The hardware structure of a computer device according to another embodiment is shown, and the computer device includes: The processor 51 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application; The memory 52 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 52 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 52, and the processor 51 calls and executes the threat event detection method based on traceability graph matching in the embodiment of this application; Input / output interface 53, used to implement information input and output; Communication interface 54, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.); A bus 55 that transmits information between the various components of the device (e.g., the processor 51, the memory 52, the input / output interface 53, and the communication interface 54); The processor 51 , the memory 52 , the input / output interface 53 and the communication interface 54 are connected to each other in communication within the device via a bus 55 .

[0164] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned threat event detection method based on traceability graph matching is implemented.

[0165] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0166] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0167] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0168] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0169] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0170] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0171] It should be understood that in the present application, "at least one (item)" and "several" refer to one or more, and "plurality" refers to two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0172] In the several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0173] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0174] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0175] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.

[0176] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.

Claims

1. A threat event detection method based on provenance graph matching, characterized in that: The method comprises: Obtain the query graph built based on network threat intelligence and the traceability graph generated based on the preset kernel audit log; Inputting the query graph and the tracing graph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the tracing graph; Through the target model, the preset target function is subjected to isomorphism distance minimization calculation to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, wherein the target function is used to characterize the isomorphism distance relationship between the second node embedding matrix of the source graph and the first node embedding matrix of the query graph under different permutation matrices; By using the target permutation matrix, the second node embedding matrix is ​​permuted into a target node embedding matrix aligned with the element positions in the first node embedding matrix, and a target isomorphism distance between the target node embedding matrix and the first node embedding matrix is ​​determined; When the target isomorphism distance is less than a preset warning distance threshold, based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix, the threat event detection result of the query graph is output.

2. The threat event detection method based on traceability graph matching according to claim 1 is characterized in that: Before inputting the query graph and the traceability graph into the pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph, the method further includes: For each query node in the query graph, obtaining a corresponding query neighborhood sequence; Perform sequence matching in the traceability graph based on the query neighborhood sequence, and when the query neighborhood sequence is a subsequence of any traceability node in the traceability graph, add the any traceability node as a candidate node to the candidate node set of the corresponding query node; According to the correlation between each candidate node in the candidate node set and the corresponding query node, the candidate node set is globally optimized to obtain a corresponding target node set; Based on multiple target nodes included in the multiple target node sets and the connection relationship between the multiple target nodes in the traceability graph, construct a corresponding traceability subgraph; Then, the query graph and the traceability graph are input into the pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph, including: The query graph and the tracing subgraph are input into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the tracing subgraph.

3. The threat event detection method based on traceability graph matching according to claim 2 is characterized in that: The globally optimizing the candidate node set according to the correlation between each candidate node in the candidate node set and the corresponding query node to obtain the corresponding target node set includes: Obtaining candidate adjacent nodes of each candidate node in the candidate node set and query adjacent nodes of the corresponding query node; Based on the candidate adjacent nodes and the query adjacent nodes, construct a bipartite graph of each candidate node and the corresponding query node; When the subsequence of the query adjacent node is a subsequence of the candidate adjacent node, adding a connecting edge between the query adjacent node and the candidate adjacent node in the bipartite graph to obtain a target bipartite graph; According to the number of connection edges of each query adjacent node corresponding to the query node in the target bipartite graph, the candidate node set is globally optimized to obtain a corresponding target node set.

4. The threat event detection method based on traceability graph matching according to claim 1 is characterized in that: Before the target model is used to perform isomorphic distance minimization calculation on a preset target function to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, the method further includes: Obtaining a preset permutation matrix variable, and performing a product operation by the permutation matrix variable and the second node embedding matrix to construct an arrangement adjustment item; Taking minimizing the difference between the first node embedding matrix and the arrangement adjustment item as a constraint target, establishing a first initial function based on the difference between the first node embedding matrix and the arrangement adjustment item; Obtaining a Lagrange multiplier matrix variable, and performing a dual transformation on the first initial function through the Lagrange multiplier matrix variable to obtain a second initial function; wherein the Lagrange multiplier matrix variable, the first node embedding matrix, and the second node embedding matrix have the same dimension; The Lagrange multiplier matrix is ​​obtained by calculating the Lagrange multiplier matrix variables, and the second initial function is adjusted based on the Lagrange multiplier matrix to obtain the objective function.

5. The threat event detection method based on traceability graph matching according to claim 4 is characterized in that: The method of performing isomorphic distance minimization calculation on a preset objective function through the target model to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix includes: Inputting the first node embedding matrix and the second node embedding matrix into the bilinear activation network of the target model respectively to obtain a corresponding first low-dimensional embedding representation and a second low-dimensional embedding representation; Performing an inner product operation on the first low-dimensional embedding representation and the second low-dimensional embedding representation to generate a corresponding node similarity matrix; Through the operator network, the node similarity matrix is ​​subjected to differentiable continuous relaxation processing to approach the permutation matrix variable of the objective function through a soft permutation matrix, and the soft permutation matrix is ​​used as the target permutation matrix corresponding to the permutation matrix variable.

6. The threat event detection method based on traceability graph matching according to claim 1 is characterized in that: The target model is trained in the following way: Obtain preset positive sample pairs and corresponding negative sample pairs, wherein each positive sample pair includes a sample provenance graph and a sample first query graph whose isomorphic distance with the sample provenance graph is less than a training warning distance threshold, and each negative sample pair includes the sample provenance graph and a sample second query graph whose isomorphic distance with the sample provenance graph is greater than a training warning distance threshold; Obtaining a first isomorphic distance between the positive sample pairs and a second isomorphic distance between the negative sample pairs through a preset model; constructing a hinge loss based on the difference between the first isomorphism distance and the second isomorphism distance; Based on the hinge loss, the network parameters of the preset model are adjusted to obtain a target model.

7. The threat event detection method based on traceability graph matching according to claim 1 is characterized in that: The step of inputting the query graph and the traceability graph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the traceability graph includes: Inputting the query graph and the traceability graph into a pre-trained target model; By each neural network layer in the target model, for each query node in the query graph, based on adjacent query node features updated by adjacent query nodes in a previous neural network layer, query node features of each query node are updated, and a first node embedding matrix is ​​generated based on a plurality of query node features corresponding to the updated query nodes of a last level; Through each neural network layer, for each tracing node in the tracing graph, the tracing node features of each tracing node are updated based on the adjacent tracing node features updated in the previous neural network layer, and based on the multiple tracing node features updated corresponding to the multiple tracing nodes of the last level, a second node embedding matrix is ​​generated.

8. A threat event detection device based on provenance graph matching, characterized in that: The device comprises: The acquisition module is used to obtain the query graph constructed based on network threat intelligence and the traceability graph generated based on the preset kernel audit log; An input module, used to input the query graph and the tracing graph into a pre-trained target model to obtain a first node embedding matrix corresponding to the query graph and a second node embedding matrix corresponding to the tracing graph; A calculation module, used to perform isomorphic distance minimization calculation on a preset objective function through the target model to obtain a target permutation matrix between the first node embedding matrix and the second node embedding matrix, wherein the objective function is used to characterize the isomorphic distance relationship between the second node embedding matrix of the traceability graph and the first node embedding matrix of the query graph under different permutation matrices; a replacement module, configured to replace the second node embedding matrix with a target node embedding matrix aligned with element positions in the first node embedding matrix through the target replacement matrix, and determine a target isomorphism distance between the target node embedding matrix and the first node embedding matrix; An output module is used to output the threat event detection result of the query graph based on the element position alignment relationship between the target node embedding matrix and the second node embedding matrix when the target isomorphism distance is less than a preset warning distance threshold.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the threat event detection method based on provenance graph matching as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the threat event detection method based on provenance graph matching according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Isomorphic subgraph query method and device, electronic device and storage medium

    CN110489607A

  • A line clearance system

    US20230351582A1

Cited By

  • Financial fraud group identification method, light quantum computer and medium

    CN121329434A