Method and Device for Constructing Dynamic Edge and Hypernode Heterogeneous Graph Based on Large Model
Through the construction method of dynamic edge and supernode heterogeneous graphs based on large models, the timeliness and accuracy of traditional heterogeneous graphs in the dynamic changes of network traffic is solved, and the rapid response to network attacks and efficient detection of complex threats is achieved, which improves the accuracy and interpretability of network security analysis.
Patent Information
- Application Number
- CN202510406010.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-02
AI Technical Summary
Traditional heterogeneous graph generation methods are difficult to adapt to the dynamic changes in network traffic, resulting in insufficient timeliness and accuracy, affecting the ability to respond quickly to dynamic events such as network attacks.
The dynamic edge and supernode heterogeneous graph construction method based on large models is adopted. Network traffic data is collected in real time, network entity feature embedding vectors are extracted, and the improved DBSCAN algorithm is used for incremental clustering, semantic similarity and causal inference scores are calculated, and the weight of dynamic edges in heterogeneous graphs is dynamically updated.
It improves the timeliness and accuracy of heterogeneous graphs, enhances the ability to respond quickly to network attacks, improves the detection accuracy of APT attacks, and identifies complex threat patterns through graph neural networks and timing analysis.
Smart Images

Figure CN119922095B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network data analysis and graph construction based on large models, and particularly to a method and device for constructing a dynamic edge and hyper-node heterogeneous graph based on large models. Background Art
[0002] In recent years, significant developments have been made in the fields of network traffic analysis and network security, especially in complex network behavior modeling and anomaly detection. By constructing a relationship model between network entities, potential threats and attack behaviors can be effectively identified, which is of great significance for ensuring network security. Currently, as an effective data structure, heterogeneous graphs are widely used in network traffic analysis to represent different types of network entities and their complex relationships.
[0003] However, traditional heterogeneous graph generation methods often rely on static rules and are difficult to adapt to the dynamic change characteristics of network traffic. In the face of increasing network data, the existing technologies lack effective mechanisms to dynamically update the representation of network entities, resulting in insufficient timeliness and accuracy of heterogeneous graphs. This problem seriously affects the ability to quickly respond to dynamic events such as network attacks. Summary of the Invention
[0004] To solve the above problems in the prior art, the present invention proposes a method and device for constructing a dynamic edge and hyper-node heterogeneous graph based on large models, improving the timeliness and accuracy of heterogeneous graphs.
[0005] In the first aspect of the present invention, a method for constructing a dynamic edge and hyper-node heterogeneous graph based on large models is proposed, and the method includes:
[0006] Collect network traffic data in real time and perform preprocessing;
[0007] Extract network entities and relationship data between network entities from the preprocessed network traffic data based on a first large model;
[0008] Extract semantic, behavioral, and context features of each network entity within the current time window based on a second large model to generate a feature embedding vector for each network entity;
[0009] Use an improved DBSCAN algorithm to perform cross-time window incremental clustering on network entities according to the feature embedding vectors to obtain hyper-nodes and noise points within the current time window;
[0010] Use the hyper-nodes within the current time window as nodes of the heterogeneous graph, and construct dynamic edges between hyper-nodes according to the relationship data to form the heterogeneous graph;
[0011] Calculate the semantic similarity, temporal correlation degree, and causal inference score for the supernodes within the current time window, and then update the weights of the dynamic edges in the heterogeneous graph;
[0012] Among them, the current time window is divided according to a preset time length.
[0013] In this application, an improved DBSCAN algorithm is adopted, that is, a sliding time window mechanism is introduced and incremental clustering is performed. Through the above method, the time dimension limitation of the traditional static graph model is broken through, the dynamic evolution tracking of the network entity relationship is realized, the timeliness and accuracy of the heterogeneous graph are improved, and it helps the network security department to improve the rapid response ability to dynamic events such as network attacks; feature embedding vectors are generated according to semantic, behavioral, and context features. Compared with the single traffic feature analysis method, the detection accuracy of complex threats such as APT attacks is improved; the supernode structure generated by incremental DBSCAN clustering effectively compresses the scale of the heterogeneous graph nodes while ensuring the retention of key information.
[0014] Preferably, the step of "using the improved DBSCAN algorithm to perform cross-time window incremental clustering on network entities according to the feature embedding vectors to obtain the supernodes and noise points within the current time window" includes:
[0015] Adjust the DBSCAN clustering parameters according to the average distance and total number of the feature embedding vectors within the current time window;
[0016] If the current time window is the initial time window, perform DBSCAN clustering on the network entities within the current time window according to the feature embedding vectors to obtain a number of initial clusters and initial noise points, and use the initial clusters and the initial noise points as historical supernodes and historical noise points for subsequent clustering respectively;
[0017] If a certain historical noise point has not been merged into any supernode and has not been upgraded to an independent supernode, and does not exist within the current time window, then delete the historical noise point;
[0018] If the current time window is a non-initial time window, perform DBSCAN clustering on the network entities within the current time window with the historical supernodes and the historical noise points according to the feature embedding vectors to obtain the supernodes and noise points within the current time window;
[0019] If a certain noise point within the current time window satisfies that it has existed continuously in K time windows, then upgrade the noise point to an independent supernode; where K is a preset value;
[0020] Update the historical super nodes and the historical noise points with the super nodes and the noise points in the current time window, and calculate the average value of the feature embedding vectors of all network entities included in each historical super node as the feature embedding vector of the historical super node.
[0021] In this application, the noise points that continuously exist in K time windows are upgraded to independent super nodes, and the historical noise points that fail to survive continuously in K time windows are deleted. This filters out the temporary device access on the network (such as one-time scanning tools, which usually only exist in 1-2 windows), while the long-term hidden attackers (such as latent malicious nodes) are retained due to their continuous existence.
[0022] Preferably, the step of "adjusting the DBSCAN clustering parameters according to the average distance and the total number of the feature embedding vectors in the current time window" includes:
[0023] In the current time window, calculate the average distance between all feature embedding vectors;
[0024] Dynamically adjust the eps value according to the average distance:
[0025] ;
[0026] Count the total number of feature embedding vectors in the current time window;
[0027] Dynamically adjust the MinPts value according to the total number:
[0028] ;
[0029] Among them, is the neighborhood radius, is a preset coefficient, is the average distance, is the minimum sample number, is a preset coefficient, is the total number.
[0030] This application introduces an eps adjustment mechanism driven by the average distance, making the neighborhood radius positively correlated with the current data density, effectively avoiding the problems of over-segmentation of sparse regions and under-segmentation of dense regions; dynamically adjusting the MinPts value helps to reduce the misjudgment rate of noise points.
[0031] Preferably, the step of "using the super nodes in the current time window as the nodes of the heterogeneous graph, and constructing dynamic edges between the super nodes according to the relationship data to form the heterogeneous graph" includes:
[0032] For any two super nodes X and Y in the current time window, traverse the network entities included in X and the network entities included in Y ;
[0033] If there is explicit relationship data between entity and entity in the pre - processed network traffic data, then entity and entity are marked as a valid entity - relationship pair , and a dynamic edge is established between super - nodes X and Y; where ; ;
[0034] According to the attribute fields of the relationship data, each of the valid entity - relationship pairs is classified into a secure connection, a suspicious connection, or a data - dependency relationship by type;
[0035] Count the number of the valid entity - relationship pairs between super - node X and super - node Y by type to obtain P1, P2, and P3; where P1, P2, and P3 respectively represent the number of the secure connection, the suspicious connection, and the data - dependency relationship;
[0036] Calculate the initial weight of the dynamic edge between super - nodes X and Y according to the number of the valid entity - relationship pairs:
[0037] ;
[0038] Among them, the values of α, β, and γ are determined by fitting historical training data and satisfy the normalization condition: α + β + γ = 1.
[0039] Preferably, the step of "calculating the semantic similarity, the temporal correlation degree, and the causal inference score for the super - nodes within the current time window, and then updating the weight of the dynamic edge in the heterogeneous graph" includes:
[0040] Use a similarity - measurement method to calculate the semantic similarity between each pair of interconnected super - nodes;
[0041] Arrange the feature - embedding vectors of each super - node in all time windows in chronological order to form time - series data;
[0042] Use a time - series analysis method to calculate the temporal correlation degree between the time - series data of each pair of interconnected super - nodes;
[0043] Use a causal - discovery algorithm to infer the causal relationship between each pair of interconnected super - nodes from the time - series data, and represent the causal inference score by the edge weight of the causal graph;
[0044] For each pair of interconnected hypernodes, update the weight of the dynamic edge according to the initial weight, the semantic similarity, the temporal correlation degree, and the causal inference score:
[0045] ;
[0046] where is the updated weight of the dynamic edge; is a preset weight coefficient and satisfies , and , retaining the baseline influence of the initial weight; are respectively the initial weight, the semantic similarity, the temporal correlation degree, and the causal inference score between the interconnected hypernodes i and the hypernode j .
[0047] By combining semantic similarity, temporal correlation, and causal relationship, this application dynamically updates the weights between hypernodes, enabling a more comprehensive capture of the complex relationships between nodes, improving the accuracy and interpretability of network analysis, and at the same time retaining the influence of the initial weight to ensure stability.
[0048] Preferably, the method further includes:
[0049] If < 0.5, determine that the dynamic edge is a weak connection and mark it as a dotted line in the heterogeneous graph;
[0050] If 0.5 <= <= 0.8, mark it as a normal solid line;
[0051] If > 0.8, mark it as a bold solid line and trigger anomaly detection.
[0052] By visually differentiating the connection strength through the weight range of the dynamic edge (dotted line, normal solid line, bold solid line), this application can intuitively identify key connections and potential anomalies, improving the understanding of the network structure and the ability to quickly respond to abnormal situations.
[0053] Preferably, after "using the improved DBSCAN algorithm to perform incremental clustering of network entities across time windows according to the feature embedding vectors to obtain the hypernodes and noise points within the current time window" and before "using the hypernodes within the current time window as the nodes of the heterogeneous graph and constructing dynamic edges between the hypernodes according to the relationship data to form the heterogeneous graph", the method further includes:
[0054] If the average distance of the feature embedding vectors of network entities within a certain supernode in the current time window exceeds a preset splitting threshold, then split the supernode into network entities and recluster them into several subclusters using the improved DBSCAN algorithm;
[0055] Generate new supernodes from each subcluster and mark the split supernode as in a failed state;
[0056] Each new supernode inherits all the dynamic edges of the split supernode and distributes weights according to the proportion of the number of inherited network entities;
[0057] Traverse the network entities included in each pair of new supernodes, establish dynamic edges between the new supernodes according to the relationship data, and assign initial weights.
[0058] In this application, by dynamically splitting supernodes and reclustering, it can more accurately reflect the changes and distributions of network entities, improving the accuracy of clustering; at the same time, inheriting and distributing dynamic edge weights and establishing new connections ensure the continuity and consistency of the heterogeneous graph, enhancing the dynamic adaptability of the network structure.
[0059] Preferably, the method further includes:
[0060] Based on the heterogeneous graph, use graph neural networks and time series analysis techniques to analyze the correlation relationships between events in real time, and identify potential attack chains and attacker groups;
[0061] According to the correlation relationships, the attack chains, and the attacker groups, use community discovery and graph embedding techniques to identify groups of tightly connected supernodes in the heterogeneous graph, and combine with the third large model to perform semantic analysis on group behaviors to identify attack clusters and attack patterns;
[0062] According to the attack clusters and the attack chains, use the fourth large model to generate natural language descriptions to explain the semantic information of the supernodes, the dynamic edges, and the attack chains in the heterogeneous graph, and combine with visualization techniques to generate a threat map.
[0063] In this application, by combining graph neural networks, time series analysis, and semantic techniques, it can identify complex attack chains and attacker groups in real time, revealing potential threat patterns; at the same time, generating natural language descriptions and visualizing threat maps enhances the interpretability of threat intelligence and decision-making support capabilities, facilitating efficient security protection.
[0064] In the second aspect of the present invention, an electronic device is proposed, including a processor and a memory, and a computer program capable of being loaded and executed by the processor as described in the above method is stored on the memory.
[0065] In a third aspect of the present invention, a computer-readable storage device is provided, which is characterized in that it stores a computer program that can be loaded and executed by a processor to perform the method described above. Description of the Drawings
[0066] Figure 1 It is a schematic diagram of the main steps of an embodiment of the method for constructing a dynamic edge and hyper-node heterogeneous graph based on a large model in this application. Detailed Embodiments
[0067] The preferred embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the protection scope of the present application.
[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0069] It should be noted that in the description of the present application, the terms "first" and "second" are only for the convenience of description and do not indicate or imply the relative importance of the devices, elements, or parameters, and therefore should not be construed as limiting the present application. In addition, the term "and / or" in the present application is only a description of the associated relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, unless otherwise specified.
[0070] Figure 1 It is a schematic diagram of the main steps of an embodiment of the method for constructing a dynamic edge and hyper-node heterogeneous graph based on a large model in this application. As Figure 1 shown, the construction method of this embodiment includes steps A1 - A7:
[0071] Step A1: Real-time collect network traffic data and perform preprocessing.
[0072] The preprocessing includes data cleaning (removing noise and redundant information), format standardization (unifying data formats), etc.
[0073] Step A2: Extract network entities and relationship data between network entities from the preprocessed network traffic data based on the first large model.
[0074] The first large model can be a rule-based engine or a machine learning model and has been trained; network entities can include devices, users, applications, IP addresses and ports, domain names and URLs, etc.; relationship data such as communication connections, data flows, etc.
[0075] Step A3: Extract semantic, behavioral, and contextual features of each network entity within the current time window based on the second large model to generate a feature embedding vector for each network entity.
[0076] Among them, the current time window is divided according to a preset time length; the second large model can be a deep learning model (such as Transformer or GNN) and has been trained. Semantic features include roles, functions, etc., behavioral features include access frequency, data volume, etc., and contextual features include time, geographical location, etc.
[0077] Step A4: Use the improved DBSCAN algorithm to perform cross-time window incremental clustering on network entities according to the feature embedding vectors to obtain supernodes and noise points within the current time window.
[0078] Specifically, this step may include steps A41 - A46:
[0079] Step A41: Adjust the DBSCAN clustering parameters according to the average distance and total number of feature embedding vectors within the current time window.
[0080] Specifically, step A41 includes steps A411 - A414:
[0081] Step A411: Calculate the average distance between all feature embedding vectors within the current time window.
[0082] Step A412: Dynamically adjust the eps value according to the average distance, as shown in formula (1):
[0083] (1)
[0084] Because the average distance can reflect the distribution density of network entities within the current time window, and the eps parameter determines the neighborhood radius between data points. When the average distance is large, it indicates that the network entities are sparsely distributed, and the eps should be appropriately increased to expand the neighborhood range, so as to discover larger clusters; conversely, when the average distance is small, it indicates that the network entities are densely distributed, and the eps should be appropriately decreased to narrow the neighborhood range, so as to discover finer clusters.
[0085] Step A413: Count the total number of feature embedding vectors within the current time window.
[0086] Step A414: Dynamically adjust the MinPts value according to the total number, as shown in formula (2):
[0087] (2)
[0088] Wherein, is the neighborhood radius, a1 is a preset coefficient, is the said average distance, is the minimum number of samples, a2 is a preset coefficient, is the said total quantity.
[0089] This total quantity can reflect the activity degree of network entities within the current time window. Dynamically adjusting the MinPts value helps reduce the misjudgment rate of noise points.
[0090] Step A42: If the current time window is the initial time window, cluster the network entities within the current time window according to the feature embedding vectors to obtain a number of initial clusters and initial noise points, and use the initial clusters and initial noise points as historical super nodes and historical noise points for subsequent clustering respectively.
[0091] Step A43: If a certain historical noise point has not been merged into any super node and has not been upgraded to an independent super node, and does not exist within the current time window, then delete the historical noise point.
[0092] This step reduces the redundancy of historical data and ensures the effectiveness of noise point management.
[0093] Step A44: If the current time window is not the initial time window, cluster the network entities within the current time window with the historical super nodes and historical noise points according to the feature embedding vectors to obtain the super nodes and noise points within the current time window.
[0094] In this step, historical super nodes may absorb new network entities, and historical noise points may be clustered into a certain super node. When clustering, the embedding feature vector of the historical super node takes the average value of the feature embedding vectors of all network entities it contains.
[0095] Step A45: If a certain noise point within the current time window satisfies that it has existed continuously in K time windows, then upgrade the noise point to an independent super node; wherein, K is a preset value.
[0096] Step A46: Update the historical super nodes and historical noise points with the super nodes and noise points within the current time window, and calculate the average value of the feature embedding vectors of all network entities contained in each historical super node as the feature embedding vector of the historical super node.
[0097] Step A5: Take the super-nodes within the current time window as the nodes of the heterogeneous graph, and construct dynamic edges between the super-nodes according to the relationship data to form a heterogeneous graph.
[0098] Specifically, this step may include steps A51 - A55:
[0099] Step A51: For any two super-nodes X and Y within the current time window, traverse the network entities included in X and the network entities included in Y .
[0100] Step A52: If there is explicit relationship (such as communication connection, data dependency relationship) data between entity and entity in the preprocessed network traffic data, then mark entity and entity as a valid entity relationship pair , and establish a dynamic edge between super-nodes X and Y;
[0101] Among them, ; ; The display relationship data may include: communication connection (such as TCP / UDP connection, HTTP request), data dependency relationship (such as database query, file transfer), etc.
[0102] Step A53: According to the attribute fields of the relationship data, divide each valid entity relationship pair into a secure connection, a suspicious connection, or a data dependency relationship by type.
[0103] The classification basis may include: communication protocol (such as HTTPS is secure, unknown protocol is suspicious), access permission (such as legitimate user is secure, unauthorized user is suspicious), behavior pattern (such as high-frequency communication is suspicious, low-frequency communication is secure), etc.
[0104] A secure connection is a relationship that conforms to the normal behavior pattern, such as authenticated communication, legitimate data access; a suspicious connection is a relationship that may have potential risks, such as communication with abnormal frequency, unauthorized access attempt; a data dependency relationship represents the data flow or dependency between entities, such as database query, file read and write operations.
[0105] Step A54: Count the number of valid entity relationship pairs between super-node X and super-node Y by type to obtain P1, P2, and P3; where P1, P2, and P3 respectively represent the number of secure connections, suspicious connections, and data dependency relationships;
[0106] Step A55: Calculate the initial weight of the dynamic edge between super-nodes X and Y according to the number of the valid entity relationship pairs, as shown in formula (3):
[0107] (3)
[0108] Among them, the values of α, β, and γ are determined by fitting historical training data and satisfy the normalization condition: α + β + γ = 1.
[0109] Step A6: Calculate the semantic similarity, temporal correlation degree, and causal inference score for the hypernodes within the current time window, and then update the weights of the dynamic edges in the heterogeneous graph.
[0110] Specifically, this step may include steps A61 - A65:
[0111] Step A61: Calculate the semantic similarity between each pair of interconnected hypernodes using a similarity measurement method.
[0112] The similarity measurement method can adopt, such as, cosine similarity, Euclidean distance, Jaccard similarity coefficient, etc.
[0113] Step A62: Arrange the feature embedding vectors of each hypernode within all time windows in chronological order to form time - series data.
[0114] Step A63: Calculate the temporal correlation degree between the time - series data of each pair of interconnected hypernodes using a temporal analysis method.
[0115] The temporal analysis method can adopt, such as, Pearson correlation coefficient, dynamic time warping (DTW), Granger causality test, etc.
[0116] Step A64: Use a causal discovery algorithm to infer the causal relationship between each pair of interconnected hypernodes from the time - series data, and represent the causal inference score through the edge weights of the causal graph.
[0117] The causal discovery algorithm can adopt, such as, PC algorithm, LiNGAM, Granger causality test, etc.
[0118] Step A65: For each pair of interconnected hypernodes, update the weight of the dynamic edge according to the initial weight, semantic similarity, temporal correlation degree, and causal inference score, as shown in formula (4):
[0119] (4)
[0120] Among them, is the updated weight of the dynamic edge; is a preset weight coefficient and satisfies , and , retaining the benchmark influence of the initial weight; are respectively the interconnected hypernodes i and hypernodej The initial weights, semantic similarity, temporal correlation, and causal inference scores of the dynamic edges among them.
[0121] Step A7: Draw the corresponding dynamic edges as dashed lines, solid lines, or bold solid lines according to the updated dynamic edge weights.
[0122] Specifically as follows:
[0123] If < 0.5, determine that the dynamic edge is a weak connection and mark it as a dashed line in the heterogeneous graph;
[0124] If 0.5 <= <= 0.8, mark it as an ordinary solid line;
[0125] If > 0.8, mark it as a bold solid line and trigger anomaly detection.
[0126] In an alternative embodiment, after step A46 and before step A5, the method may further include:
[0127] Step A47: If a certain supernode meets the splitting condition, split it.
[0128] Specifically, it may include steps A471 - A474:
[0129] Step A471: If the average distance of the feature embedding vectors of the network entities within a certain supernode in the current time window exceeds a preset splitting threshold (such as Euclidean distance > 2.0), split the supernode into network entities and recluster them into several sub - clusters using the improved DBSCAN algorithm.
[0130] Step A472: Generate new supernodes from each sub - cluster and mark the split supernode as in a failed state.
[0131] Step A473: Each new supernode inherits all the dynamic edges of the split supernode and distributes the weights according to the proportion of the number of inherited network entities.
[0132] For example, if a certain supernode A is split and reclustered into supernode A1 and supernode A2, and A1 and A2 inherit 70% and 30% of the network entities respectively, then both A1 and A2 will inherit all the dynamic edges of A, and the weight of each dynamic edge inherited by A1 is set to 70% of the original weight, and the weight of each dynamic edge inherited by A2 is set to 30% of the original weight. The weights assigned here serve as the initial weights of each dynamic edge after inheritance.
[0133] Step A474: Traverse the network entities included in each pair of new supernodes, establish dynamic edges between the new supernodes according to the relationship data, and assign initial weights.
[0134] What is inherited in step A473 above is the dynamic edges between supernode A and other external supernodes. What is obtained in this step is the dynamic edges between several new supernodes split from A.
[0135] In another alternative embodiment, after step A7, it may further include:
[0136] Step A8: Analyze network attack events and potential threats based on the heterogeneous graph.
[0137] Specifically, step A8 may include steps A81 - A83:
[0138] Step A81: Based on the heterogeneous graph, use graph neural network and time - series analysis techniques to analyze the correlation relationships between events in real - time, and identify potential attack chains and attacker groups.
[0139] Step A82: According to the correlation relationships, attack chains and attacker groups, use community discovery and graph embedding techniques to identify groups of tightly - connected supernodes in the heterogeneous graph, and combine with the third large - scale model to perform semantic analysis on the group behavior, and identify attack clusters and attack patterns.
[0140] Step A83: According to the attack clusters and attack chains, use the fourth large - scale model to generate a natural - language description to explain the semantic information of supernodes, dynamic edges and attack chains in the heterogeneous graph, and combine with visualization techniques to generate a threat map.
[0141] Although the above - mentioned embodiments describe each step in the above - mentioned sequential order, those skilled in the art can understand that in order to achieve the effects of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reversed order, and these simple changes are all within the protection scope of the present invention.
[0142] Based on the above - mentioned method embodiments, the present application also provides an embodiment of an electronic device. The electronic device of this embodiment includes a processor and a memory, and a computer program capable of being loaded and executed by the processor as described in the above - mentioned method is stored on the memory.
[0143] Furthermore, based on the above - mentioned method embodiments, the present application also provides an embodiment of a computer - readable storage device. A computer program capable of being loaded and executed by a processor as described in the above - mentioned method is stored in the storage device of this embodiment.
[0144] The computer - readable storage device may include: various media such as USB flash drives, mobile hard disks, read - only memory (ROM), random access memory (RAM), magnetic disks or optical discs that can store program codes.
[0145] Those skilled in the art should be able to realize that the method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0146] So far, the technical solution of the present invention has been described in combination with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
Claims
1. A method for constructing a dynamic edge and hyper-node heterogeneous graph based on a large model, characterized in that The method includes: Collecting network traffic data in real time and performing preprocessing; Extracting network entities and relationship data between network entities from the preprocessed network traffic data based on a first large model; Extracting semantic, behavioral, and context features of each network entity within the current time window based on a second large model to generate a feature embedding vector for each network entity; Using an improved DBSCAN algorithm to perform cross-time-window incremental clustering on network entities according to the feature embedding vectors to obtain supernodes and noise points within the current time window; Taking the supernodes within the current time window as nodes of a heterogeneous graph, and constructing dynamic edges between the supernodes according to the relationship data to form the heterogeneous graph; Calculating semantic similarity, temporal correlation degree, and causal inference score for the supernodes within the current time window, and then updating the weights of the dynamic edges in the heterogeneous graph; Wherein, the current time window is divided according to a preset time length; The step of "taking the supernodes within the current time window as nodes of a heterogeneous graph, and constructing dynamic edges between the supernodes according to the relationship data to form the heterogeneous graph" includes: For any two super nodes X and Y within the current time window, traverse the network entities {x1, x2,..., x M} contained in X and the network entities {y1, y2,..., y N} contained in Y; If entity x m and entity y n have explicit relationship data in the pre - processed network traffic data, then entity x m and entity y n are marked as a valid entity - relationship pair (x m , y n ), and a dynamic edge is established between super - nodes X and Y; where, m = 1, 2, …, M; n = 1, 2, …, N; Dividing each valid entity relationship pair into a secure connection, a suspicious connection, or a data dependency relationship by type according to the attribute fields of the relationship data; Counting the number of valid entity relationship pairs between supernode X and supernode Y by type to obtain P1, P2, and P3; where P1, P2, and P3 respectively represent the number of the secure connection, the suspicious connection, and the data dependency relationship; Calculating the initial weight of the dynamic edge between supernodes X and Y according to the number of valid entity relationship pairs: W init = α·P1 + β·P2 + γ·P3; Wherein, the values of α, β, and γ are determined by fitting historical training data and satisfy the normalization condition: α + β + γ = 1.
2. The method for constructing a dynamic edge and hyper-node heterogeneous graph based on a large model according to claim 1, wherein The step of "using an improved DBSCAN algorithm to perform cross-time-window incremental clustering on network entities according to the feature embedding vectors to obtain supernodes and noise points within the current time window" includes: Adjusting the DBSCAN clustering parameters according to the average distance and total number of feature embedding vectors within the current time window; If the current time window is the initial time window, performing DBSCAN clustering on the network entities within the current time window according to the feature embedding vectors to obtain a number of initial clusters and initial noise points, and taking the initial clusters and the initial noise points as historical supernodes and historical noise points for subsequent clustering respectively; If a certain historical noise point has not been merged into any supernode and has not been upgraded to an independent supernode and does not exist within the current time window, deleting the historical noise point; If the current time window is a non-initial time window, performing DBSCAN clustering on the network entities within the current time window with the historical supernodes and the historical noise points according to the feature embedding vectors to obtain supernodes and noise points within the current time window; If a certain noise point within the current time window satisfies that it has existed continuously in K time windows, upgrading the noise point to an independent supernode; where K is a preset value; Update the historical super nodes and the historical noise points with the super nodes and the noise points in the current time window, and calculate the average value of the feature embedding vectors of all network entities included in each historical super node as the feature embedding vector of the historical super node.
3. The method for constructing a dynamic edge and hyper-node heterogeneous graph based on a large model according to claim 2, wherein The step of "adjusting the DBSCAN clustering parameters according to the average distance and the total number of the feature embedding vectors in the current time window" includes: Calculate the average distance between all the feature embedding vectors in the current time window; Dynamically adjust the value of eps according to the average distance: eps = a1 * Distance; Count the total number of the feature embedding vectors in the current time window; Dynamically adjust the value of MinPts according to the total number: MinPts = a2 * Total; Where, eps is the neighborhood radius, a1 is a preset coefficient, Distance is the average distance, MinPts is the minimum number of samples, a2 is a preset coefficient, and Total is the total number.
4. The method for constructing a dynamic edge and hyper-node heterogeneous graph based on a large model according to claim 1, wherein The step of "calculating the semantic similarity, the temporal correlation degree and the causal inference score for the super nodes in the current time window, and then updating the weights of the dynamic edges in the heterogeneous graph" includes: Calculate the semantic similarity between each pair of interconnected super nodes by using a similarity measurement method; Arrange the feature embedding vectors of each super node in all time windows in chronological order to form time series data; Calculate the temporal correlation degree between the time series data of each pair of interconnected super nodes by using a temporal analysis method; Infer the causal relationship between each pair of interconnected super nodes from the time series data by using a causal discovery algorithm, and represent the causal inference score by the edge weight of the causal graph; For each pair of interconnected super nodes, update the weight of the dynamic edge according to the initial weight, the semantic similarity, the temporal correlation degree and the causal inference score: Wij = Q1·Winit + Q2·Sij + Q3·Tij + Q4·Cij; Among them, W ij is the updated dynamic edge weight; Q1, Q2, Q3, and Q4 are preset weight coefficients, and Q1 + Q2 + Q3 + Q4 = 1, and Q1 > 0.2 to retain the benchmark influence of the initial weight; W init、 S ij、 T ij and C ij are respectively the initial weight, the semantic similarity, the temporal correlation degree, and the causal inference score of the dynamic edge between the interconnected super-nodes i and super-node j.
5. The method for constructing a dynamic edge and hyper-node heterogeneous graph based on a large model according to claim 4, wherein The method further includes: If W ij < 0.5, determine that the dynamic edge is a weak connection and mark it as a dotted line in the heterogeneous graph; If 0.5 <= W ij <= 0.8, it is marked as a solid line; If W ij > 0.8, mark it as a bold solid line and trigger anomaly detection.
6. The method for constructing a dynamic edge and hyper-node heterogeneous graph based on a large model according to claim 1, wherein After "using the improved DBSCAN algorithm to perform incremental clustering of network entities across time windows according to the feature embedding vectors to obtain the super nodes and the noise points in the current time window", and before "using the super nodes in the current time window as the nodes of the heterogeneous graph, and constructing dynamic edges between the super nodes according to the relationship data to form the heterogeneous graph", the method further includes: If the average distance of the feature embedding vectors of the network entities in a certain super node in the current time window exceeds a preset splitting threshold, then split the super node into network entities and re-cluster them into several sub-clusters by using the improved DBSCAN algorithm; Generate new super nodes from each sub-cluster, and mark the split super node as in a failed state; Each new super node inherits all the dynamic edges of the split super node, and distributes the weights according to the proportion of the number of inherited network entities; Traverse the network entities included in each pair of new super nodes, and establish dynamic edges between the new super nodes according to the relationship data and assign initial weights.
7. The method for constructing a dynamic edge and hyper-node heterogeneous graph based on a large model according to claim 1, wherein The method further includes: Based on the heterogeneous graph, using graph neural network and time series analysis techniques, analyze the correlation relationships between events in real time, and identify potential attack chains and attacker groups; According to the correlation relationships, the attack chains and the attacker groups, using community discovery and graph embedding techniques, identify groups of tightly connected supernodes in the heterogeneous graph, and combine with the third large model to perform semantic analysis on group behaviors, and identify attack clusters and attack patterns; According to the attack clusters and the attack chains, use the fourth large model to generate a natural language description to explain the semantic information of the supernodes, the dynamic edges and the attack chains in the heterogeneous graph, and combine with visualization techniques to generate a threat map.
8. An electronic device, characterized in that, It includes a processor and a memory, and a computer program capable of being loaded and executed by the processor as described in any one of claims 1-7 is stored on the memory.
9. A computer-readable storage device, characterized in that, Stores a computer program capable of being loaded and executed by a processor as described in any one of claims 1-7.
Citation Information
Patent Citations
Vehicle using travel time prediction method and device, server and storage medium
CN115713168A
Industrial internet time series data anomaly detection method and system
CN118898045A
Network attack link tracking and threat situation reasoning method based on knowledge graph
CN119544327A