Network Attack Homology Analysis Method, Device, Computer Equipment and Storage Medium
By constructing preset weighted attack behavior diagrams and modularity divisions, the problems of high computational complexity and inaccurate results in homologous analysis of network attacks are solved, efficient and interpretable homologous analysis is achieved, and network security protection capabilities are improved.
Patent Information
- Application Number
- CN202510646517.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-20
AI Technical Summary
In the prior art, the calculation complexity of network attack homologous analysis is high, low efficiency, and the clustering results lack interpretability, resulting in misjudgment or misjudgment of homologous relationships, reducing the accuracy of the analysis results.
By constructing a preset weighted attack behavior diagram, using connection weights and modularity to divide the address cluster, calculate the characteristic center of mass, perform similarity analysis, and improve the accuracy and efficiency of homology analysis.
It realizes efficient and accurate analysis of the homologous relationship of network attacks, improves network security protection capabilities, and ensures the traceability of attack patterns and interpretability of analysis results.
Smart Images

Figure CN120165989B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network security technology, and in particular, to a method, apparatus, computer device, and storage medium for network attack homology analysis. Background Art
[0002] A network attack is an act of using computer network technology to illegally intrude into, interfere with, damage, or steal information resources in a network system to achieve specific goals. It can include, but is not limited to, various forms such as data leakage, system damage, denial-of-service attacks, and malware propagation, posing a serious threat to personal privacy, corporate interests, social order, and security.
[0003] In order to more effectively respond to complex and changing network attacks and enhance the overall protection ability of network security, the correlation between different attack events can be identified through network attack homology analysis, and it can be determined whether they come from the same attack source or organization, so as to reveal the attacker's behavior patterns and technical means, improve the defense efficiency, and reduce resource waste.
[0004] In related technologies, usually, the Internet Protocol addresses corresponding to each attack event are used as nodes, and the feature similarity between all Internet Protocol addresses is compared pairwise, and the Internet Protocol addresses with high feature similarity are clustered to identify the homologous relationship between attackers. However, on the one hand, the method of clustering after pairwise comparison has a high computational complexity and low analysis efficiency when facing a large amount of attack data; on the other hand, the clustering results lack interpretability, resulting in security personnel being unable to effectively trace the attack pattern, and thus misjudging or missing the homologous relationship, reducing the accuracy of the analysis results. Summary of the Invention
[0005] This application proposes a method, apparatus, computer device, and storage medium for network attack homology analysis, which can improve the efficiency and accuracy of network attack homology analysis.
[0006] To achieve the above object, the first aspect of the embodiments of this application proposes a method for network attack homology analysis, and the method includes:
[0007] Obtain a plurality of preset Internet Protocol addresses and a preset weighted attack behavior graph, where the preset weighted attack behavior graph includes a plurality of preset Internet Protocol addresses and a plurality of connection weights, and each connection weight represents the degree of association between two preset Internet Protocol addresses with a homologous relationship;
[0008] Divide the plurality of preset Internet Protocol addresses in the preset weighted attack behavior graph into a plurality of initial address clusters;
[0009] For each initial address cluster, determine the connection weights associated with each preset Internet protocol address included, and calculate the corresponding modularity according to the connection weights, where the modularity is used to characterize the homology compatibility degree among the multiple preset Internet protocol addresses included in the corresponding initial address cluster;
[0010] According to the modularity corresponding to each initial address cluster, re-determine the attribution relationship between each preset Internet protocol address and the initial address cluster, obtain multiple target address clusters, and determine the characteristic centroid of each target address cluster;
[0011] Obtain the Internet protocol address to be analyzed, and calculate the similarity between the target feature vector of the Internet protocol address to be analyzed and the multiple characteristic centroids of the multiple target address clusters, so as to obtain the network attack homology analysis result of the Internet protocol address to be analyzed.
[0012] Correspondingly, a second aspect of the embodiments of the present application proposes a network attack homology analysis device, and the device includes:
[0013] An acquisition module, configured to acquire a plurality of preset Internet protocol addresses and a preset weighted attack behavior graph, where the preset weighted attack behavior graph includes a plurality of preset Internet protocol addresses and a plurality of connection weights, and each connection weight represents the association degree between two preset Internet protocol addresses with a homologous relationship;
[0014] A partitioning module, configured to partition the plurality of preset Internet protocol addresses in the preset weighted attack behavior graph into a plurality of initial address clusters;
[0015] A first calculation module, configured to, for each initial address cluster, determine the connection weights associated with each preset Internet protocol address included, and calculate the corresponding modularity according to the connection weights, where the modularity is used to characterize the homology compatibility degree among the multiple preset Internet protocol addresses included in the corresponding initial address cluster;
[0016] A determination module, configured to re-determine the attribution relationship between each preset Internet protocol address and the initial address cluster according to the modularity corresponding to each initial address cluster, obtain a plurality of target address clusters, and determine the characteristic centroid of each target address cluster;
[0017] A second calculation module, configured to acquire the Internet protocol address to be analyzed, and calculate the similarity between the target feature vector of the Internet protocol address to be analyzed and the multiple characteristic centroids of the multiple target address clusters, so as to obtain the network attack homology analysis result of the Internet protocol address to be analyzed.
[0018] In some embodiments, the acquisition module is further configured to:
[0019] Obtain the multiple preset Internet protocol addresses and the homologous relationship between the multiple preset Internet protocol addresses, and determine the first feature vector and relationship label between any two preset Internet protocol addresses, where the first feature vector includes multiple first feature values of the any two preset Internet protocol addresses in multiple feature interaction dimensions;
[0020] Sequentially input each first feature vector and the corresponding relationship label into the initial gradient boosting model, and perform decision tree split gain accumulation through the initial gradient boosting model to obtain the feature importance score corresponding to each first feature value;
[0021] Based on the multiple feature importance scores, determine the connection weight between the two preset Internet protocol addresses corresponding to the first feature vector;
[0022] Based on the multiple preset Internet protocol addresses, the relationship label between any two preset Internet protocol addresses, and the corresponding connection weight, construct the preset weighted attack behavior graph corresponding to the multiple preset Internet protocol addresses.
[0023] In some embodiments, the obtaining module is further configured to:
[0024] Sequentially input each first feature vector and the corresponding relationship label into the initial gradient boosting model, and through the initial gradient boosting model, predict each first feature vector to obtain the corresponding prediction label;
[0025] Based on the difference between the prediction label and the relationship label, determine the first target loss;
[0026] Based on the first target loss, perform split gain calculation on the multiple first feature values included in each first feature vector, and determine the intermediate split feature with the largest split gain from the multiple first feature values for node splitting, and calculate the second target loss;
[0027] Based on the difference between the second target loss and the first target loss, perform split gain calculation on the multiple first feature values included in the first feature vector, and determine the intermediate split feature with the largest split gain from the multiple first feature values for node splitting, and calculate the updated second target loss;
[0028] Repeat the step of performing split gain calculation on the multiple first eigenvalues included in the first feature vector based on the difference between the updated second target loss and the first target loss, determining the intermediate split feature with the largest split gain from the multiple first eigenvalues for node splitting, and calculating to obtain the updated second target loss until the preset number of training times is reached, and obtaining the feature importance score corresponding to each first eigenvalue according to the multiple split gains iteratively obtained for each eigenvalue.
[0029] In some embodiments, the obtaining module is further configured to:
[0030] Perform normalization processing on each feature importance score to obtain the corresponding target feature importance score;
[0031] Obtain the sub-feature weight of each first eigenvalue according to the product of each first eigenvalue and the corresponding target feature importance score;
[0032] Based on the sum of the multiple sub-feature weights corresponding to the multiple first eigenvalues, obtain the connection weight between the two preset Internet protocol addresses corresponding to the first feature vector.
[0033] In some embodiments, the network attack homology analysis device further includes a comparison module, configured to:
[0034] Obtain multiple historical connection weights between any two preset Internet protocol addresses, as well as the scoring mean and scoring standard deviation of the multiple historical connection weights;
[0035] Obtain a preset adjustment parameter, and obtain a first product according to the product of the adjustment parameter and the scoring standard deviation;
[0036] Determine a first dynamic threshold according to the difference between the scoring mean and the first product;
[0037] Compare the connection weight between any two preset Internet protocol addresses with the first dynamic threshold to obtain a comparison result;
[0038] Update the relationship label between any two preset Internet protocol addresses based on the comparison result to obtain a target relationship label;
[0039] Then, constructing the preset weighted attack behavior graph corresponding to the multiple preset Internet protocol addresses based on the multiple preset Internet protocol addresses, the relationship labels between any two preset Internet protocol addresses, and the corresponding connection weights includes:
[0040] Construct a preset weighted attack behavior graph corresponding to multiple preset Internet protocol addresses based on the target relationship tags and corresponding connection weights between any two preset Internet protocol addresses.
[0041] In some embodiments, the first calculation module is further configured to:
[0042] Obtain a graph connection weight according to the sum of the connection weights corresponding to all connection edges in the preset weighted attack behavior graph;
[0043] Obtain the node degree corresponding to the target node corresponding to each preset Internet protocol address in the initial address cluster;
[0044] Obtain a second product according to the product of the node degrees between any two target nodes;
[0045] Obtain a first ratio based on the ratio of the second product to the graph connection weight;
[0046] Obtain a first difference, which is the difference between the connection weight corresponding to any two target nodes in the initial address cluster and the corresponding first ratio;
[0047] Obtain the modularity corresponding to the initial address cluster based on the graph connection weight and the multiple first differences between multiple target nodes included in the initial address cluster.
[0048] In some embodiments, the determination module is further configured to:
[0049] Determine an intermediate address cluster with a modularity less than a preset second dynamic threshold according to the modularity corresponding to each initial address cluster;
[0050] Obtain a preset attack behavior feature library, and based on the attack behavior feature library, re-determine the associated Internet protocol addresses having an attack feature relationship with each preset Internet protocol address in each intermediate address cluster, and migrate each preset Internet protocol address to the initial address cluster corresponding to the associated Internet protocol address;
[0051] Repeat the step of re-determining the associated Internet protocol addresses having the attack feature relationship with each preset Internet protocol address in each intermediate address cluster based on the attack behavior feature library, and migrating each preset Internet protocol address to the initial address cluster corresponding to the associated Internet protocol address until the modularity corresponding to each intermediate address cluster is greater than the second dynamic threshold, to obtain multiple target address clusters.
[0052] In some embodiments, the determination module is further configured to:
[0053] For each target address cluster, obtain the second feature vector of the target node corresponding to each preset Internet Protocol (IP) address in the target address cluster, and the node degree of the target node;
[0054] Obtain a third product according to the product of the second feature vector and the corresponding node degree;
[0055] Obtain a feature sum according to the sum of multiple third products corresponding to multiple preset IP addresses included in each target address cluster;
[0056] Obtain the sum of the node degrees of multiple target nodes corresponding to each target address cluster to obtain a target degree sum;
[0057] Based on the ratio of the feature sum to the target degree sum, obtain the feature centroid of each target address cluster.
[0058] In some embodiments, the second calculation module is further configured to:
[0059] Calculate the similarity between each target feature in the target feature vector corresponding to the IP address to be analyzed and each sub-feature centroid in each feature centroid to obtain a corresponding feature similarity;
[0060] Obtain the feature centroid importance score corresponding to each sub-feature centroid, and based on the product of the feature similarity and the feature centroid importance score, obtain the target feature similarity corresponding to each target feature;
[0061] Add the multiple target feature similarities corresponding to multiple target features to obtain the total feature similarity corresponding to each feature centroid;
[0062] Obtain a preset third dynamic threshold, and sequentially compare the multiple total feature similarities corresponding to the multiple feature centroids with the third dynamic threshold to obtain a comparison result;
[0063] When the comparison result indicates that there is a target total feature similarity greater than the third dynamic threshold among the multiple total feature similarities, determine the corresponding target address cluster as the target homologous cluster of the IP address to be analyzed, and based on the target homologous cluster, obtain the network attack homologous analysis result of the IP address to be analyzed.
[0064] Correspondingly, a third aspect of the embodiments of the present application proposes a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the network attack homologous analysis method according to any one of the embodiments of the first aspect of the present application.
[0065] Correspondingly, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the network attack homology analysis method according to any one of the embodiments of the first aspect of the present application.
[0066] The present application obtains a plurality of preset Internet protocol addresses and a preset weighted attack behavior graph. The preset weighted attack behavior graph includes a plurality of preset Internet protocol addresses and a plurality of connection weights, and each connection weight represents the association degree between two preset Internet protocol addresses with a homology relationship. The plurality of preset Internet protocol addresses in the preset weighted attack behavior graph are divided into a plurality of initial address clusters. For each initial address cluster, the connection weights associated with each preset Internet protocol address included are determined, and the corresponding modularity is calculated according to the connection weights. The modularity is used to characterize the homology compatibility degree among the plurality of preset Internet protocol addresses included in the corresponding initial address cluster. According to the modularity corresponding to each initial address cluster, the attribution relationship between each preset Internet protocol address and the initial address cluster is re-determined to obtain a plurality of target address clusters, and the characteristic centroid of each target address cluster is determined. An Internet protocol address to be analyzed is obtained, and the similarity between the target feature vector of the Internet protocol address to be analyzed and the plurality of characteristic centroids of the plurality of target address clusters is calculated to obtain the network attack homology analysis result of the Internet protocol address to be analyzed. In this way, the correlation of the attack behaviors between any two Internet protocol addresses can be accurately quantified through the connection weights between the plurality of Internet protocol addresses in the pre-generated preset weighted attack behavior graph, so as to improve the interpretability of the relationship between each Internet protocol address, and further improve the accuracy of homology analysis. And, by adopting the method of modularity iterative processing, the preset weighted attack behavior graph is divided into target address clusters, and the characteristic centroid of each target address cluster is calculated. In this way, a globally representative core feature vector can be accurately generated based on the homologous Internet protocol addresses included in the same address cluster, so that when there is an Internet protocol address to be analyzed that needs to perform homology analysis subsequently, the Internet protocol address to be analyzed can be directly compared with the characteristic centroids of each target address cluster (including a large number of homologous Internet protocol addresses), without having to compare and cluster with a large number of Internet protocol addresses one by one, and the homology is determined based on the centroid similarity, avoiding the low efficiency problem caused by pairwise comparison of the Internet protocol address to be analyzed with a large number of Internet protocol addresses. In this way, both the traceability of the attack pattern and the accuracy of homology analysis are retained, and the efficiency of homology analysis is improved. In summary, the present application can improve the efficiency and accuracy of network attack homology analysis, which is of great significance for enhancing network security protection capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 is a schematic structural diagram of a network attack homology analysis system provided by an embodiment of the present application;
[0068] Figure 2 is a flowchart of the network attack homology analysis method provided by an embodiment of the present application;
[0069] Figure 3 is the overall flowchart of the network attack homology analysis method provided by an embodiment of the present application;
[0070] Figure 4 is a schematic diagram of the functional modules of the network attack homology analysis device provided by an embodiment of the present application;
[0071] Figure 5 is a schematic diagram of the hardware structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0072] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0073] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart in the flowchart. The terms "first", "second", etc. in the specification, claims and the above drawings are used to distinguish similar objects and do not have to be used to describe a specific order or sequence.
[0074] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0075] Network attack is an act of using computer network technology to illegally invade, interfere with, damage or steal information resources in a network system to achieve specific goals. It can include, but is not limited to, various forms such as data leakage, system damage, denial-of-service attack, and malicious software propagation, posing a serious threat to personal privacy, corporate interests, social order and security.
[0076] In order to more effectively cope with complex and changing network attacks and enhance the overall protection ability of network security, the correlation between different attack events can be identified through network attack homology analysis, and it can be judged whether they come from the same attack source or organization, so as to reveal the attacker's behavior patterns and technical means, improve the defense efficiency and reduce resource waste.
[0077] In the related art, the Internet Protocol address corresponding to each attack event is usually used as a node, and the feature similarity between all Internet Protocol addresses is compared pairwise, and the Internet Protocol addresses with high feature similarity are clustered to identify the homologous relationship between attackers. However, on the one hand, the method of clustering after pairwise comparison has a high computational complexity and low analysis efficiency when faced with a large amount of attack data; on the other hand, the clustering result lacks interpretability, resulting in security personnel being unable to effectively trace the attack pattern, and thus misjudging or missing the homologous relationship, reducing the accuracy of the analysis result.
[0078] Based on this, the embodiments of the present application provide a method, device, computer device and storage medium for analyzing the homology of network attacks. The present application proposes a method, device, computer device and storage medium for analyzing the homology of network attacks, which can improve the efficiency and accuracy of analyzing the homology of network attacks.
[0079] The method, device, computer device and storage medium for analyzing the homology of network attacks provided by the embodiments of the present application are specifically described through the following embodiments. First, the system for analyzing the homology of network attacks in the embodiments of the present application is described.
[0080] Please refer to Figure 1 , in some embodiments, the embodiments of the present application provide a system for analyzing the homology of network attacks, including a terminal 11 and a server side 12.
[0081] Exemplarily, the terminal 11 can be a network security detection device, a personal computer or a workstation, a mobile computing device, etc. The server side 12 can be a data center server, a cloud server, a dedicated computing cluster, etc.
[0082] Further, the terminal 11 can collect original attack data from the network environment and transmit the original attack data to the server side 12 through the network for further processing. When the server side 12 receives the data from the terminal 11, it can perform a series of preprocessing operations such as cleaning and feature extraction on it to obtain a plurality of preset Internet Protocol addresses, and construct a preset weighted attack behavior graph for the homologous relationship between the plurality of preset Internet Protocol addresses. Then, the preset weighted attack behavior graph is divided to obtain a plurality of target address clusters, and the feature centroid of each target address cluster is calculated. When the terminal 11 collects the Internet Protocol address to be analyzed and needs to perform homology analysis, it can send the Internet Protocol address to be analyzed to the server side 12, so that the server side 12 can calculate the corresponding network attack homology analysis result by calculating the Internet Protocol address to be analyzed and each feature centroid one by one, and feedback the network attack homology analysis result to the terminal 11.
[0083] Further, the server 12 can regularly update the preset weighted attack behavior graph, the target address cluster, and the algorithm model, and push the latest updates to the terminal 11 to ensure that the entire system can cope with new threats. At the same time, the terminal 11 can also send new data to the server 12 or request further analysis support.
[0084] The network attack homology analysis method in the embodiments of the present application can be illustrated by the following embodiments.
[0085] It should be noted that in each specific implementation manner of the present application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.
[0086] In the embodiments of the present application, the description will be made from the dimension of the network attack homology analysis device, which can be specifically integrated in a computer device. Refer to Figure 2 , Figure 2 which is the flowchart of the steps of the network attack homology analysis method provided by the embodiments of the present application. In the embodiments of the present application, taking the network attack homology analysis device being specifically integrated in a terminal or a server as an example, when the processor on the terminal or the server executes the program instructions corresponding to the network attack homology analysis method, the specific process is as follows:
[0087] Step 101, obtain multiple preset Internet protocol addresses and a preset weighted attack behavior graph, where the preset weighted attack behavior graph includes multiple preset Internet protocol addresses and multiple connection weights, and each connection weight represents the degree of association between two preset Internet protocol addresses with a homologous relationship.
[0088] In some implementation manners, in order to quantify and visualize the degree of association between different attack Internet protocol addresses, a preset weighted attack behavior graph constructed in advance according to multiple preset Internet protocol addresses can be obtained, and the multiple preset Internet protocol addresses can be processed to provide a visual framework to help security analysts identify potential attack groups and provide the necessary data basis for subsequent in-depth analysis.
[0089] Among them, the preset Internet Protocol address is the Internet Protocol (IP) address, which can be a set of Internet Protocol addresses predefined through historical network attack data. Each preset Internet Protocol address is a unique address used to identify a device on the network, and each Internet Protocol address is marked as a potential attack source. For example, the preset Internet Protocol address can be 192.168.1.100, 203.0.113.5, etc.
[0090] Among them, the preset weighted attack behavior graph can be a graphical structure constructed based on the relationships between attackers (preset Internet Protocol addresses). Among them, the nodes represent different preset Internet Protocol addresses, the edges represent the same-source relationships between the preset Internet Protocol addresses, and each edge has a connection weight.
[0091] Among them, the connection weight can be the numerical value carried by the edge between two preset Internet Protocol addresses in the preset weighted attack behavior graph, and is used to quantify the same-source degree between these two preset Internet Protocol addresses.
[0092] Among them, the correlation degree can be used to characterize the closeness of the relationship between different preset Internet Protocol addresses, and is an index calculated comprehensively based on various features (such as spatio-temporal features, behavior fingerprints, association graphs, etc.) included in multiple feature interaction dimensions between the preset Internet Protocol addresses, and is used to measure whether the preset Internet Protocol addresses belong to the same same-source attack group.
[0093] In some embodiments, the features of multiple feature interaction dimensions such as the spatio-temporal feature dimension, the behavior fingerprint dimension, and the association graph dimension between the preset Internet Protocol addresses can be vectorized through the preset set of Internet Protocol addresses obtained in advance (such as the Internet Protocol addresses included in the firewall interception records, the Internet Protocol addresses of the advanced persistent threat attacks captured by the honeypot, etc.), to obtain the first feature vector between two Internet Protocol addresses. For example, the first feature vector can be:
[0094] [0,1,0,0,0,1,0.72,0.33,0.82,1,0,1,0.33,1];
[0095] Among them, the feature interaction dimensions of each first eigenvalue in the above first feature vector are respectively the same B segment, autonomous system number (ASN), same city + network type, time Kullback-Leibler (KL) divergence, attack type Jaccard similarity index, attack payload text similarity, special pattern, password Jaccard similarity index, honeypot access interval, service overlap degree, and number of common domain names. The above first feature vector is only an example. In actual situations, the feature interaction dimensions and first eigenvalues may be different, and can be obtained according to the actual situation.
[0096] Furthermore, the initial gradient boosting model can be used to calculate the feature importance scores of each first eigenvalue of the multiple feature interaction dimensions included in the first feature vector, and then obtain the feature importance scores of each first eigenvalue in the determination of the homologous relationship between two preset Internet protocol addresses. According to the multiple feature importance scores of the multiple first eigenvalues corresponding to the first feature vector, the connection weight between the two preset Internet protocol addresses can be calculated.
[0097] Furthermore, by taking each preset Internet protocol address as the target node of the preset weighted attack behavior graph, taking the homologous relationship between any two Internet protocol addresses as the edge, and taking the corresponding connection weight as the weight of the edge, the preset weighted attack behavior graph can be constructed.
[0098] Through the above method, the preset weighted attack behavior graph can be obtained to visually present the complex association relationships between Internet protocol addresses, which helps security analysts quickly understand the overall architecture of attack behaviors and potential attack groups, helps with subsequent division of address clusters, and improves the efficiency of homologous analysis.
[0099] In some embodiments, in order to improve the accuracy of the homologous analysis results, a weighted attack behavior graph corresponding to multiple preset Internet protocol addresses (i.e., all the collected attack Internet protocol addresses) can be established to visually display the degree of closeness of the connections between different Internet protocol addresses, and then display the logical basis for homologous determination. For example, step 101 may include:
[0100] (101.1) Obtain multiple preset Internet protocol addresses and the homologous relationships between the multiple preset Internet protocol addresses, and determine the first feature vector and relationship label between any two preset Internet protocol addresses, where the first feature vector includes multiple first eigenvalues of any two preset Internet protocol addresses in multiple feature interaction dimensions;
[0101] (101.2) Input each first eigenvector and the corresponding relationship label into the initial gradient boosting model in sequence, and perform decision tree split gain accumulation through the initial gradient boosting model to obtain the feature importance score corresponding to each first eigenvalue;
[0102] (101.3)Based on multiple feature importance scores, determine the connection weights of two preset Internet protocol addresses corresponding to the first eigenvector;
[0103] (101.4)Based on multiple preset Internet protocol addresses, the relationship labels between any two preset Internet protocol addresses, and the corresponding connection weights, construct a preset weighted attack behavior graph corresponding to the multiple preset Internet protocol addresses.
[0104] Among them, the homologous relationship can include non - homologous and homologous. When the homologous relationship indicates homology between two preset Internet protocol addresses, it represents the relationship that the two preset Internet protocol addresses are considered to belong to the same attack group due to features such as similar behavior patterns and attack methods. Otherwise, the homologous relationship between the two Internet protocol addresses is considered non - homologous.
[0105] Among them, the first eigenvector can be a feature set representing the features between any two preset Internet protocol addresses, and this feature set contains the specific numerical representations of these two preset Internet protocol addresses in multiple feature interaction dimensions.
[0106] Among them, the relationship label can be a mark of whether there is homology between any two preset Internet protocol addresses, usually represented by 0 or 1 (1 indicates homology exists, 0 indicates non - homology).
[0107] Among them, the feature interaction dimension can be different angles or aspects for evaluating the relevance between two preset Internet protocol addresses, such as spatio - temporal features, behavior fingerprints, and association graphs. Among them, spatio - temporal features, behavior fingerprints, and association graphs can be further refined.
[0108] Among them, the first eigenvalue can be an element in the first eigenvector, representing the specific similarity between any two preset Internet protocol addresses in a specific feature interaction dimension.
[0109] Among them, the initial gradient boosting model can be a machine learning model based on the gradient boosting algorithm, which is used to learn the importance of each first eigenvalue for homology determination from the first feature vector and its corresponding relationship label, and optimize the model through the decision tree splitting gain accumulation process. Finally, the target gradient boosting model is obtained, and the feature importance score corresponding to each first eigenvalue is output according to the accumulated splitting gain. Exemplarily, the initial gradient boosting model can be a Light Gradient Boosting Machine (LightGBM) model, such as the LightGBM model.
[0110] Among them, the feature importance score can be a quantitative index calculated by the initial gradient boosting model to measure the importance degree of each first eigenvalue for determining the homology relationship between two preset Internet protocol addresses.
[0111] Exemplarily, if there are 3 preset Internet protocol addresses (only for example here, in fact, multiple or a large number of preset Internet protocol addresses can be used to construct a comprehensive preset weighted attack behavior graph): IP1 = 192.168.1.1, IP2 = 192.168.1.2, IP3 = 192.168.2.1, the homology relationships of these 3 preset Internet protocol addresses have been pre-labeled. IP1 and IP2 are homologous (thus determining the relationship label = 1), IP1 and IP3 are not homologous (thus determining the relationship label = 0), and IP2 and IP3 are not homologous (thus determining the relationship label = 0).
[0112] In some embodiments, the first feature vector between any two preset Internet protocol addresses can be obtained. Specifically, multiple first sub-features corresponding to the first preset Internet protocol address in multiple feature interaction dimensions, and multiple second sub-features corresponding to the second preset Internet protocol address in multiple feature interaction dimensions can be obtained, where the multiple feature interaction dimensions include a spatio-temporal feature dimension, a behavior fingerprint dimension, and an association graph dimension; then, the behavior overlap degree between each first sub-feature and the corresponding second sub-feature is obtained, and the behavior overlap degree is encoded to obtain the corresponding first eigenvalue. Based on the multiple first eigenvalues corresponding to the multiple feature interaction dimensions, the first feature vector between any two preset Internet protocol addresses can be obtained.
[0113] Exemplarily, an example is given for obtaining the first eigenvector between any two preset Internet Protocol (IP) addresses. Specifically, in the process of determining the behavior overlap degree between each first sub-feature and its corresponding second sub-feature, taking two preset IP addresses 192.168.1.1 and 192.168.1.2 as an example, during the example process, IP1 is denoted as IP1, and IP2 is denoted as IP2. The spatio-temporal features can characterize the attacker's behavior pattern from three dimensions: network layer features, geographical features, and time pattern features. Among them, the network layer features can include same C-segment features and same B-segment features, that is, to determine whether two preset IP addresses belong to the same subnet segment (for example, IP1 and IP2 are in the same C-segment). The network layer features can also include the autonomous system number to determine whether two preset IP addresses belong to the same autonomous system (for example, both IP1 and IP2 belong to AS12345), and can also include the IP nature (for example, IP1 is the preset IP address of a data center) to determine whether two preset IP addresses are of specific types such as data centers and enterprise dedicated lines; the geographical features can include the features obtained by performing correlation analysis by combining geographical location and network type (such as home broadband, enterprise dedicated line). For example, both IP1 and IP2 are located in City T and are enterprise dedicated line users; the time pattern features can be calculated through the similarity of the attack time distribution (for example, the KL divergence of the attack time distribution of two preset IP addresses). For example, the KL divergence of the attack time distribution of IP1 and IP2 is 0.2, indicating that their time patterns are very similar.
[0114] Exemplarily, the behavior fingerprint can be calculated from three dimensions: attack vector, attack payload features, and password features. Among them, the attack vector can include the Jaccard similarity of attack types (for example, IP1 and IP2 use the same SQL injection attack method, and its Jaccard similarity is 0.9) and the malicious domain name overlap degree (for example, both IP1 and IP2 use the malicious domain name "mmm.com"); the payload features can include text similarity (for example, the payload text similarity of IP1 and IP2 is 0.85) and special patterns (for example, both IP1 and IP2 contain the same URL parameter structure); the password features can include the Jaccard similarity of the top 100 passwords. For example, 70 out of the first 100 passwords tried by IP1 and IP2 are the same, and the Jaccard similarity is 0.7.
[0115] Further, the associated graph dimension can be determined from three dimensions: honeypot linkage, service association, and domain name mapping. Among them, honeypot linkage can be determined by identifying that the access interval to the same honeypot is less than <X> hours. If the time interval between IP1 and IP2 accessing the same honeypot is less than 2 hours, it indicates that there may be a connection between the two. Service association can be determined by the overlap degree of the attacked service system sets. For example, both IP1 and IP2 attacked the information systems of Bank A and Hospital B. Domain name mapping can be determined by the number of co-occurring domain names in threat intelligence. For example, both IP1 and IP2 are associated with the domain name "mx.com".
[0116] Further, after obtaining the behavior overlap degree of the first sub-feature and the second sub-feature, the behavior overlap degree can be encoded in the following ways. For numerical features (such as KL divergence of time distribution, Jaccard similarity, etc.), the original values can be directly used. For categorical features (such as IP nature, ASN number, etc.), One-Hot encoding can be adopted. For text features (such as payload text, etc.), the cosine similarity can be calculated after vectorization based on TF-IDF. For boolean features (such as in the same C segment, in the same B segment, etc.), they can be converted into binary values (1 means yes, 0 means no, etc.).
[0117] For example, when encoding, in the same C segment: 1 (in the same C segment); ASN number: 1 (belonging to the same autonomous system AS12345); IP nature: 1 (both are enterprise dedicated lines); Jaccard similarity of attack type: 0.9 (using the same method for injection attacks); malicious domain name overlap degree: 1 (both use "mmm.com"); text similarity: 0.85 (high payload text similarity); special mode: 1 (both use "username=admin&password=1234"); access interval to the same honeypot: 1 (the time interval for accessing the same honeypot is less than 2 hours); overlap degree of the attacked service system sets: 1 (both attacked Bank A and Hospital B); number of co-occurring domain names in threat intelligence: 1 (both are associated with "badactor.com"). Thus, the first feature vector can be obtained as [1, 1, 1, 0.9, 1, 0.85, 1, 1, 1, 1].
[0118] In some embodiments, a training set for the initial gradient boosting model can be constructed. The training set can include positive samples and negative samples. Among them, the positive samples are the first feature vectors with a confidence level > 0.8 for two preset Internet protocol addresses of the same origin and with correct manual review, and the label , and the negative samples are randomly sampled pairs of two preset Internet protocol addresses and cases of false positives by rules, and the label 0;
[0119] Then the training set can be expressed as:
[0120] ;
[0121] Among them, represents the first eigenvector of two preset Internet protocol address pairs (IPi, IPj), represents the relationship label of the IP pair (1 indicates the same origin, 0 indicates different origins).
[0122] Furthermore, an initial gradient boosting model (such as the LightGBM model) can be used for training to learn the contribution degree of each first eigenvalue to the homologous determination result. After the initial gradient boosting model is trained, a target gradient boosting model can be obtained. At the same time, the model will output the feature importance score corresponding to each first eigenvalue , representing the weight corresponding to the k-th first eigenvalue.
[0123] In some embodiments, for each first eigenvector, each first eigenvalue it contains can be multiplied by the corresponding feature importance score to obtain the sub-feature weight of each first eigenvalue, and based on the sum of the multiple sub-feature weights corresponding to the multiple first eigenvalues, the connection weight of the two preset Internet protocol addresses corresponding to the first eigenvector can be obtained. In this way, the association strength between any two preset Internet protocol addresses can be accurately reflected.
[0124] Furthermore, a preset weighted attack behavior graph can be constructed based on all preset Internet protocol addresses as target nodes, the relationship label (i.e., the same origin or different origins) between any two preset Internet protocol addresses as the connection edge between the target nodes, and the corresponding connection weight as the weight of the connection edge. In this way, the association strength between preset Internet protocol addresses can be quantified, and potential homologous attack groups can be intuitively displayed, thereby helping security analysts to more efficiently and accurately identify and understand complex network attack patterns and improve the overall network security protection ability.
[0125] By constructing a preset weighted attack behavior graph, a comprehensive and detailed attack behavior data basis can be provided for subsequent analysis, ensuring the accuracy and comprehensiveness of the analysis, and also providing strong support for the formulation of network security defense strategies, helping to detect and respond to potential network security threats in a timely manner.
[0126] In some embodiments, to improve the accuracy of the model in identifying the homologous relationship of network attacks, the initial gradient boosting model can be trained to accurately quantify the importance of the first eigenvalue of each first feature vector in different feature interaction dimensions in determining the homologous relationship between two preset Internet protocol addresses, so that the model can continuously improve the rationality of the assigned weights, thereby facilitating the traceability of homologous analysis and further enhancing the accuracy and reliability of homologous analysis. For example, (101.2) may include:
[0127] (101.2.1) Sequentially input each first feature vector and the corresponding relationship label into the initial gradient boosting model, and through the initial gradient boosting model, predict each first feature vector to obtain the corresponding predicted label;
[0128] (101.2.2) Determine the first target loss based on the difference between the predicted label and the relationship label;
[0129] (101.2.3) Based on the first target loss, calculate the split gain for the multiple first eigenvalues included in each first feature vector, and determine the intermediate split feature with the largest split gain from the multiple first eigenvalues for node splitting, and calculate the second target loss;
[0130] (101.2.4) Based on the difference between the second target loss and the first target loss, calculate the split gain for the multiple first eigenvalues included in the first feature vector, and determine the intermediate split feature with the largest split gain from the multiple first eigenvalues for node splitting, and calculate the updated second target loss;
[0131] (101.2.5) Repeat the step of calculating the split gain for the multiple first eigenvalues included in the first feature vector based on the difference between the updated second target loss and the first target loss, and determining the intermediate split feature with the largest split gain from the multiple first eigenvalues for node splitting, and calculating the updated second target loss until the preset number of training times is reached. According to the multiple split gains iteratively obtained for each eigenvalue, obtain the feature importance score corresponding to each first eigenvalue.
[0132] Among them, the predicted label can be the label obtained by predicting each first feature vector through the initial gradient boosting model, and is used to indicate whether there is a homologous relationship (such as 0 or 1) between the two preset Internet protocol addresses judged by the model.
[0133] Among them, the first target loss can be used to measure the error between the current prediction result (predicted label) of the model and the actual situation (relationship label).
[0134] Among them, the split gain can be the information gain or the reduced loss value brought by the feature value when the intermediate split feature is selected as the split point in the decision tree. The larger the split gain, the more effectively this feature value can distinguish data of different categories, that is, the larger the feature importance score of the corresponding intermediate split feature.
[0135] Among them, the intermediate split feature can be the feature value that can maximize the split gain calculated according to the split gain during each node splitting process, that is, the optimal split point.
[0136] Among them, the second objective loss can be the loss value after each node splitting, which reflects the performance of the model after using the selected intermediate split feature, is used to evaluate the effect of splitting, and provides a basis for further optimization.
[0137] Exemplarily, the model can be initialized first. For example, the maximum tree depth of the model can be set to 8, the minimum number of samples in the leaf node is 10, and the loss function is set to the logarithmic loss function. Specifically, in the first round of iteration, the initial gradient boosting model (generally a single tree) calculates the prediction probability for each first feature vector:
[0138] ;
[0139] Among them, is the Sigmoid function, is the output of the t-th tree.
[0140] Furthermore, if the initial gradient boosting model outputs the prediction label as 0.75 according to the current first feature vector, it means that there is a 75% probability of determining homology. At this time, the difference between the prediction label and the true relationship label can be quantified by calculating the first objective loss to drive the optimization of the initial gradient boosting model. Exemplarily, the first objective loss and the second objective loss can be calculated using loss functions such as logarithmic loss and cross-entropy loss. This application does not limit this too much.
[0141] Furthermore, after the initial gradient boosting model is predicted, if the predicted label predicted by the initial gradient boosting model is significantly different from the relationship label, or the training rounds have not reached the preset rounds, the feature and split point with the largest split gain can be selected through a greedy strategy to gradually optimize the tree structure. Specifically, the value of each feature can be discretized into a histogram. For example, the first eigenvalue corresponding to the feature interaction dimension of time KL divergence can be divided into intervals [0, 0.1), [0.1, 0.2), etc., and the split gain of each candidate split point (that is, the feature interaction dimension that can be used for splitting) is calculated, and all first eigenvalues and their split points are traversed to select the splitting method with the largest gain. For example, the splitting gain of the feature with the same C segment is 12.5, and the splitting gain of the feature time KL divergence is 8.3. Then, the segment with the largest gain can be selected for splitting, and the tree structure is updated and the second target loss after the split is calculated.
[0142] Furthermore, each new tree can be fitted based on the residual of the previous model to gradually approach the true label, and the second target loss after the split is compared with the first target loss. If the second target loss is less than the first target loss, and the difference between the second target loss and the first target loss is greater than the preset threshold, the split will continue; otherwise, early stopping will be triggered.
[0143] For example, after the first split, the loss dropped to 0.58. The second split selected the time KL divergence as the intermediate split feature. Specifically, the time KL divergence > 0.2, and the loss dropped to 0.52. The iteration continued until the loss converged or the maximum number of trees was reached. At this point, the sum of the split gains of each feature interaction dimension (that is, the feature interaction dimension corresponding to each first eigenvalue) in all decision trees can be accumulated and normalized to obtain the importance score corresponding to each first eigenvalue.
[0144] Through the above method, the model can automatically learn the contribution of the feature interaction dimension corresponding to each first eigenvalue to homology judgment, thereby providing a scientific feature weight distribution for network attack homology analysis, which is convenient for subsequent homology analysis of network attacks and improves the accuracy and reliability of the analysis results.
[0145] In some embodiments, in order to enhance the rationality and interpretability of the connection weight calculation, the sub-feature weight can be calculated by multiplying each first feature value by the corresponding target feature importance score, quantifying the specific impact of a single feature in the connection weight, and finally determining the connection weight between two preset Internet protocol addresses, so as to ensure the fairness and rationality of the feature weight allocation and improve the reliability and interpretability of the analysis results. For example, (101.3) may include:
[0146] (101.3.1) Normalize each feature importance score to obtain the corresponding target feature importance score;
[0147] (101.3.2) Obtain the sub - feature weight of each first eigenvalue according to the product of each first eigenvalue and the corresponding target feature importance score.
[0148] (101.3.3) Based on the sum of the multiple sub - feature weights corresponding to the multiple first eigenvalues, obtain the connection weight between the two preset Internet protocol addresses corresponding to the first eigenvector.
[0149] Among them, the target feature importance score can be the result obtained by normalizing the feature importance score of each first eigenvalue. The purpose is to enable the importance scores of different features to be compared on the same scale, so as to more accurately evaluate the relative importance of each feature in the model.
[0150] Among them, the sub - feature weight can be the product of each first eigenvalue and its corresponding target feature importance score, indicating the proportion of this eigenvalue in determining the connection strength between the two preset Internet protocol addresses.
[0151] In some embodiments, the feature importance score of each feature can be normalized through the following formula to obtain the corresponding target feature importance score:
[0152] ;
[0153] Among them, represents the feature importance score of the k - th first eigenvalue, represents the normalized target feature importance score, and n represents the total number of first eigenvalues included in the first eigenvector.
[0154] In some embodiments, the sub - feature weight, that is, the feature contribution degree of the corresponding first eigenvalue in the homologous determination of the first eigenvector, can be calculated through the following formula:
[0155] ;
[0156] Among them, represents the feature contribution degree (that is, the sub - feature weight) of the k - th feature, represents the normalized target feature importance score, represents the corresponding first eigenvalue. By calculating the feature contribution degree corresponding to each first eigenvalue, it is convenient to trace the results and perform visual analysis during the subsequent process of network attack homologous analysis.
[0157] Further, the sub - feature weight (that is, the feature contribution degree) can be normalized:
[0158] ;
[0159] wherein, represents the sub-feature weight of the i-th feature, and n represents the total number of the first eigenvalues, which is also the total number of the sub-feature weights.
[0160] Furthermore, based on the sum of the normalized sub-feature weights corresponding to multiple first eigenvalues, the connection weight between two preset Internet protocol addresses corresponding to the first eigenvector can be obtained:
[0161] ;
[0162] represents the normalized sub-feature weight.
[0163] Through the above method, the contribution degree of different first eigenvalues to determining the homology relationship between two preset Internet protocol addresses can be effectively quantified, so as to obtain a comprehensive connection weight. In this way, not only the accuracy of homology analysis is improved, but also the interpretability of the preset weighted attack behavior graph is enhanced, which is convenient for result tracing and visual analysis.
[0164] In some embodiments, in order to update the target relationship label in a timely manner according to the latest data, the first dynamic threshold can be determined by introducing statistics such as historical connection weights, scoring means, and scoring standard deviations, and flexibly using adjustment parameters, so as to obtain the homology relationship between any two preset Internet protocol addresses. In this way, it can be ensured that even in the case of rapid changes in the network environment, the effective tracking and analysis of network attack behaviors can be maintained, and the robustness and adaptability of the overall network security defense system are enhanced. Exemplarily, before constructing the preset weighted attack behavior graph corresponding to multiple preset Internet protocol addresses, that is, before (101.4), it may further include:
[0165] (A.1) Obtain multiple historical connection weights between any two preset Internet protocol addresses, as well as the scoring mean and scoring standard deviation of the multiple historical connection weights;
[0166] (A.2) Obtain a preset adjustment parameter, and obtain a first product according to the product of the adjustment parameter and the scoring standard deviation;
[0167] (A.3) Determine the first dynamic threshold according to the difference between the scoring mean and the first product;
[0168] (A.4) Compare the connection weight between any two preset Internet protocol addresses with the first dynamic threshold to obtain a comparison result;
[0169] (A.5) Based on the comparison result, update the relationship label between any two preset Internet protocol addresses to obtain the target relationship label;
[0170] Then, based on multiple preset Internet protocol addresses, the relationship labels between any two preset Internet protocol addresses, and the corresponding connection weights, construct a preset weighted attack behavior graph corresponding to the multiple preset Internet protocol addresses, including:
[0171] Based on the multiple preset Internet protocol addresses, the target relationship labels between any two preset Internet protocol addresses, and the corresponding connection weights, construct a preset weighted attack behavior graph corresponding to the multiple preset Internet protocol addresses.
[0172] Among them, the historical connection weight can be the connection weight calculated for two preset Internet protocol addresses at multiple past time points.
[0173] Among them, the average score can be the average value of multiple historical connection weights, used to represent the general level or central tendency of the historical connection weights.
[0174] Among them, the standard deviation of the score can be used to measure the degree of dispersion of the distribution of multiple historical connection weights, and is used to reflect the fluctuation of the historical connection weights around the average score.
[0175] Among them, the adjustment parameter can be a preset parameter used to adjust the sensitivity of the dynamic threshold. After multiplying the adjustment parameter by the standard deviation of the score, it can control the range of the first dynamic threshold to adapt to the requirements of different application scenarios.
[0176] Among them, the first product can be the result obtained by multiplying the adjustment parameter by the standard deviation of the score, used to provide a flexibility factor when determining the first dynamic threshold.
[0177] Among them, the first dynamic threshold can be a threshold obtained based on the difference between the average score and the first product, used to determine whether the current connection weight significantly deviates from the historical average level, so as to decide whether to update the relationship label.
[0178] Among them, the comparison result can be the result of comparing the current connection weight between any two preset Internet protocol addresses with the first dynamic threshold, indicating whether the relevance between the two preset Internet protocol addresses has changed significantly.
[0179] Among them, the target relationship label can be a new label (such as 0 or 1) indicating whether there is a homologous relationship between any two preset Internet protocol addresses updated according to the comparison result, used to reflect the latest analysis conclusion.
[0180] In some embodiments, the first dynamic threshold has the following calculation formula:
[0181] ;
[0182] wherein, represents the average score, represents the standard deviation of the score, represents the adjustment parameter.
[0183] Exemplarily, if the multiple historical connection weights between any two preset Internet protocol addresses (such as IP1 and IP2) include: w1 = 0.8, w2 = 0.7, w3 = 0.9, w4 = 0.6, w5 = 0.8. Then, the average score can be calculated as 0.76, and the standard deviation of the score is 0.098.
[0184] Further, if the adjustment parameter is set to 2, according to the product of the preset adjustment parameter and the standard deviation of the score, the first product = 2 × 0.098 = 0.196 can be obtained. Then, according to the difference between the average score and the first product, the first dynamic threshold = 0.76 - 0.196 = 0.564 can be determined.
[0185] Exemplarily, the first dynamic threshold calculated from these historical connection weights can be compared with the connection weight between its corresponding two preset Internet protocol addresses to determine whether it exceeds the first dynamic threshold. Assume that the current connection weights of IP1 and IP2 are 0.7, 0.7 > 0.564, then the comparison result is exceeded, and the relationship label is updated to homologous (relationship label is 1); otherwise, it is updated to non - homologous (label is 0).
[0186] It should be noted that the method for determining the relationship label of two Internet protocol addresses in (A.1) to (A.5) can also be applied to "obtaining multiple preset Internet protocol addresses and the homologous relationship between multiple preset Internet protocol addresses" in (101.1). When it is necessary to determine the relationship label of two Internet protocol addresses at any time, the above formula can be used to determine the homologous relationship. When there is no historical connection weight, the first dynamic threshold for determining the relationship label between two preset Internet protocol addresses can be set manually or through other algorithms, and this application does not limit it too much.
[0187] Through the above method, the threshold can be dynamically adjusted and the relationship label can be updated in real - time to adapt to the changes in network attack behaviors, improving the accuracy and flexibility of homologous analysis.
[0188] Step 102, divide the multiple preset Internet protocol addresses in the preset weighted attack behavior graph into multiple initial address clusters.
[0189] In some embodiments, in order to enable the system to quickly identify pre-set Internet protocol addresses with close connections, that is, pre-set Internet protocol addresses that may belong to the same attack group, multiple pre-set Internet protocol addresses in the pre-set weighted attack behavior graph can be first divided into multiple initial address clusters, so as to facilitate subsequent continuous optimization of each cluster by calculating modularity, and then enable the subsequent community discovery algorithm to operate more efficiently.
[0190] Among them, the initial address clusters can be multiple groups or clusters initially divided from multiple pre-set Internet protocol addresses according to the connection weights and feature similarities of the pre-set Internet protocol addresses.
[0191] In some embodiments, the initial address clusters can also be obtained by randomly dividing the pre-set weighted attack behavior graph.
[0192] Exemplarily, each target node (that is, the pre-set Internet protocol address) in the pre-set weighted attack behavior graph can be divided according to the similarity relationship between any two pre-set Internet protocol addresses in the attack behavior feature library. For example, if two pre-set Internet protocol addresses have similar features in the attack behavior feature library, then these two pre-set Internet protocol addresses can be divided into the same initial address cluster.
[0193] In some embodiments, the attack behavior feature library mainly includes three parts: spatio-temporal features, behavior fingerprints, and association graphs. Their meanings have been introduced in detail when introducing the first feature vector between two pre-set Internet protocol addresses above, and will not be elaborated here one by one. Next, the process of constructing the attack behavior feature library will be briefly introduced. First, the spatio-temporal features can depict the behavior patterns of attackers from the network layer (such as the same C segment, the same B segment, ASN number, etc.), geographical features (such as city + network type, etc.), and time patterns (such as KL divergence of attack time distribution, etc.); second, the behavior fingerprints describe the specific technical means and habits of attackers by analyzing attack vectors (such as Jaccard similarity of attack types, malicious domain name overlap), payload features (such as text similarity, special patterns), and password features (such as Jaccard similarity of simple passwords); finally, the association graphs reveal the relationships between attackers from three dimensions: honeypot linkage (such as access interval of the same honeypot point), service association (overlap degree of the attacked service system set), and domain name mapping (number of domain names that appear together in threat intelligence). The process of constructing this attack behavior feature library depends on the experience design of network security experts to ensure its high interpretability and adaptability.
[0194] In some embodiments, multiple initial address clusters may also be obtained by partitioning based on the connection weights of multiple preset Internet Protocol (IP) addresses in a preset weighted attack behavior graph. Specifically, when the connection weight between any two preset IP addresses in the preset weighted attack behavior graph is greater than a preset threshold, these two preset IP addresses may be partitioned into the same initial address cluster. The preset threshold may be set to 0.6, 0.7, etc.
[0195] In the above manner, the initial address clusters can be determined as the starting point for community evolution, facilitating subsequent updates of the address clusters.
[0196] Step 103: For each initial address cluster, determine the connection weights associated with each preset IP address included therein, and calculate the corresponding modularity according to the connection weights. The modularity is used to characterize the degree of homology and compatibility among the multiple preset IP addresses included in the corresponding initial address cluster.
[0197] In some embodiments, in order to enable security analysts to more accurately locate specific attack groups, the modularity can be used as an evaluation criterion to identify address clusters with high internal connectivity and consistent behavior patterns, or to identify clusters with poor compatibility, so as to dynamically adjust and optimize the cluster structure, thereby improving the accuracy of network attack homology analysis.
[0198] Among them, the modularity can be an index used to quantify the connection strength and internal correlation among the preset IP addresses within each initial address cluster. A high modularity means a stronger homology relationship among the initial IP addresses within the initial address cluster.
[0199] Among them, the degree of homology and compatibility can be the degree of mutual compatibility and consistency shown among all preset IP addresses based on the connection weights in a specific initial address cluster. That is to say, the degree of homology and compatibility is used to characterize whether the preset IP addresses in the same initial address cluster have similar behavior patterns or attack methods, and thus the possibility of being classified into the same attack group.
[0200] In some embodiments, for each preset Internet Protocol address (target node) in the preset weighted attack behavior graph, the sum of all connection weights with other target nodes in the graph (i.e., the node degree) can be statistically calculated. The specific method for obtaining the connection weights has been elaborated above and will not be repeated here. After that, the graph connection weight is calculated, that is, half of the sum of all connection weights in the entire preset weighted attack behavior graph (to avoid double counting of weights) is used as the normalization reference value. Thus, for each initial address cluster, the differences between the following two situations can be compared: the actual connection weights between any two target nodes within the same initial address cluster, and the expected connection weights between target nodes within the same initial address cluster assuming that the network connections are completely random (determined by the product of the node degrees of the two target nodes and the ratio of the graph connection weight). Thus, the differences between the actual connection weights and the expected connection weights in each initial address cluster can be accumulated and then divided by the graph connection weight to obtain the modularity corresponding to the initial address cluster.
[0201] By calculating the modularity of each initial address cluster, the tightness of the cluster division can be quantified, providing a basis for subsequent optimization (such as node membership adjustment), so as to better understand and analyze the homology of network attacks.
[0202] In some embodiments, in order to accurately quantify the connection strength and internal consistency between preset Internet Protocol addresses within a cluster, the modularity of each address cluster obtained by division can be calculated, so that the system can more accurately divide the community structure in the network, and then iteratively obtain potential attack groups with high homology compatibility. For example, "calculating the corresponding modularity according to the connection weight" in step 103 may include:
[0203] (103.1) Obtaining the graph connection weight according to the sum of the connection weights corresponding to all connection edges in the preset weighted attack behavior graph;
[0204] (103.2) Obtaining the node degree corresponding to the target node corresponding to each preset Internet Protocol address in the initial address cluster;
[0205] (103.3) Obtaining a second product according to the product of the node degrees between any two target nodes;
[0206] (103.4) Obtaining a first ratio based on the ratio of the second product to the graph connection weight;
[0207] (103.5) Obtaining the difference between the connection weight corresponding to any two target nodes in the initial address cluster and the corresponding first ratio to obtain a first difference;
[0208] (103.6) Based on the graph connection weights and the multiple first differences between multiple target nodes included in the initial address cluster, obtain the modularity corresponding to the initial address cluster.
[0209] Among them, the graph connection weight can be the sum of the connection weights corresponding to all connection edges in the preset weighted attack behavior graph, which reflects the total association strength between all Internet protocol addresses in the entire graph.
[0210] Among them, the node degree can be the sum of the connection weights representing each preset Internet protocol address between the target node corresponding to the initial address cluster and all other target nodes, and is used to measure the importance of this target node in the network.
[0211] Among them, the second product can be the result obtained based on the product of the node degrees between any two target nodes, which reflects the relative importance of these two target nodes in the network and their possible degree of mutual influence.
[0212] Among them, the first ratio can be a value obtained based on the ratio of the second product to the graph connection weight, which represents the expected connection strength between any two target nodes, that is, in the random network model, the connection weight that these two target nodes should have.
[0213] Among them, the first difference can be the result obtained by obtaining the difference between the connection weight corresponding to any two target nodes in the initial address cluster and the corresponding first ratio, which is used to characterize the difference between the actual connection weight and the expected connection weight, and is used to determine whether the connection between these two target nodes is significantly stronger or weaker than the expected value in the random case.
[0214] In some embodiments, the calculation formula of the modularity is as follows:
[0215] ;
[0216] Among them represents the connection weight of the connection edge between target node i and target node j, represents the node degree of target node i, represents the sum of all connection weights in the preset weighted attack behavior graph, is an indicator function, which takes the value of 1 when target node i and target node j belong to the same address cluster, and 0 otherwise.
[0217] In some embodiments, , that is, half of the sum of all connection weights. This formula can be used to calculate the global benchmark value for subsequent normalization. Assuming that there are 3 edges in the graph with weights of 0.8, 0.5, and 0.3 respectively, then m = (0.8 + 0.5 + 0.3) / 2 = 0.8.
[0218] Exemplarily, to calculate the node degree corresponding to each target node, it can be done by summing up all the adjacency connection weights of the target node, that is . For example, if the target node A is connected to the target nodes B and C, and the connection weights are 0.8 and 0.5 respectively, then the node degree of the target node A is 0.8 + 0.5 = 1.3.
[0219] Furthermore, the product of the node degrees between any two target nodes in the initial address cluster can be calculated. For example, if the node degree of the target node A is 1.3 and the node degree of the target node B is 0.8, then the product of the node degrees of the target node A and the target node B, that is, the second product is 1.3 × 0.8 = 1.04.
[0220] Furthermore, the first weight That is, the proportion of the connection weight of the expected random connection. For example, if the second product is 1.04 and the graph connection weight m = 0.8, then the first ratio = 1.04 / (2 × 0.8) = 0.65. Then, the deviation between the actual connection strength between target nodes and the random expectation can be measured by the difference between the connection weight of any two target nodes in the preset weighted attack behavior graph and the first ratio. For example, if the actual connection weight between the target node A and the target node B is 0.8 and the first ratio is 0.65, then the first difference = 0.8 - 0.65 = 0.15.
[0221] In some embodiments, the first differences of all the connection edges (that is, the edges corresponding to any two target nodes) in the same initial address cluster can be accumulated and normalized by dividing by 2m to obtain the modularity corresponding to the initial address cluster. For example, if the initial address cluster includes the target nodes A, B, and C, if the first difference between the target node A and the target node B is 0.15, the first difference between the target node A and the target node C is, and the first difference between the target node B and the target node C is -0.1, then the modularity Q = (0.15 + 0.2 - 0.1) / (2 × 0.8) ≈ 0.156.
[0222] Furthermore, if the modularity > 0, it indicates that the connection strength within the address cluster is higher than the random expectation (such as Q = 0.15 indicating that the address cluster has significant homologous characteristics); if Q is close to 0 or negative, re - partitioning is required (such as the difference of the B - C edge being negative, which may belong to noise, and the address clusters to which the target nodes B and C belong can be re - partitioned).
[0223] In the above manner, the structural features of the preset weighted attack behavior graph can be transformed into multiple quantifiable modularities, providing theoretical support for dynamic cluster division and attack homology determination, facilitating more accurate identification and analysis of network attack behaviors, and ultimately achieving the goal of accurately mining potential attack groups from massive data.
[0224] Step 104: According to the modularity corresponding to each initial address cluster, re-determine the affiliation relationship between each preset Internet protocol address and the initial address cluster to obtain multiple target address clusters, and determine the characteristic centroid of each target address cluster.
[0225] In some embodiments, in order to form more optimized and accurate target address clusters, the affiliation relationship between the preset Internet protocol addresses and these clusters can be re-evaluated and adjusted based on the modularity corresponding to each initial address cluster, so that the system can more precisely identify groups of attack Internet protocol addresses with high homology, and ensure stronger relevance and consistency among the Internet protocol addresses within each address cluster, thereby improving the efficiency and accuracy of the entire network attack homology analysis.
[0226] Among them, the target address cluster can be a more optimized set of Internet protocol addresses (also referred to as a community) formed after re-determining the affiliation relationship between each preset Internet protocol address and the initial address cluster according to the modularity. The Internet protocol addresses within each target address cluster are considered to have a high degree of homology compatibility, that is, there are significant similarities in their behavior patterns or attack techniques, belonging to the same attack group.
[0227] Among them, the characteristic centroid can be the center point or average representative of the characteristic values of all preset Internet protocol addresses within each target address cluster. It is a comprehensive characteristic vector obtained by weighted averaging the characteristic vectors of all Internet protocol addresses within the target address cluster, and is used to describe the main characteristic attributes of the entire target address cluster.
[0228] In some embodiments, the affiliation relationship of the preset Internet protocol addresses can be dynamically adjusted based on the principle of modularity maximization, so that the internal connection density of the address cluster is significantly higher than the expected random network, and the characteristic centroid is extracted to support subsequent homology determination.
[0229] In some embodiments, when calculating the modularity of each initial address cluster, it is possible to determine the preset Internet protocol addresses for which the first difference between any two target nodes in the initially calculated address cluster is negative, and re-partition the home address clusters of these preset Internet protocol addresses. Specifically, after re-partitioning these preset Internet protocol copies into other initial address clusters, the modularity is recalculated. In this way, the efficiency and accuracy of address cluster partitioning can be improved. Alternatively, when the modularity of an initial address cluster is relatively low (e.g., lower than a preset modularity threshold), the address clusters to which each preset Internet protocol address in the initial address cluster belongs can be re-partitioned until the current modularity obtained from the partitioning is lower than the modularity threshold, at which point the re-partitioning process can be stopped to obtain the partitioned target address clusters.
[0230] In some embodiments, after partitioning the target address clusters, for each target address cluster, the eigenvectors of the multiple target address clusters it contains can be extracted and weighted averaged to obtain the eigen-centroid of each target address cluster. When a new Internet protocol address is added to a target address cluster or the community structure changes, only the centroid of the affected target address cluster needs to be updated, without full-scale calculation.
[0231] Through the above method, the homologous association strength (modularity) of the preset Internet protocol addresses within the address cluster can be maximized, and by calculating the eigen-centroid as the cluster fingerprint of each target address cluster, it is convenient to quickly and accurately match the Internet protocol addresses to be analyzed subsequently, so as to perform accurate homologous analysis on each Internet protocol address to be analyzed.
[0232] In some embodiments, to ensure that each finally formed target address cluster has a relatively high modularity, that is, there is a stronger association and consistency among internal members, the initial address clusters can be iteratively optimized until the modularity of all address clusters exceeds a second dynamic threshold, forming a more accurate target address cluster. In this way, not only the accuracy of identifying potential attack groups is improved, but also the adaptability and response speed of the system are enhanced, providing more solid support for network security protection. For example, "According to the modularity corresponding to each initial address cluster, re-determine the affiliation relationship between each preset Internet protocol address and the initial address cluster to obtain multiple target address clusters" in step 104 may include:
[0233] (104.a1) According to the modularity corresponding to each initial address cluster, determine the intermediate address clusters whose modularity is less than the preset second dynamic threshold;
[0234] (104.a2) Obtain a preset attack behavior feature library, and based on the attack behavior feature library, re-determine the associated Internet protocol addresses that have an attack feature relationship with each preset Internet protocol address in each intermediate address cluster, and migrate each preset Internet protocol address to the initial address cluster corresponding to the associated Internet protocol address;
[0235] (104.a3) Repeatedly execute the steps of re-determining the associated Internet protocol addresses that have an attack feature relationship with each preset Internet protocol address in each intermediate address cluster based on the attack behavior feature library, and migrating each preset Internet protocol address to the initial address cluster corresponding to the associated Internet protocol address, until the modularity corresponding to each intermediate address cluster is greater than the second dynamic threshold, to obtain multiple target address clusters.
[0236] Among them, the second dynamic threshold can be a threshold for evaluating whether an address cluster needs to be further adjusted. When the modularity of an address cluster is lower than the second dynamic threshold, it indicates that the association between the preset Internet protocol addresses within the address cluster is not strong enough, and the cluster structure needs to be optimized by reallocating the preset Internet protocol addresses it contains.
[0237] Among them, the intermediate address cluster can be an address cluster that still needs to be further divided.
[0238] Among them, the attack behavior feature library can be a data set containing various network attack behavior features. Based on the attack behavior feature library, the similarity and potential homologous relationship between different preset Internet protocol addresses can be analyzed and identified. The content of the attack behavior feature library has been introduced above and will not be elaborated here.
[0239] Among them, the attack feature relationship can describe the similarity or association shown between two preset Internet protocol addresses based on specific attack behavior features. For example, by querying the attack behavior feature library, it can be determined that the password features and domain names of two preset Internet protocol addresses overlap, etc. Then, one of the two preset Internet protocol addresses can be divided into the address cluster where the other one is located.
[0240] Among them, the associated Internet protocol address can be another preset Internet protocol address in other address clusters (such as the initial address cluster or the intermediate address cluster) that has an attack feature relationship with the preset Internet protocol address to be divided.
[0241] In some embodiments, the second dynamic threshold can be set according to the actual situation.
[0242] Exemplarily, when the modularity of the initial address cluster is less than a preset second dynamic threshold, for example, the modularity of the initial address cluster is -0.1 while the second dynamic threshold is 0.1, then it is determined that the initial address cluster is an intermediate address cluster for re-partitioning. For example, according to the association rules in the attack behavior feature library (such as "same-source attack features" including sharing the same C2 server, consistent Payload hash, etc.), the association strength between each preset Internet protocol address in the intermediate address cluster and other address clusters (which may include intermediate address clusters and initial address clusters) can be re-evaluated. If the rules of the attack behavior feature library are satisfied, for example, IP1 in intermediate address cluster A and IP2 in intermediate address cluster B belong to the same subnet segment and the same autonomous system, then IP1 can be tried to be migrated to intermediate address cluster B, and the modularity of intermediate address cluster A and intermediate address cluster B can be recalculated.
[0243] Furthermore, if after a preset Internet protocol address is migrated to a new address cluster, the modularity of the address cluster it belongs to decreases, then the address cluster to which the preset Internet protocol address belongs can be re-determined.
[0244] In some embodiments, for each preset Internet protocol address in each initial address cluster, the modularity can also be calculated after iterative movement among multiple address clusters. If it contributes to the modularity of the migrated address cluster, for example, after migrating IP1 from address cluster A to address cluster B, both the modularity of address cluster A and address cluster B increase, then it can be determined that the target address cluster corresponding to IP1 is address cluster B; otherwise, continue to try to migrate IP1.
[0245] Furthermore, it can also be stopped iterating for each address cluster until the modularity of all address clusters is greater than the second dynamic threshold, and multiple target address clusters are obtained. If there is a new preset Internet protocol address that does not contribute to the modularity of any address cluster (or reduces the modularity of the corresponding address cluster after joining any address cluster), then the new preset Internet protocol address can be used as a separate target address cluster.
[0246] It should be noted that for other preset Internet protocol addresses, the above method can also be used to determine the target address cluster of each preset Internet protocol address, so as to improve the accuracy of target address cluster partitioning, and further improve the efficiency, accuracy, and interpretability of network attack same-source analysis.
[0247] In some embodiments, when a newly added preset Internet Protocol (IP) address is joined, instead of recalculating the entire graph, a local address cluster is recalculated through an incremental algorithm to avoid reconstructing the entire graph. For example, when there is a new preset IP address (such as IPx), its related IP addresses with attack feature relationships can be determined through an attack behavior feature library, such as IP1 and IP2. If IP1 corresponds to target address cluster A and IP2 corresponds to target address cluster B, and after IPx is added to target address cluster A, the modularity of target address cluster A increases by 0.1, and after IPx is added to target address cluster B, the modularity of target address cluster B increases by 0.02. Then, by comparing the modularity gains of target address cluster A and target address cluster B, it can be determined that target address cluster A with a larger modularity gain is the address cluster to which IPx belongs. At this time, IPx can be added to target address cluster A. In this way, only the affected address clusters can be re-partitioned, ensuring the accuracy of the partition while improving the calculation efficiency.
[0248] In some embodiments, a sliding window mechanism can be set to recalculate the address cluster partition regularly (such as daily or weekly) using the latest attack data such as preset IP addresses to ensure the real-time performance and accuracy of the structure of each target address cluster.
[0249] By continuously optimizing each address cluster through an iterative process, it can be ensured that the preset IP addresses within each finally partitioned target address cluster have a high degree of homology and a close attack feature association. In this way, not only the understanding and recognition ability of network attack behaviors are enhanced, but also it helps to construct a more accurate and effective network security defense system.
[0250] In some embodiments, in order to facilitate the rapid traceability of the IP address to be analyzed, the characteristic centroid of each target address cluster can be calculated to quantify and characterize the main attack behavior characteristics of each target address cluster, so as to facilitate the subsequent rapid comparison of the IP address to be analyzed with the characteristic centroid, and then achieve the rapid and accurate traceability of the IP address to be analyzed. Exemplarily, "determining the characteristic centroid of each target address cluster" in step 104 may include:
[0251] (104.b1) For each target address cluster, obtain the second feature vector of the target node corresponding to each preset IP address in the target address cluster, and the node degree of the target node;
[0252] (104.b2) Obtain the third product according to the product of the second feature vector and the corresponding node degree;
[0253] (104.b3) obtaining a characteristic sum according to the sum of multiple third products corresponding to multiple preset Internet Protocol addresses included in each target address cluster;
[0254] (104.b4) Obtain the sum of multiple node degrees of multiple target nodes corresponding to each target address cluster to obtain the target degree sum;
[0255] (104.b5) Based on the ratio of the feature sum to the target degree sum, obtain the feature centroid of each target address cluster.
[0256] Among them, the second feature vector can be a feature vector of the target node corresponding to each preset Internet Protocol address, which may include multi-dimensional features such as attack time interval, protocol distribution, load entropy value, etc., which is used to describe the behavior pattern and attributes of the preset Internet Protocol address.
[0257] Among them, the node degree can be a representation of the connection strength between each preset Internet Protocol address and other target nodes (i.e., other preset Internet Protocol addresses) in the target address cluster, which can be obtained by adding the connection weights between the corresponding preset Internet Protocol address and all other preset Internet Protocol addresses of the same source, wherein the method for obtaining the connection weights has been introduced in detail above and will not be repeated here.
[0258] The third product may be a result obtained by multiplying the second eigenvector of the target node by its corresponding node degree, which reflects the weighted contribution of the eigenvector of the target node in the target address cluster to which it belongs.
[0259] The feature sum may be the sum of multiple third products corresponding to all preset Internet Protocol addresses contained in each target address cluster, which integrates the weighted feature vectors of all target nodes in the target address cluster to calculate the overall feature performance of the target address cluster.
[0260] The target degree sum may be the sum of the node degrees corresponding to all target nodes in each target address cluster, and is used to normalize the feature sum.
[0261] In some implementations, the feature centroid of each target address cluster may be calculated using the following formula:
[0262] ;
[0263] in, is the feature centroid of the jth target address cluster, is the node degree of the i-th target node in the target address cluster, is the node feature vector of the i-th target node, Then it represents the third product, Represents the sum of features, represents the sum of target degrees.
[0264] Specifically, the central features of a target address cluster can be described by calculating the feature centroid of each target address cluster, so as to quickly classify new Internet protocol addresses to be analyzed. Specifically, taking the calculation of the feature centroid in target address cluster A as an example, for each target node in target address cluster A, obtain the second feature vector and node degree of the target node. Taking target node 1 as an example, its second feature vector can be [0.2, 0.5, 0.8], and the node degree is 3.
[0265] Furthermore, by multiplying the feature vector (i.e., the second feature vector) of each target node by its node degree, the corresponding third product can be obtained. Still taking target node 1 as an example, its third product is [0.2×3, 0.5×3, 0.8×3], that is, [0.6, 1.5, 2.4].
[0266] Furthermore, the sums of the multiple third products corresponding to all target nodes included in the target address cluster can be added up to obtain the sum of features of the target address cluster. Taking target address cluster A as an example, if it has 3 target nodes, and the third products of each target node are [0.6, 1.5, 2.4], [0.3, 0.9, 1.2], [0.4, 1.2, 1.6] in sequence, then the multiple third products can be added to obtain the corresponding sum of features as [1.3, 3.6, 5.2].
[0267] In some embodiments, the sums of the node degrees corresponding to all target nodes in the target address cluster can be added up to obtain the sum of target degrees. Taking target address cluster A as an example, if it has 3 target nodes, and the node degrees of each target node are 3, 4, and 5 respectively, then the sum of target degrees corresponding to target address cluster A is 12.
[0268] Furthermore, by the ratio of the sum of features to the sum of target degrees, the feature centroid of each target address cluster can be obtained. For example, if the sum of features of target address cluster A is [1.3, 3.6, 5.2] and the sum of target degrees is 12, then the feature centroid is [0.1083, 0.3, 0.4333].
[0269] Through the above method, the feature centroid of each target address cluster can be obtained, so as to transform complex network behavior features into cluster fingerprints that can be efficiently compared, providing core support for large-scale attack homology analysis.
[0270] Step 105: Obtain the Internet Protocol address to be analyzed, calculate the similarity between the target feature vector of the Internet Protocol address to be analyzed and the multiple feature centroids of multiple target address clusters, and obtain the network attack homology analysis result of the Internet Protocol address to be analyzed.
[0271] In some embodiments, in order to quickly identify newly emerging or unknown attack sources, the similarity between the target feature vector of the Internet Protocol address to be analyzed and the feature centroids of multiple determined target address clusters can be calculated to determine whether the Internet Protocol address to be analyzed belongs to a certain known attack group, and thus obtain its network attack homology analysis result. In this way, the response speed and accuracy to potential threats can be effectively improved, and the overall network security protection ability can be enhanced.
[0272] Among them, the Internet Protocol address to be analyzed can be a newly discovered or unclassified IP address that requires homology analysis, and the attack group to which it belongs is unknown, and network attack homology analysis needs to be performed to determine it.
[0273] Among them, the target feature vector can be a set of features possessed by each Internet Protocol address to be analyzed.
[0274] Among them, the network attack homology analysis result can be the result obtained based on the similarity calculation between the target feature vector of the Internet Protocol address to be analyzed and the feature centroids of each target address cluster, which characterizes which known attack group the Internet Protocol address to be analyzed is most likely to belong to, or indicates that it does not belong to any known group, thereby providing a basis for subsequent security policy formulation.
[0275] In some embodiments, after obtaining the Internet Protocol address to be analyzed, its target feature vector can be extracted to obtain multiple feature values. Next, this target feature vector is compared with the multiple feature centroids of multiple target address clusters calculated in advance one by one. Specifically, the comparison can be performed by calculating the similarity between the target feature vector and each feature centroid.
[0276] Furthermore, the similarity between each feature in the target feature vector and the feature in the corresponding feature centroid can be measured. The features for which the similarity is measured should belong to the same dimension, such as all belonging to the payload entropy value, etc.
[0277] Specifically, methods such as cosine similarity and Jaccard similarity can be used for similarity measurement. When calculating the similarity between the target feature vector and each feature centroid, the similarity between each feature in the target feature vector and the sub-feature centroid in the corresponding dimension of the feature centroid can be calculated. After obtaining the corresponding feature similarity, the calculated result is multiplied by the feature centroid importance score corresponding to the feature to obtain the target feature similarity corresponding to each target feature.
[0278] Exemplarily, similar to calculating the importance score of the first eigenvector, the importance score of the feature centroid can also be calculated by the initial gradient boosting model. Specifically, the loss value is calculated by comparing the predicted importance score with the actual importance score to train the initial gradient boosting model. After the model training is completed, the gains corresponding to each feature dimension can be accumulated to obtain the importance score corresponding to the feature of each feature dimension. Alternatively, the importance score of the feature centroid can also be set manually, and the embodiments of the present application do not make specific limitations in this regard.
[0279] Furthermore, the target feature similarities corresponding to the multiple target features included in the target feature vector can be added together to obtain the total feature similarity between each feature centroid and the Internet protocol address to be analyzed.
[0280] Furthermore, after calculating multiple total feature similarities, each total feature similarity can be compared with a preset third dynamic threshold. When there is a total feature similarity greater than the third dynamic threshold, the feature centroid corresponding to the total feature similarity can be used as the target feature centroid corresponding to the Internet protocol address to be analyzed, and the target address cluster corresponding to the target feature centroid can be used as the target homologous cluster of the Internet protocol address to be analyzed, so as to obtain the network attack homologous analysis result of the Internet protocol address to be analyzed.
[0281] In some embodiments, when there are multiple total feature similarities greater than the third dynamic threshold, the target address cluster corresponding to the feature centroid with the largest total feature similarity can be determined as the target homologous cluster of the Internet protocol address to be analyzed. For example, if the target total feature similarity between the feature centroid 1 of the target address cluster A and the Internet protocol address to be analyzed is 0.9, the target total feature similarity between the feature centroid 2 of the target address cluster B and the Internet protocol address to be analyzed is 0.75, and the third dynamic threshold is 0.7, then it can be determined that the target address cluster A is the target homologous cluster of the Internet protocol address to be analyzed.
[0282] This application obtains multiple preset Internet Protocol (IP) addresses and a preset weighted attack behavior graph. The preset weighted attack behavior graph includes multiple preset IP addresses and multiple connection weights, where each connection weight represents the degree of association between two corresponding preset IP addresses with the same source. The multiple preset IP addresses in the preset weighted attack behavior graph are divided into multiple initial address clusters. For each initial address cluster, the connection weights associated with each preset IP address included are determined, and the corresponding modularity is calculated based on the connection weights. The modularity is used to characterize the degree of compatibility with the same source among the multiple preset IP addresses included in the corresponding initial address cluster. According to the modularity corresponding to each initial address cluster, the attribution relationship between each preset IP address and the initial address cluster is re-determined to obtain multiple target address clusters, and the characteristic centroid of each target address cluster is determined. The IP address to be analyzed is obtained, and the similarity between the target feature vector of the IP address to be analyzed and the multiple characteristic centroids of the multiple target address clusters is calculated to obtain the network attack homology analysis result of the IP address to be analyzed. In this way, the correlation of attack behaviors between any two IP addresses can be accurately quantified through the connection weights between the multiple IP addresses in the pre-generated preset weighted attack behavior graph, so as to improve the interpretability of the relationships between IP addresses, and further improve the accuracy of homology analysis. Moreover, by adopting the method of modularity iterative processing, the preset weighted attack behavior graph is divided into target address clusters, and the characteristic centroid of each target address cluster is calculated. In this way, a globally representative core feature vector can be accurately generated based on the IP addresses with the same source included in the same address cluster. When there is an IP address to be analyzed that requires homology analysis subsequently, the IP address to be analyzed can be directly compared with the characteristic centroids of each target address cluster (including a large number of IP addresses with the same source), without having to compare and cluster with each of the large number of IP addresses one by one. Homology determination is performed based on the centroid similarity, avoiding the low efficiency problem caused by pairwise comparison of the IP address to be analyzed with the large number of IP addresses. In this way, both the traceability of the attack pattern and the accuracy of homology analysis are retained, and the efficiency of homology analysis is improved. In summary, this application can improve the efficiency and accuracy of network attack homology analysis, which is of great significance for enhancing network security protection capabilities.
[0283] In some embodiments, in order to improve the efficiency and accuracy of homology analysis, the homology cluster attribution of the Internet protocol address to be analyzed can be determined by calculating the similarity between the feature vector of the Internet protocol address to be analyzed and each known feature centroid, so as to improve the interpretability and analysis efficiency of homology analysis. For example, "calculating the similarity between the target feature vector of the Internet protocol address to be analyzed and the multiple feature centroids of multiple target address clusters to obtain the network attack homology analysis result of the Internet protocol address to be analyzed" in step 105 may include:
[0284] (105.1)Calculating the similarity between each target feature in the target feature vector corresponding to the Internet protocol address to be analyzed and each sub-feature centroid in each feature centroid to obtain the corresponding feature similarity;
[0285] (105.2)Obtaining the feature centroid importance score corresponding to each sub-feature centroid, and obtaining the target feature similarity corresponding to each target feature based on the product of the feature similarity and the feature centroid importance score;
[0286] (105.3)Adding the multiple target feature similarities corresponding to the multiple target features to obtain the total feature similarity corresponding to each feature centroid;
[0287] (105.4)Obtaining a preset third dynamic threshold, and sequentially comparing the multiple total feature similarities corresponding to the multiple feature centroids with the third dynamic threshold to obtain a comparison result;
[0288] (105.5)When the comparison result indicates that there is a target total feature similarity greater than the third dynamic threshold among the multiple total feature similarities, determining the corresponding target address cluster as the target homology cluster of the Internet protocol address to be analyzed, and obtaining the network attack homology analysis result of the Internet protocol address to be analyzed based on the target homology cluster.
[0289] Among them, the target features can be the specific features in the target feature vector of the Internet protocol address to be analyzed, such as the feature values corresponding to the feature dimensions such as attack time interval, protocol distribution, and payload entropy value. The multiple target features can form the target feature vector corresponding to the Internet protocol address to be analyzed.
[0290] Among them, the sub-feature centroid can be the specific feature included in each feature centroid, which is used to characterize the feature value of the target homology cluster in a specific dimension.
[0291] Among them, the feature similarity can be the degree of similarity between each target feature of the Internet protocol address to be analyzed and the corresponding sub-feature centroid. It can be calculated by methods such as cosine similarity or Jaccard similarity.
[0292] Among them, the characteristic centroid importance score can be a numerical value reflecting the relative importance of each sub - characteristic centroid in its affiliated characteristic centroid. It can be obtained through historical data statistics, evaluation by technical personnel, or calculation using an initial gradient boosting model.
[0293] Among them, the target feature similarity can be the result obtained based on the product of the feature similarity and the characteristic centroid importance score, representing the weighted similarity between a certain target feature of the Internet Protocol address to be analyzed and a certain sub - characteristic centroid.
[0294] Among them, the total feature similarity can be the numerical value obtained by adding up the target feature similarities corresponding to multiple target features, representing the overall similarity between the Internet Protocol address to be analyzed and the corresponding characteristic centroid.
[0295] Among them, the third dynamic threshold can be a threshold used to determine whether the total feature similarity reaches the attribution standard. If there is a total feature similarity greater than this threshold, it is considered that the Internet Protocol address to be analyzed has a significant similarity with the corresponding characteristic centroid.
[0296] Among them, the comparison result can be the result obtained by comparing multiple total feature similarities with the second dynamic threshold, used to determine whether there is a situation where the target total feature similarity is greater than the second dynamic threshold.
[0297] Among them, the target homologous cluster can be when the comparison result indicates that there is a target total feature similarity greater than the second dynamic threshold, the target address cluster to which the corresponding characteristic centroid belongs is the target homologous cluster, that is, it indicates that the probability that the Internet Protocol address to be analyzed belongs to this target homologous cluster is the highest.
[0298] In some embodiments, the total feature similarity between the Internet Protocol address to be analyzed and the corresponding characteristic centroid can be calculated by the following formula:
[0299] ;
[0300] Among them, represents the Internet Protocol address to be analyzed, represents the characteristic centroid of the j - th target address cluster, represents the characteristic centroid importance score, represents the k - th target feature, represents the k - th sub - characteristic centroid, represents the feature similarity.
[0301] Exemplarily, if the target feature vector of the Internet Protocol address to be analyzed is [0.8, 1, 0.7], and the feature centroid 1 of the target address cluster A is [0.9, 0, 0.6], by calculating the similarity between each target feature in the target vector and the corresponding sub-feature centroid, the corresponding feature similarity can be obtained. Taking the feature in the first dimension as an example, the feature similarity between the target feature 0.8 and the sub-feature centroid 0.9 can be calculated. For example, if the cosine similarity is used for calculation, then the feature similarity can be obtained as 0.95. Thus, the feature similarity between the target feature and the sub-feature centroid under each feature dimension can be calculated.
[0302] In some embodiments, the calculation methods of the feature similarities under each feature dimension can be the same or different, and can be specifically set according to the actual situation.
[0303] Further, if in the feature centroid 1, the weight w1 of the first dimension is 0.5, the weight w2 of the second dimension is 0.3, and the weight w3 of the third dimension is 0.2, then, if through the above calculation method, the feature similarities of the three dimensions are 0.95, 0, and 0.9 respectively. Thus, through the product of each feature similarity and the corresponding feature centroid importance score, the target feature similarity corresponding to each target feature can be calculated. For example, the target feature similarity of the first dimension is 0.95×0.5 = 0.475.
[0304] Further, the multiple target feature similarities corresponding to the target feature vector can be added together to obtain the total feature similarity corresponding to the corresponding feature centroid and the Internet Protocol address to be analyzed: for the first dimension, it is 0.95×0.5 = 0.475; for the second dimension, it is 0×0.3 = 0; for the third dimension, it is 0.9×0.2 = 0.18. Adding the target feature similarities corresponding to the 3 dimensions, that is, 0.475 + 0 + 0.18, the final total feature similarity is obtained as 0.655.
[0305] Exemplarily, if the preset third dynamic threshold is 0.6, then the total feature similarity can be compared with the third dynamic threshold. For example, comparing 0.655 with 0.6, 0.655 is greater than 0.6, indicating that the Internet Protocol address to be analyzed is homologous to the target address cluster A corresponding to the feature centroid 1. At the same time, if the total feature similarity between the Internet Protocol address to be analyzed and other target address clusters is lower than the third dynamic threshold, it indicates that the homologous relationship between the Internet Protocol address to be analyzed and these target address clusters is non-homologous.
[0306] Through the above method, a comprehensive search of the entire database can be avoided, effectively simplifying the homologous analysis process, realizing efficient, accurate, and interpretable homologous attack analysis, and providing key technical support for real-time threat response.
[0307] Please refer to Figure 3 , Figure 3 which is the overall flowchart of the network attack homology analysis method provided by the embodiments of this application. Exemplarily, multiple preset Internet protocol addresses can be extracted from the attack logs, and the preset Internet protocol addresses can be a large number of IP addresses covering various attack groups. Then, according to the multiple extracted preset Internet protocol addresses, an attack behavior feature library can be constructed to facilitate subsequent operations such as the division of address clusters, etc.
[0308] Specifically, the attack behavior feature library can include spatio-temporal features, behavior fingerprints, and association graphs. Exemplarily, the spatio-temporal features can depict the attacker's behavior patterns from three dimensions: the network layer, geographical features, and time patterns. The behavior fingerprints can depict the attacker's technical means and behavior habits from three dimensions: attack vectors, payload features, and password features. The association graphs can depict the association relationships between attackers from three dimensions: honeypot linkage, service association, and domain name mapping.
[0309] Furthermore, the features in the attack behavior feature library can be encoded to facilitate obtaining the first feature vector. The ways of encoding the features can include the original values of numerical features, One-Hot encoding (i.e., one-hot encoding) of categorical features, vectorization and cosine similarity of text features, binary values of boolean features, and so on. And based on the data encoded from any two preset Internet protocol addresses, the relationship between the two preset Internet protocol addresses is modeled, and the connection weight between the two preset Internet protocol addresses is determined. Finally, with the preset Internet protocol address as the target node, the homology relationship between two preset Internet protocol addresses as the connection edge, and the connection weight as the edge weight, a preset weighted attack behavior graph can be constructed.
[0310] Furthermore, multiple preset Internet protocol addresses can be divided into multiple initial address clusters, and for each initial address cluster, the modularity is calculated according to the connection weights between the preset Internet protocol addresses in the preset weighted attack behavior graph to evaluate the homology compatibility between the preset Internet protocol addresses within the address cluster. Then, according to the magnitudes of the modularity corresponding to each initial address cluster, the attribution relationship between each preset Internet protocol address and the initial address cluster is re-determined to obtain multiple target address clusters.
[0311] Further, for each target address cluster, the feature centroid representing the average features of the entire address cluster can be calculated. In this way, when an Internet Protocol address to be analyzed for homologous analysis appears, the similarity between the target feature vector corresponding to the Internet Protocol address to be analyzed and the feature centroids of multiple target address clusters can be calculated respectively, and the result of homologous analysis of network attacks can be obtained, avoiding the low efficiency problem caused by pairwise comparison of the Internet Protocol address to be analyzed with a large number of Internet Protocol addresses. In this way, the traceability of the attack pattern is retained, the accuracy of homologous analysis is ensured, and the efficiency of homologous analysis is improved.
[0312] In some embodiments, after determining that the Internet Protocol address to be analyzed belongs to a certain target address cluster, the Internet Protocol address to be analyzed can be updated to the corresponding target address cluster, and the feature centroid of the target address cluster can be updated to ensure the real-time performance and accuracy of homologous determination.
[0313] Please refer to Figure 4 , the embodiment of the present application further provides a device for homologous analysis of network attacks, which can implement the above-mentioned method for homologous analysis of network attacks. The device for homologous analysis of network attacks includes:
[0314] An acquisition module 41, configured to acquire a plurality of preset Internet Protocol addresses and a preset weighted attack behavior graph, where the preset weighted attack behavior graph includes a plurality of preset Internet Protocol addresses and a plurality of connection weights, and each connection weight represents the association degree between two corresponding preset Internet Protocol addresses having a homologous relationship;
[0315] A partitioning module 42, configured to partition the plurality of preset Internet Protocol addresses in the preset weighted attack behavior graph into a plurality of initial address clusters;
[0316] A first calculation module 43, configured to, for each initial address cluster, determine the connection weights associated with each preset Internet Protocol address included, and calculate the corresponding modularity according to the connection weights, where the modularity is used to characterize the homologous compatibility degree between the plurality of preset Internet Protocol addresses included in the corresponding initial address cluster;
[0317] A determination module 44, configured to re-determine the attribution relationship between each preset Internet Protocol address and the initial address cluster according to the modularity corresponding to each initial address cluster, obtain a plurality of target address clusters, and determine the feature centroid of each target address cluster;
[0318] A second calculation module 45, configured to acquire the Internet Protocol address to be analyzed, and calculate the similarity between the target feature vector of the Internet Protocol address to be analyzed and the plurality of feature centroids of the plurality of target address clusters, so as to obtain the result of homologous analysis of network attacks of the Internet Protocol address to be analyzed.
[0319] The specific implementation manner of the network attack homology analysis device is basically the same as the specific embodiments of the above-mentioned network attack homology analysis method, and will not be elaborated here. On the premise of meeting the requirements of the embodiments of the present application, other functional modules can also be set in the network attack homology analysis device to implement the network attack homology analysis method in the above embodiments.
[0320] The embodiments of the present application also provide a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned network attack homology analysis method. The computer device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0321] Please refer to Figure 5 , Figure 5 which shows the hardware structure of a computer device in another embodiment. The computer device includes:
[0322] A processor 51, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0323] A memory 52, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 52 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 52, and the processor 51 is called to execute the network attack homology analysis method of the embodiments of the present application;
[0324] An input / output interface 53, which is used to implement information input and output;
[0325] A communication interface 54, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);
[0326] A bus 55, which transmits information between various components of the device (such as the processor 51, the memory 52, the input / output interface 53, and the communication interface 54);
[0327] Among them, the processor 51, the memory 52, the input / output interface 53, and the communication interface 54 are communicatively connected to each other inside the device through the bus 55.
[0328] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned network attack homology analysis method is implemented.
[0329] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0330] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0331] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine some steps, or different steps.
[0332] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0333] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0334] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0335] It should be understood that in this application, "at least one (item)" and "several" mean one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression is any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0336] In several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0337] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0338] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0339] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.
[0340] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.
Claims
1. A method for analyzing the homology of network attacks, characterized in that, The method includes: Obtaining a plurality of preset Internet protocol addresses and a preset weighted attack behavior graph, wherein the preset weighted attack behavior graph includes a plurality of preset Internet protocol addresses and a plurality of connection weights, and each connection weight represents the association degree between two preset Internet protocol addresses with a homologous relationship; Dividing the plurality of preset Internet protocol addresses in the preset weighted attack behavior graph into a plurality of initial address clusters; For each initial address cluster, determining the connection weights associated with each preset Internet protocol address included, and calculating the corresponding modularity according to the connection weights, where the modularity is used to characterize the homologous compatibility degree among the plurality of preset Internet protocol addresses included in the corresponding initial address cluster; According to the modularity corresponding to each initial address cluster, re-determining the attribution relationship between each preset Internet protocol address and the initial address cluster, obtaining a plurality of target address clusters, and determining the characteristic centroid of each target address cluster; Obtaining the Internet protocol address to be analyzed, calculating the similarity between each target feature in the target feature vector corresponding to the Internet protocol address to be analyzed and each sub-feature centroid in each feature centroid to obtain the corresponding feature similarity; obtaining the feature centroid importance score corresponding to each sub-feature centroid, and based on the product of the feature similarity and the feature centroid importance score, obtaining the target feature similarity corresponding to each target feature; adding the multiple target feature similarities corresponding to the multiple target features to obtain the total feature similarity corresponding to each feature centroid; obtaining a preset third dynamic threshold, and sequentially comparing the multiple total feature similarities corresponding to the multiple feature centroids with the third dynamic threshold to obtain a comparison result; when the comparison result indicates that there is a target total feature similarity greater than the third dynamic threshold among the multiple total feature similarities, determining the corresponding target address cluster as the target homologous cluster of the Internet protocol address to be analyzed, and based on the target homologous cluster, obtaining the network attack homologous analysis result of the Internet protocol address to be analyzed.
2. The network attack homology analysis method according to claim 1, characterized in that The obtaining the plurality of preset Internet protocol addresses and the preset weighted attack behavior graph includes: Obtaining the plurality of preset Internet protocol addresses and the homologous relationship between the plurality of preset Internet protocol addresses, and determining the first feature vector and the relationship label between any two preset Internet protocol addresses, where the first feature vector includes a plurality of first feature values of the any two preset Internet protocol addresses in a plurality of feature interaction dimensions; Sequentially inputting each first feature vector and the corresponding relationship label into the initial gradient boosting model, and performing decision tree split gain accumulation through the initial gradient boosting model to obtain the feature importance score corresponding to each first feature value; Based on the plurality of feature importance scores, determining the connection weight between the two preset Internet protocol addresses corresponding to the first feature vector; Based on the plurality of preset Internet protocol addresses, the relationship label between any two preset Internet protocol addresses, and the corresponding connection weights, constructing the preset weighted attack behavior graph corresponding to the plurality of preset Internet protocol addresses.
3. The network attack homology analysis method according to claim 2, characterized in that Successively input each first feature vector and the corresponding relationship label into the initial gradient boosting model, and perform decision tree split gain accumulation through the initial gradient boosting model to obtain the feature importance score corresponding to each first eigenvalue, including: Successively input each first feature vector and the corresponding relationship label into the initial gradient boosting model, and through the initial gradient boosting model, predict each first feature vector to obtain the corresponding predicted label; Determine the first target loss based on the difference between the predicted label and the relationship label; Based on the first target loss, calculate the split gain for the multiple first eigenvalues included in each first feature vector, and determine the intermediate split feature with the largest split gain from the multiple first eigenvalues for node splitting, and calculate the second target loss; Based on the difference between the second target loss and the first target loss, calculate the split gain for the multiple first eigenvalues included in the first feature vector, and determine the intermediate split feature with the largest split gain from the multiple first eigenvalues for node splitting, and calculate the updated second target loss; Repeat the step of calculating the split gain for the multiple first eigenvalues included in the first feature vector based on the difference between the updated second target loss and the first target loss, and determining the intermediate split feature with the largest split gain from the multiple first eigenvalues for node splitting, and calculating the updated second target loss until the preset number of training times is reached, and obtain the feature importance score corresponding to each first eigenvalue according to the multiple split gains iteratively obtained for each first eigenvalue.
4. The network attack homology analysis method according to claim 2, wherein The determining the connection weights of the two preset Internet protocol addresses corresponding to the first feature vector based on the multiple feature importance scores includes: Perform normalization processing on each feature importance score to obtain the corresponding target feature importance score; Obtain the sub-feature weight of each first eigenvalue according to the product of each first eigenvalue and the corresponding target feature importance score; Based on the sum of the multiple sub-feature weights corresponding to the multiple first eigenvalues, obtain the connection weights of the two preset Internet protocol addresses corresponding to the first feature vector.
5. The network attack homology analysis method according to claim 2, wherein Before constructing the preset weighted attack behavior graph corresponding to the multiple preset Internet protocol addresses based on the multiple preset Internet protocol addresses, the relationship labels between any two preset Internet protocol addresses, and the corresponding connection weights, further include: Obtain the multiple historical connection weights between any two preset Internet protocol addresses, and the scoring mean and scoring standard deviation of the multiple historical connection weights; Obtain a preset adjustment parameter, and obtain the first product according to the product of the adjustment parameter and the scoring standard deviation; Determine the first dynamic threshold according to the difference between the scoring mean and the first product; Compare the connection weight between any two preset Internet protocol addresses with the first dynamic threshold to obtain a comparison result; Based on the comparison result, update the relationship label between any two preset Internet protocol addresses to obtain the target relationship label; Then, constructing the preset weighted attack behavior graph corresponding to the multiple preset Internet protocol addresses, including: based on the multiple preset Internet protocol addresses, the relationship tags between any two preset Internet protocol addresses, and the corresponding connection weights. Constructing the preset weighted attack behavior graph corresponding to the multiple preset Internet protocol addresses based on the multiple preset Internet protocol addresses, the target relationship tags between any two preset Internet protocol addresses, and the corresponding connection weights.
6. The network attack homology analysis method according to claim 1, wherein The calculating the corresponding modularity according to the connection weights includes: Obtaining the graph connection weight according to the sum of the connection weights corresponding to all the connection edges in the preset weighted attack behavior graph; Obtaining the node degree corresponding to the target node of each preset Internet protocol address in the corresponding initial address cluster; Obtaining the second product according to the product of the node degrees between any two target nodes; Obtaining the first ratio based on the ratio of the second product to the graph connection weight; Obtaining the first difference, which is the difference between the connection weight corresponding to any two target nodes in the initial address cluster and the corresponding first ratio; Obtaining the modularity corresponding to the initial address cluster based on the graph connection weight and the multiple first differences between the multiple target nodes included in the initial address cluster.
7. The method for analyzing the homology of network attacks according to claim 1, wherein The re-determining the belonging relationship between each preset Internet protocol address and the initial address cluster according to the modularity corresponding to each initial address cluster to obtain multiple target address clusters includes: Determining the intermediate address clusters with modularity less than a preset second dynamic threshold according to the modularity corresponding to each initial address cluster; Obtaining a preset attack behavior feature library, and based on the attack behavior feature library, re-determining the associated Internet protocol addresses having an attack feature relationship with each preset Internet protocol address in each intermediate address cluster, and migrating each preset Internet protocol address to the initial address cluster corresponding to the associated Internet protocol address; Repeating the step of re-determining the associated Internet protocol addresses having the attack feature relationship with each preset Internet protocol address in each intermediate address cluster based on the attack behavior feature library, and migrating each preset Internet protocol address to the initial address cluster corresponding to the associated Internet protocol address until the modularity corresponding to each intermediate address cluster is greater than the second dynamic threshold, to obtain multiple target address clusters.
8. The method for analyzing the homology of network attacks according to claim 1, characterized in that, The determining the feature centroid of each target address cluster includes: For each target address cluster, obtaining the second feature vector of the target node corresponding to each preset Internet protocol address in the target address cluster, and the node degree of the target node; Obtaining the third product according to the product of the second feature vector and the corresponding node degree; Obtaining the feature sum according to the sum of the multiple third products corresponding to the multiple preset Internet protocol addresses included in each target address cluster; Obtaining the sum of the multiple node degrees of the multiple target nodes corresponding to each target address cluster to obtain the target degree sum; Obtaining the feature centroid of each target address cluster based on the ratio of the feature sum to the target degree sum.
9. A network attack homology analysis device, characterized in that, The device includes: an acquisition module, configured to acquire a plurality of preset Internet protocol addresses and a preset weighted attack behavior graph, where the preset weighted attack behavior graph includes a plurality of preset Internet protocol addresses and a plurality of connection weights, and each connection weight represents the association degree between two preset Internet protocol addresses with a homologous relationship; a division module, configured to divide the plurality of preset Internet protocol addresses in the preset weighted attack behavior graph into a plurality of initial address clusters; a first calculation module, configured to, for each initial address cluster, determine the connection weights associated with each preset Internet protocol address included therein, and calculate the corresponding modularity according to the connection weights, where the modularity is used to characterize the homologous compatibility degree among the plurality of preset Internet protocol addresses included in the corresponding initial address cluster; a determination module, configured to re-determine the belonging relationship between each preset Internet protocol address and the initial address cluster according to the modularity corresponding to each initial address cluster, obtain a plurality of target address clusters, and determine the characteristic centroid of each target address cluster; a second calculation module, configured to acquire an Internet protocol address to be analyzed, calculate the similarity between each target feature in the target feature vector corresponding to the Internet protocol address to be analyzed and each sub-feature centroid in each feature centroid to obtain the corresponding feature similarity; obtain the feature centroid importance score corresponding to each sub-feature centroid, and obtain the target feature similarity corresponding to each target feature based on the product of the feature similarity and the feature centroid importance score; add the plurality of target feature similarities corresponding to the plurality of target features to obtain the total feature similarity corresponding to each feature centroid; obtain a preset third dynamic threshold, and sequentially compare the plurality of total feature similarities corresponding to the plurality of feature centroids with the third dynamic threshold to obtain a comparison result; when the comparison result indicates that there is a target total feature similarity greater than the third dynamic threshold among the plurality of total feature similarities, determine the corresponding target address cluster as the target homologous cluster of the Internet protocol address to be analyzed, and obtain the network attack homologous analysis result of the Internet protocol address to be analyzed based on the target homologous cluster.
10. A computer device, characterized in that, The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the network attack homologous analysis method according to any one of claims 1 to 8 is implemented.
11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the network attack homologous analysis method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
User risk assessment method and device
CN117009879A
Human-cluster interaction method and system based on augmented reality
CN117075725A