Network attack homologous analysis method and device, computer equipment and storage medium

By using preset weighted attack behavior diagrams and modular calculation methods in the homologous analysis of network attacks, the problems of high computational complexity and lack of interpretability in the existing technology are solved, efficient and accurate homologous analysis is achieved, and network security protection capabilities are improved.

CN120165989AActive Publication Date: 2025-06-17PENG CHENG LAB

Patent Information

Application Number
CN202510646517.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-17
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The existing technology has high computational complexity, low analysis efficiency in the homologous analysis of network attacks, and lacks interpretability of clustering results, making it difficult for security personnel to effectively trace the attack mode, misjudgment or misjudgment of homologous relationships.

Method used

A method of homologous analysis of network attacks is proposed. By obtaining multiple preset Internet protocol addresses and preset weighted attack behavior diagrams, the initial address cluster is divided, the module degree is calculated, the address cluster is redetermined, the target address cluster is obtained, the characteristic center of mass is calculated, and the similarity calculation is performed to obtain the homologous analysis results.

Benefits of technology

It improves the efficiency and accuracy of homologous analysis of network attacks, enhances the traceability of attack mode, reduces resource waste, and improves network security protection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120165989A_ABST
    Figure CN120165989A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a network attack homologous analysis method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring a plurality of preset internet protocol addresses and a preset weighted attack behavior graph; dividing a plurality of preset internet protocol addresses in the preset weighted attack behavior graph into a plurality of initial address clusters; for each initial address cluster, determining a connection weight associated with each preset internet protocol address contained in the initial address cluster, and calculating a corresponding modularity according to the connection weight; according to the modularity corresponding to each initial address cluster, a plurality of target address clusters are obtained again, and the feature centroid of each target address cluster is determined; and obtaining a to-be-analyzed internet protocol address, and performing similarity calculation on the target feature vector of the to-be-analyzed internet protocol address and the plurality of feature centroids of the plurality of target address clusters to obtain a network attack homologous analysis result of the to-be-analyzed internet protocol address. Therefore, the efficiency and accuracy of network attack homologous analysis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technology, and in particular, to a method, device, computer device and storage medium for network attack homology analysis. Background Art

[0002] A network attack is an act of using computer network technology to illegally invade, interfere with, damage or steal information resources in a network system to achieve specific goals. It can include, but is not limited to, various forms such as data leakage, system damage, denial-of-service attacks, and malicious software propagation, posing a serious threat to personal privacy, corporate interests, social order and security.

[0003] In order to more effectively respond to complex and ever-changing network attacks and enhance the overall protection ability of network security, the correlation between different attack events can be identified through network attack homology analysis, and it can be judged whether they come from the same attack source or organization, so as to reveal the attacker's behavior patterns and technical means, improve the defense efficiency, and reduce resource waste.

[0004] In related technologies, usually, the Internet Protocol address corresponding to each attack event is used as a node, and the feature similarity between all Internet Protocol addresses is compared pairwise, and the Internet Protocol addresses with high feature similarity are clustered to identify the homology relationship between attackers. However, on the one hand, the method of clustering after pairwise comparison has a high computational complexity and low analysis efficiency when facing a large amount of attack data; on the other hand, the clustering result lacks interpretability, resulting in security personnel being unable to effectively trace the attack pattern, and thus misjudging or missing the homology relationship, reducing the accuracy of the analysis result. Summary of the Invention

[0005] This application proposes a method, device, computer device and storage medium for network attack homology analysis, which can improve the efficiency and accuracy of network attack homology analysis.

[0006] To achieve the above object, the first aspect of the embodiments of this application proposes a method for network attack homology analysis, and the method includes: Obtain a plurality of preset Internet Protocol addresses and a preset weighted attack behavior graph, where the preset weighted attack behavior graph includes a plurality of preset Internet Protocol addresses and a plurality of connection weights, and each connection weight represents the association degree between two preset Internet Protocol addresses with a homology relationship; Divide the plurality of preset Internet Protocol addresses in the preset weighted attack behavior graph into a plurality of initial address clusters; For each initial address cluster, determine the connection weights associated with each preset Internet protocol address included, and calculate the corresponding modularity according to the connection weights, where the modularity is used to characterize the homology compatibility degree among the multiple preset Internet protocol addresses included in the corresponding initial address cluster; According to the modularity corresponding to each initial address cluster, re-determine the attribution relationship between each preset Internet protocol address and the initial address cluster, obtain multiple target address clusters, and determine the characteristic centroid of each target address cluster; Obtain the Internet protocol address to be analyzed, and calculate the similarity between the target feature vector of the Internet protocol address to be analyzed and the multiple characteristic centroids of the multiple target address clusters, so as to obtain the network attack homology analysis result of the Internet protocol address to be analyzed.

[0007] Correspondingly, a second aspect of the embodiments of the present application proposes a network attack homology analysis device, and the device includes: An acquisition module, configured to acquire a plurality of preset Internet protocol addresses and a preset weighted attack behavior graph, where the preset weighted attack behavior graph includes a plurality of preset Internet protocol addresses and a plurality of connection weights, and each connection weight represents the association degree between two preset Internet protocol addresses with a homologous relationship; A division module, configured to divide the plurality of preset Internet protocol addresses in the preset weighted attack behavior graph into a plurality of initial address clusters; A first calculation module, configured to, for each initial address cluster, determine the connection weights associated with each preset Internet protocol address included, and calculate the corresponding modularity according to the connection weights, where the modularity is used to characterize the homology compatibility degree among the multiple preset Internet protocol addresses included in the corresponding initial address cluster; A determination module, configured to re-determine the attribution relationship between each preset Internet protocol address and the initial address cluster according to the modularity corresponding to each initial address cluster, obtain a plurality of target address clusters, and determine the characteristic centroid of each target address cluster; A second calculation module, configured to obtain the Internet protocol address to be analyzed, and calculate the similarity between the target feature vector of the Internet protocol address to be analyzed and the multiple characteristic centroids of the multiple target address clusters, so as to obtain the network attack homology analysis result of the Internet protocol address to be analyzed.

[0008] In some embodiments, the acquisition module is further configured to: Acquire the plurality of preset Internet protocol addresses and the homologous relationships between the plurality of preset Internet protocol addresses, and determine the first feature vector and the relationship label between any two preset Internet protocol addresses, where the first feature vector includes a plurality of first feature values of the any two preset Internet protocol addresses in a plurality of feature interaction dimensions; Each first feature vector and its corresponding relationship label are sequentially input into the initial gradient boosting model, and decision tree split gain accumulation is performed through the initial gradient boosting model to obtain the feature importance score corresponding to each first eigenvalue; Based on the multiple feature importance scores, determine the connection weights of the two preset Internet protocol addresses corresponding to the first feature vector; Based on the multiple preset Internet protocol addresses, the relationship labels between any two preset Internet protocol addresses, and the corresponding connection weights, construct the preset weighted attack behavior graph corresponding to the multiple preset Internet protocol addresses.

[0009] In some embodiments, the obtaining module is further configured to: Each first feature vector and its corresponding relationship label are sequentially input into the initial gradient boosting model, and through the initial gradient boosting model, each first feature vector is predicted to obtain the corresponding predicted label; Based on the difference between the predicted label and the relationship label, determine the first target loss; Based on the first target loss, perform split gain calculation on the multiple first eigenvalues included in each first feature vector, and determine the intermediate split feature with the largest split gain from the multiple first eigenvalues for node splitting, and calculate the second target loss; Based on the difference between the second target loss and the first target loss, perform split gain calculation on the multiple first eigenvalues included in the first feature vector, and determine the intermediate split feature with the largest split gain from the multiple first eigenvalues for node splitting, and calculate the updated second target loss; Repeat the step of performing split gain calculation on the multiple first eigenvalues included in the first feature vector based on the difference between the updated second target loss and the first target loss, and determining the intermediate split feature with the largest split gain from the multiple first eigenvalues for node splitting, and calculating the updated second target loss until the preset number of training times is reached. According to the multiple split gains iteratively obtained for each eigenvalue, obtain the feature importance score corresponding to each first eigenvalue.

[0010] In some embodiments, the obtaining module is further configured to: Perform normalization processing on each feature importance score to obtain the corresponding target feature importance score; According to the product of each first eigenvalue and the corresponding target feature importance score, obtain the sub-feature weight of each first eigenvalue; Based on the sum of the multiple sub-feature weights corresponding to the multiple first eigenvalues, obtain the connection weight between the two preset Internet protocol addresses corresponding to the first eigenvector.

[0011] In some embodiments, the network attack homology analysis device further includes a comparison module for: Obtain multiple historical connection weights between any two preset Internet protocol addresses, as well as the scoring mean and scoring standard deviation of the multiple historical connection weights; Obtain a preset adjustment parameter, and obtain a first product according to the product of the adjustment parameter and the scoring standard deviation; Determine a first dynamic threshold according to the difference between the scoring mean and the first product; Compare the connection weight between any two preset Internet protocol addresses with the first dynamic threshold to obtain a comparison result; Based on the comparison result, update the relationship label between any two preset Internet protocol addresses to obtain a target relationship label; Then, constructing the preset weighted attack behavior graph corresponding to the multiple preset Internet protocol addresses based on the multiple preset Internet protocol addresses, the relationship labels between any two preset Internet protocol addresses, and the corresponding connection weights includes: Construct the preset weighted attack behavior graph corresponding to the multiple preset Internet protocol addresses based on the multiple preset Internet protocol addresses, the target relationship labels between any two preset Internet protocol addresses, and the corresponding connection weights.

[0012] In some embodiments, the first calculation module is further configured to: Obtain a graph connection weight according to the sum of the connection weights corresponding to all connection edges in the preset weighted attack behavior graph; Obtain the node degree corresponding to the target node of each preset Internet protocol address in the corresponding target node of the initial address cluster; Obtain a second product according to the product of the node degrees between any two target nodes; Obtain a first ratio based on the ratio of the second product to the graph connection weight; Obtain a first difference, which is the difference between the connection weight corresponding to any two target nodes in the initial address cluster and the corresponding first ratio; Based on the graph connection weight and the multiple first differences between the multiple target nodes included in the initial address cluster, obtain the modularity corresponding to the initial address cluster.

[0013] In some embodiments, the determination module is further configured to: Determine intermediate address clusters with modularity less than a preset second dynamic threshold according to the modularity corresponding to each initial address cluster; Obtain a preset attack behavior feature library, and based on the attack behavior feature library, re-determine associated Internet protocol addresses that have an attack feature relationship with each preset Internet protocol address in each intermediate address cluster, and migrate each preset Internet protocol address to the initial address cluster corresponding to the associated Internet protocol address; Repeat the step of re-determining, based on the attack behavior feature library, associated Internet protocol addresses that have the attack feature relationship with each preset Internet protocol address in each intermediate address cluster, and migrating each preset Internet protocol address to the initial address cluster corresponding to the associated Internet protocol address until the modularity corresponding to each intermediate address cluster is greater than the second dynamic threshold, to obtain multiple target address clusters.

[0014] In some embodiments, the determining module is further configured to: For each target address cluster, obtain the second feature vector of the target node corresponding to each preset Internet protocol address in the target address cluster, and the node degree of the target node; Obtain a third product according to the product of the second feature vector and the corresponding node degree; Obtain a feature sum according to the sum of multiple third products corresponding to multiple preset Internet protocol addresses included in each target address cluster; Obtain the sum of the node degrees of multiple target nodes corresponding to each target address cluster to obtain a target degree sum; Based on the ratio of the feature sum to the target degree sum, obtain the feature centroid of each target address cluster.

[0015] In some embodiments, the second calculation module is further configured to: Calculate the similarity between each target feature in the target feature vector corresponding to the Internet protocol address to be analyzed and each sub-feature centroid in each feature centroid to obtain a corresponding feature similarity; Obtain the feature centroid importance score corresponding to each sub-feature centroid, and based on the product of the feature similarity and the feature centroid importance score, obtain the target feature similarity corresponding to each target feature; Add the multiple target feature similarities corresponding to multiple target features to obtain the total feature similarity corresponding to each feature centroid; Obtain a preset third dynamic threshold, and compare the multiple total feature similarities corresponding to the multiple feature centroids with the third dynamic threshold in sequence to obtain a comparison result; When the comparison result indicates that there is a target total feature similarity greater than the third dynamic threshold among the multiple total feature similarities, the corresponding target address cluster is determined as the target homologous cluster of the Internet protocol address to be analyzed, and based on the target homologous cluster, the network attack homologous analysis result of the Internet protocol address to be analyzed is obtained.

[0016] Correspondingly, a third aspect of the embodiments of the present application proposes a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the network attack homologous analysis method according to any one of the embodiments of the first aspect of the present application.

[0017] Correspondingly, a fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the network attack homologous analysis method according to any one of the embodiments of the first aspect of the present application.

[0018] This application obtains multiple preset Internet Protocol (IP) addresses and a preset weighted attack behavior graph. The preset weighted attack behavior graph includes multiple preset IP addresses and multiple connection weights, where each connection weight represents the degree of association between two preset IP addresses with a homologous relationship. The multiple preset IP addresses in the preset weighted attack behavior graph are divided into multiple initial address clusters. For each initial address cluster, the connection weights associated with each preset IP address included are determined, and the corresponding modularity is calculated based on the connection weights. The modularity is used to characterize the homologous compatibility degree among the multiple preset IP addresses included in the corresponding initial address cluster. According to the modularity corresponding to each initial address cluster, the attribution relationship between each preset IP address and the initial address cluster is re-determined to obtain multiple target address clusters, and the characteristic centroid of each target address cluster is determined. An IP address to be analyzed is obtained, and the similarity between the target feature vector of the IP address to be analyzed and the multiple characteristic centroids of the multiple target address clusters is calculated to obtain the network attack homologous analysis result of the IP address to be analyzed. In this way, the connection weights between multiple IP addresses in the preset weighted attack behavior graph generated in advance can be used to accurately quantify the relevance of attack behaviors between any two IP addresses, so as to improve the interpretability of the relationships between various IP addresses, and further improve the accuracy of homologous analysis. Moreover, by adopting the method of modularity iterative processing, the preset weighted attack behavior graph is divided into target address clusters, and the characteristic centroid of each target address cluster is calculated. In this way, a globally representative core feature vector can be accurately generated based on the homologous IP addresses included in the same address cluster. When there is an IP address to be analyzed that requires homologous analysis subsequently, the IP address to be analyzed can be directly compared with the characteristic centroids of each target address cluster (including a large number of homologous IP addresses), without the need to compare and cluster with each of the large number of IP addresses one by one. The homologous determination is made based on the centroid similarity, avoiding the low efficiency problem caused by pairwise comparison of the IP address to be analyzed with a large number of IP addresses. In this way, both the traceability of the attack pattern and the accuracy of homologous analysis are retained, and the efficiency of homologous analysis is improved. In summary, this application can improve the efficiency and accuracy of network attack homologous analysis, which is of great significance for enhancing network security protection capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a schematic diagram of the architecture of the network attack homologous analysis system provided by an embodiment of this application; Figure 2 is a flowchart of the network attack homologous analysis method provided by an embodiment of this application; Figure 3 is the overall flowchart of the network attack homologous analysis method provided by an embodiment of this application; Figure 4 It is a schematic diagram of the functional modules of the network attack homology analysis device provided by an embodiment of the present application; Figure 5 It is a schematic diagram of the hardware structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0020] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0021] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application, and are not intended to limit this application.

[0023] A network attack is an act that uses computer network technology to illegally invade, interfere with, damage or steal information resources in a network system to achieve specific goals through illegal means. It can include, but is not limited to, various forms such as data leakage, system damage, denial-of-service attacks, and malicious software propagation, which pose a serious threat to personal privacy, corporate interests, social order and security.

[0024] In order to more effectively respond to complex and changing network attacks and enhance the overall protection ability of network security, the correlation between different attack events can be identified through network attack homology analysis, and it can be judged whether they come from the same attack source or organization, so as to reveal the attacker's behavior patterns and technical means, improve the defense efficiency, and reduce resource waste.

[0025] In the related art, the Internet protocol address corresponding to each attack event is usually used as a node, and the feature similarity between all Internet protocol addresses is compared pairwise, and the Internet protocol addresses with high feature similarity are clustered to identify the homologous relationship between attackers. However, on the one hand, the method of clustering after pairwise comparison has a high computational complexity and low analysis efficiency when facing a large amount of attack data; on the other hand, the clustering result lacks interpretability, resulting in security personnel being unable to effectively trace the attack pattern, and thus misjudging or missing the homologous relationship, reducing the accuracy of the analysis result.

[0026] Based on this, the embodiments of the present application provide a method, apparatus, computer device, and storage medium for network attack homology analysis. The present application proposes a method, apparatus, computer device, and storage medium for network attack homology analysis, which can improve the efficiency and accuracy of network attack homology analysis.

[0027] The method, apparatus, computer device, and storage medium for network attack homology analysis provided by the embodiments of the present application will be specifically described through the following embodiments. First, the network attack homology analysis system in the embodiments of the present application will be described.

[0028] Please refer to Figure 1 , in some embodiments, the embodiments of the present application provide a network attack homology analysis system, including a terminal 11 and a server side 12.

[0029] Exemplarily, the terminal 11 may be a network security detection device, a personal computer or workstation, a mobile computing device, etc. The server side 12 may be a data center server, a cloud server, a dedicated computing cluster, etc.

[0030] Further, the terminal 11 may collect original attack data from the network environment and transmit the original attack data to the server side 12 through the network for further processing. When the server side 12 receives data from the terminal 11, it can perform a series of preprocessing operations such as cleaning and feature extraction on it to obtain multiple preset Internet protocol addresses, and construct a preset weighted attack behavior graph for the homology relationship between the multiple preset Internet protocol addresses. Then, the preset weighted attack behavior graph is divided to obtain multiple target address clusters, and the feature centroid of each target address cluster is calculated. When the terminal 11 collects an Internet protocol address to be analyzed for homology analysis, it can send the Internet protocol address to be analyzed to the server side 12, so that the server side 12 can calculate the Internet protocol address to be analyzed and each feature centroid one by one to obtain the corresponding network attack homology analysis result, and feedback the network attack homology analysis result to the terminal 11.

[0031] Further, the server side 12 can regularly update the preset weighted attack behavior graph, target address clusters, and algorithm models, and push the latest updates to the terminal 11 to ensure that the entire system can cope with new threats. At the same time, the terminal 11 can also send new data or request further analysis support to the server side 12.

[0032] The network attack homology analysis method in the embodiments of the present application will be described through the following embodiments.

[0033] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0034] In the embodiments of the present application, a description will be made from the dimension of a network attack homology analysis device, which can be specifically integrated in a computer device. Refer to Figure 2 , Figure 2 which is a flowchart of the steps of the network attack homology analysis method provided by the embodiments of the present application. In the embodiments of the present application, taking the network attack homology analysis device being specifically integrated in a terminal or a server as an example, when the processor on the terminal or the server executes the program instructions corresponding to the network attack homology analysis method, the specific process is as follows: Step 101, obtain a plurality of preset Internet Protocol addresses and a preset weighted attack behavior graph, where the preset weighted attack behavior graph includes a plurality of preset Internet Protocol addresses and a plurality of connection weights, and each connection weight represents the degree of association between two corresponding preset Internet Protocol addresses with a homology relationship.

[0035] In some embodiments, in order to quantify and visualize the degree of association between different attacking Internet Protocol addresses, a preset weighted attack behavior graph pre-constructed according to a plurality of preset Internet Protocol addresses can be obtained, and the plurality of preset Internet Protocol addresses can be processed to provide a visual framework to help security analysts identify potential attack groups and provide the necessary data basis for subsequent in-depth analysis.

[0036] Among them, the preset Internet Protocol address is the Internet Protocol (IP) address, which can be a set of Internet Protocol addresses predefined through historical network attack data. Each preset Internet Protocol address is a unique address used to identify a device on the network, and each Internet Protocol address is marked as a potential attack source. For example, the preset Internet Protocol address can be 192.168.1.100, 203.0.113.5, etc.

[0037] Among them, the preset weighted attack behavior graph can be a graphical structure constructed based on the relationships between attackers (preset Internet protocol addresses), where nodes represent different preset Internet protocol addresses, edges represent the same-source relationships between preset Internet protocol addresses, and each edge has a connection weight.

[0038] Among them, the connection weight can be the value carried by the edge between two preset Internet protocol addresses in the preset weighted attack behavior graph, used to quantify the degree of the same-source relationship between these two preset Internet protocol addresses.

[0039] Among them, the correlation degree can be used to characterize the tightness of the relationship between different preset Internet protocol addresses. It is an index calculated comprehensively based on various features (such as spatio-temporal features, behavior fingerprints, association graphs, etc.) included in multiple feature interaction dimensions between preset Internet protocol addresses, and is used to measure whether the preset Internet protocol addresses belong to the same same-source attack group.

[0040] In some embodiments, it is possible to vectorize the features in multiple feature interaction dimensions such as the spatio-temporal feature dimension, behavior fingerprint dimension, and association graph dimension between preset Internet protocol addresses through a pre-acquired set of preset Internet protocol addresses (such as the Internet protocol addresses included in firewall interception records, the Internet protocol addresses of advanced persistent threat attacks captured by honeypots, etc.), to obtain a first feature vector between two Internet protocol addresses. For example, the first feature vector can be: [0,1,0,0,0,1,0.72,0.33,0.82,1,0,1,0.33,1]; Among them, the feature interaction dimensions of each first eigenvalue in the above first feature vector are respectively the same B segment, autonomous system number (ASN), same city + network type, time Kullback-Leibler (KL) divergence, attack type Jaccard similarity index, attack payload text similarity, special pattern, password Jaccard similarity index, honeypot access interval, service overlap degree, and number of common domain names. The above first feature vector is only an example. In actual situations, the feature interaction dimensions and first eigenvalues may be different and can be obtained according to the actual situation.

[0041] Furthermore, it is possible to use an initial gradient boosting model to calculate the feature importance scores of each first eigenvalue in multiple feature interaction dimensions included in the first feature vector, and then obtain the feature importance scores for the determination of the same-source relationship between two preset Internet protocol addresses for each first eigenvalue. According to the multiple feature importance scores of the multiple first eigenvalues corresponding to the first feature vector, the connection weight between two preset Internet protocol addresses can be calculated.

[0042] Further, by taking each preset Internet protocol address as the target node of the preset weighted attack behavior graph, taking the homology relationship between any two Internet protocol addresses as the edge, and taking the corresponding connection weight as the weight of the edge, the preset weighted attack behavior graph can be constructed.

[0043] In the above way, the preset weighted attack behavior graph can be obtained to visually present the complex association relationships between Internet protocol addresses, which helps security analysts quickly understand the overall architecture of attack behaviors and potential attack groups, helps subsequent division of address clusters, and improves the efficiency of homology analysis.

[0044] In some embodiments, in order to improve the accuracy of homology analysis results, weighted attack behavior graphs corresponding to multiple preset Internet protocol addresses (i.e., all collected attack Internet protocol addresses) can be established to visually display the degree of closeness of the connections between different Internet protocol addresses, and further display the logical basis for homology determination. For example, step 101 may include: (101.1) Obtain multiple preset Internet protocol addresses and the homology relationships between the multiple preset Internet protocol addresses, and determine the first feature vector and relationship label between any two preset Internet protocol addresses, where the first feature vector includes multiple first feature values of any two preset Internet protocol addresses in multiple feature interaction dimensions; (101.2) Input each first feature vector and the corresponding relationship label into the initial gradient boosting model in sequence, and perform decision tree split gain accumulation through the initial gradient boosting model to obtain the feature importance score corresponding to each first feature value; (101.3) Based on multiple feature importance scores, determine the connection weight of the two preset Internet protocol addresses corresponding to the first feature vector; (101.4) Based on multiple preset Internet protocol addresses, the relationship labels between any two preset Internet protocol addresses, and the corresponding connection weights, construct the preset weighted attack behavior graphs corresponding to the multiple preset Internet protocol addresses.

[0045] Among them, the homology relationship may include non-homology and homology. When the homology relationship represents homology between two preset Internet protocol addresses, it represents the relationship that the two preset Internet protocol addresses are considered to belong to the same attack group due to features such as similar behavior patterns and attack methods, otherwise the homology relationship between the two Internet protocol addresses is considered non-homologous.

[0046] Among them, the first feature vector may be a feature set representing the features between any two preset Internet protocol addresses, and this feature set contains the specific numerical manifestations of these two preset Internet protocol addresses in multiple feature interaction dimensions.

[0047] Among them, the relationship tag can be a mark indicating whether any two preset Internet protocol addresses are of the same origin, usually represented by 0 or 1 (1 indicates the existence of the same origin, and 0 indicates non-same origin).

[0048] Among them, the feature interaction dimension can be different perspectives or aspects for evaluating the correlation between two preset Internet protocol addresses, such as spatio-temporal features, behavior fingerprints, and association graphs, etc. Among them, spatio-temporal features, behavior fingerprints, and association graphs can be further refined.

[0049] Among them, the first eigenvalue can be an element in the first eigenvector, representing the specific similarity between any two preset Internet protocol addresses in a specific feature interaction dimension.

[0050] Among them, the initial gradient boosting model can be a machine learning model based on the gradient boosting algorithm, used to learn the importance of each first eigenvalue for the same-origin determination from the first eigenvector and its corresponding relationship tag, and optimize the model through the decision tree splitting gain accumulation process, finally obtaining the target gradient boosting model, and outputting the feature importance score corresponding to each first eigenvalue according to the accumulated splitting gain. Exemplarily, the initial gradient boosting model can be a Lightweight Gradient Boosting Decision Tree (GBDT) model, such as the LightGBM model.

[0051] Among them, the feature importance score can be a quantitative index calculated by the initial gradient boosting model to represent the importance degree of each first eigenvalue for determining the same-origin relationship between two preset Internet protocol addresses.

[0052] Exemplarily, if there are 3 preset Internet protocol addresses (only for example here, in fact, multiple or a large number of preset Internet protocol addresses can be used to construct a comprehensive preset weighted attack behavior graph): IP1 = 192.168.1.1, IP2 = 192.168.1.2, IP3 = 192.168.2.1, the same-origin relationships of these 3 preset Internet protocol addresses have been pre-annotated. IP1 and IP2 are of the same origin (thus determining the relationship tag = 1), IP1 and IP3 are not of the same origin (thus determining the relationship tag = 0), and IP2 and IP3 are not of the same origin (thus determining the relationship tag = 0).

[0053] In some embodiments, a first feature vector between any two preset Internet protocol addresses can be obtained. Specifically, a plurality of first sub-features corresponding to the first preset Internet protocol address in a plurality of feature interaction dimensions, and a plurality of second sub-features corresponding to the second preset Internet protocol address in a plurality of feature interaction dimensions can be obtained, where the plurality of feature interaction dimensions include a spatio-temporal feature dimension, a behavior fingerprint dimension, and an association graph dimension; then, the behavior overlap degree between each first sub-feature and the corresponding second sub-feature is obtained, and the behavior overlap degree is encoded to obtain a corresponding first feature value. Based on the plurality of first feature values corresponding to the plurality of feature interaction dimensions, the first feature vector between any two preset Internet protocol addresses can be obtained.

[0054] Exemplarily, an example is given for obtaining the first feature vector between any two preset Internet protocol addresses. Specifically, in the process of determining the behavior overlap degree between each first sub-feature and the corresponding second sub-feature, taking two preset Internet protocol addresses 192.168.1.1 and 192.168.1.2 as an example, in the example process, IP1 is denoted as IP1, and IP2 is denoted as IP2. The spatio-temporal features can characterize the attacker's behavior pattern from three dimensions: network layer features, geographical features, and time pattern features. Among them, the network layer features can include the same C-segment feature and the same B-segment feature, that is, to determine whether two preset Internet protocol addresses belong to the same subnet segment (for example, IP1 and IP2 are in the same C-segment). The network layer features can also include the autonomous system number to determine whether two preset Internet protocol addresses belong to the same autonomous system (for example, IP1 and IP2 both belong to AS12345), and can also include the IP property (for example, IP1 is the preset Internet protocol address of a data center) to determine whether two preset Internet protocol addresses are of specific types such as data centers and enterprise dedicated lines; the geographical features can include the features obtained by associating the geographical location and the network type (such as home broadband, enterprise dedicated line), for example, both IP1 and IP2 are located in City T and are enterprise dedicated line users; the time pattern features can be calculated by the similarity of the attack time distribution (for example, the KL divergence of the attack time distribution of two preset Internet protocol addresses). For example, the KL divergence of the attack time distribution of IP1 and IP2 is 0.2, indicating that their time patterns are very similar.

[0055] Exemplarily, the behavior fingerprint can be calculated from three dimensions: attack vector, attack payload feature, and password feature. Among them, the attack vector can include the Jaccard similarity of attack types (for example, the Jaccard similarity of IP1 and IP2 using the same SQL injection attack method is 0.9) and the overlap degree of malicious domains (for example, both IP1 and IP2 use the malicious domain name "mmm.com"); the payload feature can include text similarity (for example, the text similarity of the payloads of IP1 and IP2 is 0.85) and special patterns (for example, both IP1 and IP2 contain the same URL parameter structure); the password feature can include the Jaccard similarity of the top 100 passwords. For example, 70 out of the first 100 passwords tried by IP1 and IP2 are the same, and the Jaccard similarity is 0.7.

[0056] Furthermore, the associated graph dimension can be determined from three dimensions: honeypot linkage, business association, and domain name mapping. Among them, the honeypot linkage can be determined by determining that the access interval of the same honeypot point is less than <X> hours. The time interval between IP1 and IP2 accessing the same honeypot is less than 2 hours, indicating that there may be a connection between the two; the business association can be determined by the overlap degree of the attacked business system sets. For example, both IP1 and IP2 attacked the information systems of Bank A and Hospital B; the domain name mapping can be determined by the number of co-occurring domain names in threat intelligence. For example, both IP1 and IP2 are associated with the domain name "mx.com".

[0057] Furthermore, after obtaining the behavior overlap degree of the first sub-feature and the second sub-feature, the behavior overlap degree can be encoded in the following way. For numerical features (such as KL divergence of time distribution, Jaccard similarity, etc.), the original value can be directly used; for categorical features (such as IP nature, ASN number, etc.), One-Hot encoding can be adopted; for text features (such as payload text, etc.), the cosine similarity can be calculated after vectorization based on TF-IDF. For boolean features (such as in the same C segment, in the same B segment, etc.), they can be converted into binary values (1 means yes, 0 means no, etc.).

[0058] For example, during encoding, the following conditions are met: same C segment: 1 (in the same C segment); ASN number: 1 (belonging to the same autonomous system AS12345); IP property: 1 (both are enterprise dedicated lines); Jaccard similarity of attack types: 0.9 (using the same method to inject attacks); overlap degree of malicious domains: 1 (both use "mmm.com"); text similarity: 0.85 (high payload text similarity); special mode: 1 (both use "username=admin&password=1234"); same honeypot access interval: 1 (the access interval to the same honeypot is less than 2 hours); overlap degree of attacked business system sets: 1 (both attack Bank A and Hospital B); number of co-occurring domains in threat intelligence: 1 (both are associated with "badactor.com"). Thus, the first feature vector can be obtained as [1, 1, 1, 0.9, 1, 0.85, 1, 1, 1, 1].

[0059] In some embodiments, a training set for the initial gradient boosting model can be constructed. The training set can include positive samples and negative samples. Among them, the positive sample is the first feature vector with a confidence level > 0.8 for two preset Internet protocol addresses of the same origin and correct manual review, and the label , and the negative sample is two randomly sampled preset Internet protocol address pairs and rule false alarm cases, and the label 0; Then the training set can be expressed as: ; Among them, represents the first feature vector of two preset Internet protocol address pairs (IPi, IPj), represents the relationship label of the IP pair (1 indicates the same origin, 0 indicates non - same origin).

[0060] Furthermore, the initial gradient boosting model (such as the LightGBM model) can be used for training to learn the contribution degree of each first feature value to the same - origin determination result. After the initial gradient boosting model is trained, the target gradient boosting model can be obtained. At the same time, the model will output the feature importance score corresponding to each first feature value , representing the weight corresponding to the k - th first feature value.

[0061] In some embodiments, for each first feature vector, each first feature value it contains can be multiplied by the corresponding feature importance score to obtain the sub - feature weight of each first feature value. Based on the sum of the multiple sub - feature weights corresponding to the multiple first feature values, the connection weight of the two preset Internet protocol addresses corresponding to the first feature vector can be obtained. Thus, the association strength between any two preset Internet protocol addresses can be accurately reflected.

[0062] Furthermore, a preset weighted attack behavior graph can be constructed based on all preset Internet protocol addresses as target nodes, the relationship tags (i.e., homologous or non-homologous) between any two preset Internet protocol addresses as the connection edges between the target nodes, and the corresponding connection weights as the weights of the connection edges. In this way, the association strength between preset Internet protocol addresses can be quantified, and potential homologous attack groups can be intuitively displayed, thus helping security analysts to more efficiently and accurately identify and understand complex network attack patterns and improve the overall network security protection ability.

[0063] By constructing a preset weighted attack behavior graph, a comprehensive and detailed data basis for attack behavior can be provided for subsequent analysis, ensuring the accuracy and comprehensiveness of the analysis, and also providing strong support for the formulation of network security defense strategies, helping to detect and respond to potential network security threats in a timely manner.

[0064] In some embodiments, to improve the accuracy of the model in identifying the homologous relationship of network attacks, the initial gradient boosting model can be trained to accurately quantify the importance of each first eigenvalue of each first feature vector in different feature interaction dimensions in determining the homologous relationship between two preset Internet protocol addresses, so that the model can continuously improve the rationality of the assigned weights, thereby facilitating the traceability of homologous analysis and further improving the accuracy and reliability of homologous analysis. For example, (101.2) may include: (101.2.1) Sequentially input each first feature vector and the corresponding relationship tag into the initial gradient boosting model, and through the initial gradient boosting model, predict each first feature vector to obtain the corresponding predicted tag; (101.2.2) Determine the first target loss based on the difference between the predicted tag and the relationship tag; (101.2.3) Based on the first target loss, calculate the split gain of each of the multiple first eigenvalues included in each first feature vector, and determine the intermediate split feature with the largest split gain from the multiple first eigenvalues for node splitting, and calculate the second target loss; (101.2.4) Based on the difference between the second target loss and the first target loss, calculate the split gain of each of the multiple first eigenvalues included in the first feature vector, and determine the intermediate split feature with the largest split gain from the multiple first eigenvalues for node splitting, and calculate the updated second target loss; (101.2.5) Repeatedly execute the step of calculating the split gain for multiple first eigenvalues included in the first feature vector based on the difference between the updated second objective loss and the first objective loss, determining the intermediate split feature with the largest split gain from the multiple first eigenvalues for node splitting, and calculating to obtain the updated second objective loss until the preset number of training times is reached. According to the multiple split gains iteratively obtained for each eigenvalue, obtain the feature importance score corresponding to each first eigenvalue.

[0065] Among them, the predicted label can be the label obtained by predicting each first feature vector through the initial gradient boosting model, and is used to represent whether there is a homologous relationship (such as 0 or 1) between two preset Internet protocol addresses judged by the model.

[0066] Among them, the first objective loss can be used to measure the error between the current prediction result (predicted label) of the model and the actual situation (relationship label).

[0067] Among them, the split gain can be, in a decision tree, when the intermediate split feature is selected as the split point, the information gain or the reduced loss value brought by this eigenvalue. The larger the split gain, the more effectively this eigenvalue can distinguish data of different categories, that is, the larger the feature importance score of the corresponding intermediate split feature.

[0068] Among them, the intermediate split feature can be the eigenvalue that can maximize the split gain calculated according to the split gain in each node splitting process, that is, the optimal split point.

[0069] Among them, the second objective loss can be the loss value after each node splitting, reflecting the performance of the model after using the selected intermediate split feature, used to evaluate the effect of splitting, and provide a basis for further optimization.

[0070] Exemplarily, the model can be initialized first. For example, the maximum tree depth of the model can be set to 8, the minimum number of samples in the leaf node is 10, and the loss function is set to the logarithmic loss function. Specifically, in the first round of iteration, the initial gradient boosting model (generally a single tree) calculates the prediction probability for each first feature vector: ; Among them, is the Sigmoid function, is the output of the t-th tree.

[0071] Furthermore, if the initial gradient boosting model outputs the predicted label It is 0.75, indicating a 75% probability of determining homology. At this time, the difference between the predicted label and the true relationship label can be quantified by calculating the first objective loss to drive the optimization of the initial gradient boosting model. Exemplarily, the first objective loss and the second objective loss can be calculated using loss functions such as logarithmic loss, cross-entropy loss, etc., and this application does not limit this too much.

[0072] Further, after the initial gradient boosting model finishes predicting, if there is a significant deviation between the predicted label of the initial gradient boosting model and the relationship label, or the number of training rounds has not reached the preset number of rounds, the feature and split point with the largest split gain can be selected through a greedy strategy to gradually optimize the tree structure. Specifically, the values of each feature can be discretized into a histogram. For example, when the feature interaction dimension is the time KL divergence, the first feature value can be divided into intervals such as [0, 0.1), [0.1, 0.2), etc., and the split gain of each candidate split point (i.e., the feature interaction dimension that can be used for splitting) is calculated, and all the first feature values and their split points are traversed to select the split method with the largest gain. For example, the split gain of the same C segment of the feature is 12.5, and the split gain of the feature time KL divergence is 8.3, then the split with the largest gain, i.e., the same C segment, can be selected to update the tree structure and calculate the second objective loss after splitting.

[0073] Further, for each newly added tree, it can be fitted based on the residuals of the previous model to gradually approximate the true label. The second objective loss after splitting is compared with the first objective loss. If the second objective loss is less than the first objective loss and the difference between the second objective loss and the first objective loss is greater than the preset threshold, continue splitting; otherwise, trigger early stopping.

[0074] For example, after the first split, the loss drops to 0.58. For the second split, the time KL divergence is selected as the intermediate split feature, specifically, the time KL divergence > 0.2, and the loss drops to 0.52. This continues to iterate until the loss converges or reaches the maximum number of trees. At this point, the sum of the split gains of each feature interaction dimension (i.e., the feature interaction dimension corresponding to each first feature value) in all decision trees can be accumulated, and the importance score corresponding to each first feature value is obtained by normalization.

[0075] Through the above method, the model can automatically learn the contribution degree of the feature interaction dimension corresponding to each first feature value to homology determination, so as to provide a scientific feature weight assignment for network attack homology analysis, facilitating improving the accuracy and reliability of the analysis results when conducting homology analysis on network attacks subsequently.

[0076] In some embodiments, to enhance the rationality and interpretability of connection weight calculation, the sub-feature weight can be calculated by multiplying each first eigenvalue by the corresponding target feature importance score, quantifying the specific impact of a single feature on the connection weight, and finally determining the connection weight between two preset Internet protocol addresses, so as to ensure the fairness and rationality of feature weight allocation and improve the reliability and interpretability of the analysis results. For example, (101.3) may include: (101.3.1) Normalize each feature importance score to obtain the corresponding target feature importance score; (101.3.2) Obtain the sub-feature weight of each first eigenvalue according to the product of each first eigenvalue and the corresponding target feature importance score; (101.3.3) Based on the sum of the sub-feature weights corresponding to multiple first eigenvalues, obtain the connection weight between two preset Internet protocol addresses corresponding to the first eigenvector.

[0077] Among them, the target feature importance score can be the result obtained by normalizing the feature importance score of each first eigenvalue, aiming to make the importance scores of different features comparable on the same scale, so as to more accurately evaluate the relative importance of each feature in the model.

[0078] Among them, the sub-feature weight can be the product of each first eigenvalue and its corresponding target feature importance score, indicating the proportion of this eigenvalue in determining the connection strength between two preset Internet protocol addresses.

[0079] In some embodiments, each feature importance score can be normalized by the following formula to obtain the corresponding target feature importance score: ; Wherein, represents the feature importance score of the k-th first eigenvalue, represents the normalized target feature importance score, and n represents the total number of first eigenvalues included in the first eigenvector.

[0080] In some embodiments, the sub-feature weight is also the feature contribution degree of the corresponding first eigenvalue in the homologous determination of the first eigenvector, and the sub-feature weight can be calculated by the following formula: ; Wherein, represents the feature contribution degree (i.e., sub-feature weight) of the k-th feature, represents the normalized target feature importance score, Represents the corresponding first eigenvalue. By calculating the feature contribution degree corresponding to each first eigenvalue, it is convenient to perform result tracing and visual analysis during the subsequent process of network attack homology analysis.

[0081] Furthermore, the sub-feature weights (i.e., feature contribution degrees) can be normalized: ; where represents the sub-feature weight of the i-th feature, and n represents the total number of first eigenvalues, which is also the total number of sub-feature weights.

[0082] Furthermore, based on the sum of the normalized sub-feature weights corresponding to multiple first eigenvalues, the connection weight between two preset Internet protocol addresses corresponding to the first eigenvector can be obtained: ; represents the normalized sub-feature weight.

[0083] Through the above method, the contribution degree of different first eigenvalues to determining the homology relationship between two preset Internet protocol addresses can be effectively quantified, thereby obtaining a comprehensive connection weight. In this way, not only the accuracy of homology analysis is improved, but also the interpretability of the preset weighted attack behavior graph is enhanced, facilitating result tracing and visual analysis.

[0084] In some embodiments, in order to update the target relationship label in a timely manner according to the latest data, the first dynamic threshold can be determined by introducing statistics such as historical connection weights, average scores, and standard deviations of scores, as well as flexibly using adjustment parameters, so as to obtain the homology relationship between any two preset Internet protocol addresses. In this way, it can ensure that even in the case of rapid changes in the network environment, the effective tracking and analysis of network attack behaviors can be maintained, enhancing the robustness and adaptability of the overall network security defense system. Exemplarily, before constructing the preset weighted attack behavior graph corresponding to multiple preset Internet protocol addresses, that is, before (101.4), it may further include: (A.1) Obtain multiple historical connection weights between any two preset Internet protocol addresses, as well as the average score and standard deviation of the scores of the multiple historical connection weights; (A.2) Obtain a preset adjustment parameter, and obtain the first product according to the product of the adjustment parameter and the standard deviation of the score; (A.3) Determine the first dynamic threshold according to the difference between the average score and the first product; (A.4) Compare the connection weight between any two preset Internet protocol addresses with the first dynamic threshold to obtain a comparison result; (A.5) Based on the comparison result, update the relationship label between any two preset Internet protocol addresses to obtain the target relationship label; Then, based on multiple preset Internet protocol addresses, the relationship labels between any two preset Internet protocol addresses, and the corresponding connection weights, construct a preset weighted attack behavior graph corresponding to the multiple preset Internet protocol addresses, including: Based on multiple preset Internet protocol addresses, the target relationship labels between any two preset Internet protocol addresses, and the corresponding connection weights, construct a preset weighted attack behavior graph corresponding to the multiple preset Internet protocol addresses.

[0085] Among them, the historical connection weight can be the connection weight calculated for two preset Internet protocol addresses at multiple past time points.

[0086] Among them, the average score can be the average of multiple historical connection weights, which is used to represent the general level or central tendency of the historical connection weights.

[0087] Among them, the standard deviation of the score can be used to measure the degree of dispersion of the distribution of multiple historical connection weights, and is used to reflect the fluctuation of the historical connection weights around the average score.

[0088] Among them, the adjustment parameter can be a preset parameter used to adjust the sensitivity of the dynamic threshold. After multiplying the adjustment parameter by the standard deviation of the score, it can control the range of the first dynamic threshold to adapt to the requirements of different application scenarios.

[0089] Among them, the first product can be the result obtained by multiplying the adjustment parameter by the standard deviation of the score, which is used to provide a flexibility factor when determining the first dynamic threshold.

[0090] Among them, the first dynamic threshold can be a threshold obtained based on the difference between the average score and the first product, which is used to determine whether the current connection weight significantly deviates from the historical average level, so as to decide whether to update the relationship label.

[0091] Among them, the comparison result can be the result of comparing the current connection weight between any two preset Internet protocol addresses with the first dynamic threshold, indicating whether the correlation between the two preset Internet protocol addresses has changed significantly.

[0092] Among them, the target relationship label can be a new label (such as 0 or 1) indicating whether there is a homologous relationship between any two preset Internet protocol addresses updated according to the comparison result, which is used to reflect the latest analysis conclusion.

[0093] In some embodiments, the first dynamic threshold has the following calculation formula: ; Among them, represents the average score, represents the standard deviation of the score, represents the adjustment parameter.

[0094] Exemplarily, if the multiple historical connection weights between any two preset Internet protocol addresses (such as IP1 and IP2) include: w1 = 0.8, w2 = 0.7, w3 = 0.9, w4 = 0.6, w5 = 0.8. Then, the average score can be calculated as 0.76, and the standard deviation of the score is 0.098.

[0095] Furthermore, if the adjustment parameter is set to 2, according to the product of the preset adjustment parameter and the standard deviation of the score, the first product = 2 × 0.098 = 0.196. Then, according to the difference between the average score and the first product, the first dynamic threshold can be determined as 0.76 - 0.196 = 0.564.

[0096] Exemplarily, the first dynamic threshold calculated from these historical connection weights can be compared with the connection weight between its corresponding two preset Internet protocol addresses to determine whether it exceeds the first dynamic threshold. Assume that the current connection weights of IP1 and IP2 are 0.7, 0.7 > 0.564, then the comparison result is exceeded, and the relationship label is updated to homologous (relationship label is 1); otherwise, it is updated to non - homologous (label is 0).

[0097] It should be noted that the method for determining the relationship label of two Internet protocol addresses in (A.1) to (A.5) can also be applied to "obtaining multiple preset Internet protocol addresses and the homologous relationships between multiple preset Internet protocol addresses" in (101.1). When it is necessary to determine the relationship label of two Internet protocol addresses at any time, the above formula can be used to determine the homologous relationship. When there is no historical connection weight, the first dynamic threshold for determining the relationship label between two preset Internet protocol addresses can be set manually or through other algorithms, and the present application does not limit this too much.

[0098] Through the above method, the threshold can be dynamically adjusted, and the relationship label can be updated in real - time to adapt to the changes in network attack behaviors, improving the accuracy and flexibility of homologous analysis.

[0099] Step 102, divide the multiple preset Internet protocol addresses in the preset weighted attack behavior graph into multiple initial address clusters.

[0100] In some embodiments, in order to enable the system to quickly identify preset Internet protocol addresses with close connections, that is, preset Internet protocol addresses that may belong to the same attack group, multiple preset Internet protocol addresses in the preset weighted attack behavior graph can be first divided into multiple initial address clusters, so as to facilitate subsequent continuous optimization of each cluster by calculating modularity, and then enable the subsequent community discovery algorithm to operate more efficiently.

[0101] Among them, the initial address clusters can be multiple groups or clusters initially obtained by dividing multiple preset Internet protocol addresses according to the connection weights and feature similarities of the preset Internet protocol addresses.

[0102] In some embodiments, the initial address clusters can also be obtained by randomly dividing the preset weighted attack behavior graph.

[0103] Exemplarily, each target node (i.e., the preset Internet protocol address) in the preset weighted attack behavior graph can be divided according to the similarity relationship between any two preset Internet protocol addresses in the attack behavior feature library. For example, if two preset Internet protocol addresses have similar features in the attack behavior feature library, then these two preset Internet protocol addresses can be divided into the same initial address cluster.

[0104] In some embodiments, the attack behavior feature library mainly includes three parts: spatio-temporal features, behavior fingerprints, and association graphs. Their meanings have been introduced in detail when introducing the first feature vector between two preset Internet protocol addresses above, and will not be elaborated here one by one. Next, the process of constructing the attack behavior feature library will be briefly introduced. First, the spatio-temporal features can describe the behavior patterns of attackers from the network layer (such as the same C segment, the same B segment, ASN number, etc.), geographical features (such as city + network type, etc.), and time patterns (such as KL divergence of attack time distribution, etc.); secondly, the behavior fingerprints describe the specific technical means and habits of attackers by analyzing attack vectors (such as Jaccard similarity of attack types, overlap degree of malicious domains), payload features (such as text similarity, special patterns), and password features (such as Jaccard similarity of simple passwords); finally, the association graph reveals the relationships between attackers from three dimensions: honeypot linkage (such as access interval of the same honeypot point), service association (overlap degree of the attacked service system set), and domain name mapping (number of domains that co-occur in threat intelligence). The process of constructing this attack behavior feature library relies on the experience design of network security experts to ensure its high interpretability and adaptability.

[0105] In some embodiments, multiple initial address clusters can also be obtained by partitioning based on the connection weights of multiple preset Internet protocol addresses in a preset weighted attack behavior graph. Specifically, when the connection weight between any two preset Internet protocol addresses in the preset weighted attack behavior graph is greater than a preset threshold, these two preset Internet protocol addresses can be partitioned into the same initial address cluster. The preset threshold can be set to 0.6, 0.7, etc.

[0106] Through the above method, the initial address cluster can be determined as the starting point of community evolution, facilitating subsequent updates of the address cluster.

[0107] Step 103: For each initial address cluster, determine the connection weights associated with each preset Internet protocol address included, and calculate the corresponding modularity according to the connection weights. The modularity is used to characterize the degree of homology and compatibility among the multiple preset Internet protocol addresses included in the corresponding initial address cluster.

[0108] In some embodiments, in order to enable security analysts to more accurately locate specific attack groups, the modularity can be used as an evaluation criterion to identify address clusters with high internal connectivity and consistent behavior patterns, or to identify clusters with poor compatibility, so as to dynamically adjust and optimize the cluster structure, thereby improving the accuracy of network attack homology analysis.

[0109] Among them, the modularity can be an index used to quantify the connection strength and internal relevance among the preset Internet protocol addresses within each initial address cluster. A high modularity means a stronger homologous relationship among the initial Internet protocol addresses within the initial address cluster.

[0110] Among them, the degree of homology and compatibility can be the degree of mutual compatibility and consistency shown among all preset Internet protocol addresses based on the connection weights in a specific initial address cluster. That is to say, the degree of homology and compatibility is used to characterize whether the preset Internet protocol addresses in the same initial address cluster have similar behavior patterns or attack methods, and thus the possibility of being classified into the same attack group.

[0111] In some embodiments, for each preset Internet Protocol address (target node) in the preset weighted attack behavior graph, the sum of all connection weights with other target nodes in the graph (i.e., the node degree) can be statistically calculated. The specific method for obtaining the connection weights has been elaborated above and will not be repeated here. Then, the graph connection weight is calculated, that is, half of the sum of all connection weights in the entire preset weighted attack behavior graph (to avoid double counting of weights) is used as the normalization reference value. Thus, for each initial address cluster, the differences between the following two situations can be compared: the actual connection weights between any two target nodes within the same initial address cluster, and the expected connection weights between target nodes within the same initial address cluster assuming a completely random network connection (determined by the product of the node degrees of the two target nodes and the ratio of the graph connection weight). Thus, the differences between the actual connection weights and the expected connection weights in each initial address cluster can be accumulated and then divided by the graph connection weight to obtain the modularity corresponding to the initial address cluster.

[0112] By calculating the modularity of each initial address cluster, the tightness of the cluster division can be quantified, providing a basis for subsequent optimization (such as node attribution adjustment), so as to better understand and analyze the homology of network attacks.

[0113] In some embodiments, in order to accurately quantify the connection strength and internal consistency between preset Internet Protocol addresses within a cluster, the modularity of each address cluster obtained by division can be calculated, so that the system can more accurately divide the community structure in the network, and then iteratively obtain potential attack groups with a high degree of homology and compatibility. For example, "calculating the corresponding modularity according to the connection weight" in step 103 may include: (103.1) Obtaining the graph connection weight based on the sum of the connection weights corresponding to all connection edges in the preset weighted attack behavior graph; (103.2) Obtaining the node degree corresponding to the target node corresponding to each preset Internet Protocol address in the initial address cluster; (103.3) Obtaining a second product based on the product of the node degrees between any two target nodes; (103.4) Obtaining a first ratio based on the ratio of the second product to the graph connection weight; (103.5) Obtaining the difference between the connection weight corresponding to any two target nodes in the initial address cluster and the corresponding first ratio, to obtain a first difference; (103.6) Obtaining the modularity corresponding to the initial address cluster based on the graph connection weight and the multiple first differences between multiple target nodes included in the initial address cluster.

[0114] Among them, the graph connection weight can be the sum of the connection weights corresponding to all the connection edges in the preset weighted attack behavior graph, which reflects the total association strength among all Internet protocol addresses in the entire graph.

[0115] Among them, the node degree can be the sum of the connection weights indicating the connection between each preset Internet protocol address and all other target nodes corresponding to the target node in the initial address cluster, and is used to measure the importance of the target node in the network.

[0116] Among them, the second product can be the result obtained according to the product of the node degrees between any two target nodes, which reflects the relative importance of these two target nodes in the network and the possible degree of their mutual influence.

[0117] Among them, the first ratio can be a value obtained based on the ratio of the second product to the graph connection weight, which represents the expected connection strength between any two target nodes, that is, under the random network model, the connection weight that these two target nodes should have.

[0118] Among them, the first difference can be the result obtained by obtaining the difference between the connection weight corresponding to any two target nodes in the initial address cluster and the corresponding first ratio, which is used to characterize the difference between the actual connection weight and the expected connection weight, and is used to judge whether the connection between these two target nodes is significantly stronger or weaker than the expected value under random conditions.

[0119] In some embodiments, the modularity calculation formula is as follows: ; Among them represents the connection weight of the connection edge between target node i and target node j, represents the node degree of target node i, represents the sum of all connection weights in the preset weighted attack behavior graph, is an indicator function, which takes the value of 1 when target node i and target node j belong to the same address cluster, and 0 otherwise.

[0120] In some embodiments, , that is, half of the sum of all connection weights. This formula can be used to calculate the global benchmark value for subsequent normalization. Assume that there are 3 edges in the graph with weights of 0.8, 0.5, and 0.3 respectively, then m = (0.8 + 0.5 + 0.3) / 2 = 0.8.

[0121] Exemplarily, to calculate the node degree corresponding to each target node, it can be obtained by adding up all the adjacent connection weights of the target node, that is For example, if the target node A is connected to the target node B and the target node C, and the connection weights are 0.8 and 0.5 respectively, then the node degree of the target node A is 0.8 + 0.5 = 1.3.

[0122] Furthermore, the product of the node degrees between any two target nodes in the initial address cluster can be calculated. For example, if the node degree of the target node A is 1.3 and the node degree of the target node B is 0.8, then the product of the node degrees of the target node A and the target node B, that is, the second product, is 1.3×0.8 = 1.04.

[0123] Furthermore, the first weight That is, the proportion of the connection weight of the expected random connection. For example, if the second product is 1.04 and the graph connection weight m = 0.8, then the first ratio = 1.04 / (2×0.8) = 0.65. Then, the deviation between the actual connection strength between target nodes and the random expectation can be measured by the difference between the connection weight of any two target nodes in the preset weighted attack behavior graph and the first ratio. For example, if the actual connection weight between the target node A and the target node B is 0.8 and the first ratio is 0.65, then the first difference = 0.8 - 0.65 = 0.15.

[0124] In some embodiments, the first differences of all connection edges (that is, the edges corresponding to any two target nodes) in the same initial address cluster can be accumulated and normalized by dividing by 2m to obtain the modularity corresponding to the initial address cluster. For example, if the initial address cluster includes the target nodes A, B, and C, if the first difference between the target node A and the target node B is 0.15, the first difference between the target node A and the target node C is, and the first difference between the target node B and the target node C is -0.1, then the modularity Q = (0.15 + 0.2 - 0.1) / (2×0.8) ≈ 0.156.

[0125] Furthermore, if the modularity > 0, it indicates that the connection strength within the address cluster is higher than the random expectation (for example, Q = 0.15 indicates that the address cluster has significant homologous characteristics); if Q is close to 0 or negative, re - partitioning is required (for example, if the difference of the B - C edge is negative, it may belong to noise, and the address clusters to which the target nodes B and C belong can be re - partitioned).

[0126] Through the above method, the structural characteristics of the preset weighted attack behavior graph can be transformed into multiple quantifiable modularities, providing theoretical support for dynamic cluster partitioning and attack homology determination, so as to more accurately identify and analyze network attack behaviors, and ultimately achieve the goal of accurately mining potential attack groups from massive data.

[0127] Step 104: According to the modularity corresponding to each initial address cluster, re-determine the affiliation relationship between each preset Internet protocol address and the initial address cluster, obtain multiple target address clusters, and determine the characteristic centroid of each target address cluster.

[0128] In some embodiments, in order to form more optimized and accurate target address clusters, the affiliation relationship between the preset Internet protocol addresses and these clusters can be re-evaluated and adjusted based on the modularity corresponding to each initial address cluster, so that the system can more accurately identify groups of attack Internet protocol addresses with high homology, and ensure stronger relevance and consistency among the Internet protocol addresses within each address cluster, thereby improving the efficiency and accuracy of the entire network attack homology analysis.

[0129] Among them, the target address cluster can be a more optimized set of Internet protocol addresses (also called a community) formed after re-determining the affiliation relationship between each preset Internet protocol address and the initial address cluster according to the modularity. The Internet protocol addresses within each target address cluster are considered to have a high degree of homology compatibility, that is, there is a significant similarity in their behavior patterns or attack techniques, and they belong to the same attack group.

[0130] Among them, the characteristic centroid can be the center point or average representative of the characteristic values of all preset Internet protocol addresses within each target address cluster. It is a comprehensive characteristic vector obtained by weighted averaging the characteristic vectors of all Internet protocol addresses within the target address cluster, and is used to describe the main characteristic attributes of the entire target address cluster.

[0131] In some embodiments, the affiliation relationship of the preset Internet protocol addresses can be dynamically adjusted based on the principle of maximizing modularity, so that the internal connection density of the address cluster is significantly higher than the expected value of a random network, and the characteristic centroid is extracted to support subsequent homology determination.

[0132] In some embodiments, when calculating the modularity of each initial address cluster, it is possible to determine the preset Internet protocol addresses for which the first difference between any two target nodes calculated above in the initial address cluster is negative, and re-partition the affiliated address clusters of these preset Internet protocol addresses. Specifically, after copying and re-partitioning these preset Internet protocol addresses into other initial address clusters, recalculate the modularity. In this way, the efficiency and accuracy of the address cluster partitioning can be improved. Or, when the modularity of the initial address cluster is relatively low (for example, lower than a preset modularity threshold), re-partition the affiliated address clusters of each preset Internet protocol address in the initial address cluster until the current modularity obtained by the partitioning is lower than the modularity threshold, and then the re-partitioning process can be stopped to obtain the partitioned target address clusters.

[0133] In some embodiments, after obtaining the target address clusters, for each target address cluster, the feature vectors of the multiple target address clusters it contains can be extracted and weighted averaged to obtain the feature centroid of each target address cluster. When a new Internet Protocol address is added to a target address cluster or the community structure changes, only the centroid of the affected target address cluster needs to be updated, without full-scale calculation.

[0134] Through the above method, the homologous association strength (modularity) of the preset Internet Protocol addresses within the address cluster can be maximized, and by calculating the feature centroid as the cluster fingerprint of each target address cluster, it is convenient to quickly and accurately match the Internet Protocol addresses to be analyzed subsequently, so as to perform accurate homologous analysis on each Internet Protocol address to be analyzed.

[0135] In some embodiments, in order to ensure that each finally formed target address cluster has a high modularity, that is, there is a stronger association and consistency among internal members, the initial address clusters can be iteratively optimized until the modularity of all address clusters exceeds the second dynamic threshold, forming more accurate target address clusters. In this way, not only the accuracy of identifying potential attack groups is improved, but also the adaptability and response speed of the system are enhanced, providing more solid support for network security protection. For example, "re-determine the attribution relationship between each preset Internet Protocol address and the initial address cluster according to the modularity corresponding to each initial address cluster to obtain multiple target address clusters" in step 104 may include: (104.a1) Determine the intermediate address clusters with modularity less than the preset second dynamic threshold according to the modularity corresponding to each initial address cluster; (104.a2) Obtain the preset attack behavior feature library, and based on the attack behavior feature library, re-determine the associated Internet Protocol addresses having an attack feature relationship with each preset Internet Protocol address in each intermediate address cluster, and migrate each preset Internet Protocol address to the initial address cluster corresponding to the associated Internet Protocol address; (104.a3) Repeat the step of re-determining the associated Internet Protocol addresses having an attack feature relationship with each preset Internet Protocol address in each intermediate address cluster based on the attack behavior feature library, and migrating each preset Internet Protocol address to the initial address cluster corresponding to the associated Internet Protocol address until the modularity corresponding to each intermediate address cluster is greater than the second dynamic threshold, obtaining multiple target address clusters.

[0136] Among them, the second dynamic threshold can be a threshold for evaluating whether an address cluster needs further adjustment. When the modularity of an address cluster is lower than the second dynamic threshold, it indicates that the association between the preset Internet Protocol addresses within the address cluster is not strong enough, and the cluster structure needs to be optimized by reallocating the preset Internet Protocol addresses it contains.

[0137] Among them, the intermediate address cluster can be an address cluster that still needs to be further divided.

[0138] Among them, the attack behavior feature library can be a data set containing various network attack behavior features. Based on the attack behavior feature library, the similarity and potential homologous relationship between different preset Internet protocol addresses can be analyzed and identified. The content of the attack behavior feature library has been introduced above and will not be elaborated here.

[0139] Among them, the attack feature relationship can describe the similarity or relevance shown between two preset Internet protocol addresses based on specific attack behavior features. For example, by querying the attack behavior feature library, it can be determined that the password features and domain names of two preset Internet protocol addresses overlap, etc. Then, one of the two preset Internet protocol addresses can be divided into the address cluster where the other is located.

[0140] Among them, the associated Internet protocol address can be another preset Internet protocol address in other address clusters (such as the initial address cluster or the intermediate address cluster) that has an attack feature relationship with the preset Internet protocol address to be divided.

[0141] In some embodiments, the second dynamic threshold can be set according to the actual situation.

[0142] Exemplarily, when the modularity corresponding to the initial address cluster is less than the preset second dynamic threshold, for example, when the modularity of the initial address cluster is -0.1 and the second dynamic threshold is 0.1, then it is determined that the initial address cluster is an intermediate address cluster for re - division. For example, according to the association rules in the attack behavior feature library (such as "homologous attack features" including sharing the same C2 server, consistent Payload hash, etc.), the association strength between each preset Internet protocol address in the intermediate address cluster and other address clusters (which can include the intermediate address cluster and the initial address cluster) can be re - evaluated. If it meets the rules of the attack behavior feature library, for example, IP1 in intermediate address cluster A and IP2 in intermediate address cluster B belong to the same subnet segment and the same autonomous system, then IP1 can be tried to be migrated to intermediate address cluster B, and the modularity of intermediate address cluster A and intermediate address cluster B can be recalculated.

[0143] Furthermore, if after a preset Internet protocol address is migrated to a new address cluster, the modularity of the address cluster where it is located decreases, then the address cluster to which the preset Internet protocol address belongs can be re - determined.

[0144] In some embodiments, for each preset Internet Protocol (IP) address in each initial address cluster, modularity can also be calculated after iterative movement among multiple address clusters. If it contributes to the modularity of the migrated address cluster, for example, after migrating IP1 from address cluster A to address cluster B, both the modularity of address cluster A and address cluster B are improved, then it can be determined that the target address cluster corresponding to IP1 is address cluster B; otherwise, continue the migration attempt for IP1.

[0145] Furthermore, it is also possible to stop the iteration for each address cluster and obtain multiple target address clusters only after the modularity of all address clusters is greater than the second dynamic threshold. If there is a new preset IP address that does not contribute to the modularity of any address cluster (or reduces the modularity of the corresponding address cluster after joining any address cluster), then this new preset IP address can be used as a separate target address cluster.

[0146] It should be noted that for other preset IP addresses, the above method can also be used to determine the target address cluster corresponding to each preset IP address, so as to improve the accuracy of target address cluster division, and further improve the efficiency, accuracy, and interpretability of network attack homology analysis.

[0147] In some embodiments, when a new preset IP address is added, there is no need to recalculate the entire graph. Instead, an incremental algorithm can be used to recalculate the local address clusters to avoid reconstructing the entire graph. For example, when there is a new preset IP address (such as IPx), its associated IP addresses with attack feature relationships can be determined through the attack behavior feature library, such as IP1 and IP2. If IP1 corresponds to target address cluster A and IP2 corresponds to target address cluster B, and after IPx joins target address cluster A, the modularity of target address cluster A increases by 0.1, and after IPx joins target address cluster B, the modularity of target address cluster B increases by 0.02. Then, by comparing the modularity gains of target address cluster A and target address cluster B, it can be determined that target address cluster A with a larger modularity gain is the address cluster to which IPx belongs. At this time, IPx can be added to target address cluster A. In this way, only the affected address clusters need to be re-divided, which improves the calculation efficiency while ensuring the accuracy of the division.

[0148] In some embodiments, a sliding window mechanism can be set to recalculate the address cluster division regularly (such as daily or weekly) using the latest attack data of such preset IP addresses to ensure the real-time performance and accuracy of the structure of each target address cluster.

[0149] By continuously optimizing each address cluster through an iterative process, it can be ensured that the preset Internet Protocol addresses within each resulting target address cluster have a high degree of homology and a strong association of attack characteristics. Thereby, not only is the ability to understand and identify network attack behaviors enhanced, but it also helps to build a more accurate and effective network security defense system.

[0150] In some embodiments, to facilitate the rapid traceability of the Internet Protocol addresses to be analyzed, the characteristic centroid of each target address cluster can be calculated to quantify and characterize the main attack behavior characteristics of each target address cluster, so as to facilitate the subsequent rapid comparison of the Internet Protocol addresses to be analyzed with the characteristic centroid, and thus achieve the rapid and accurate traceability of the Internet Protocol addresses to be analyzed. Exemplarily, "determining the characteristic centroid of each target address cluster" in step 104 may include: (104.b1) For each target address cluster, obtain the second feature vector of the target node corresponding to each preset Internet Protocol address in the target address cluster, and the node degree of the target node; (104.b2) Obtain a third product according to the product of the second feature vector and the corresponding node degree; (104.b3) Obtain a feature sum according to the sum of the multiple third products corresponding to the multiple preset Internet Protocol addresses included in each target address cluster; (104.b4) Obtain the sum of the multiple node degrees of the multiple target nodes corresponding to each target address cluster to obtain a target degree sum; (104.b5) Based on the ratio of the feature sum to the target degree sum, obtain the characteristic centroid of each target address cluster.

[0151] Among them, the second feature vector may be the feature vector possessed by the target node corresponding to each preset Internet Protocol address, which may include multi-dimensional features such as attack time interval, protocol distribution, payload entropy value, etc., and is used to describe the behavior pattern and attributes of the preset Internet Protocol address.

[0152] Among them, the node degree may represent the connection strength between each preset Internet Protocol address and other target nodes (i.e., other preset Internet Protocol addresses) in the target address cluster, and it can be obtained by summing the connection weights between the corresponding preset Internet Protocol address and all other homologous preset Internet Protocol addresses. The method for obtaining the connection weight has been introduced in detail above and will not be elaborated here.

[0153] Among them, the third product may be the result obtained according to the product of the second feature vector of the target node and its corresponding node degree, which reflects the weighted contribution of the feature vector of the target node in its target address cluster.

[0154] Among them, the sum of features can be the sum of multiple third products corresponding to all preset Internet protocol addresses included in each target address cluster, which synthesizes the weighted feature vectors of all target nodes within the target address cluster and is used to calculate the overall feature performance of the target address cluster.

[0155] Among them, the sum of target degrees can be the sum of the node degrees corresponding to all target nodes in each target address cluster, which is used to normalize the sum of features.

[0156] In some embodiments, the feature centroid of each target address cluster can be calculated by the following formula: ; Among them, is the feature centroid of the j-th target address cluster, is the node degree of the i-th target node in the target address cluster, is the node feature vector of the i-th target node, represents the third product, represents the sum of features, represents the sum of target degrees.

[0157] Specifically, the central features of each target address cluster can be described by calculating the feature centroid of the target address cluster, so as to quickly classify new Internet protocol addresses to be analyzed. Specifically, taking the calculation of the feature centroid in target address cluster A as an example, for each target node in target address cluster A, obtain the second feature vector and node degree of the target node. Taking target node 1 as an example, its second feature vector can be [0.2, 0.5, 0.8], and the node degree is 3.

[0158] Furthermore, through the product of the feature vector (i.e., the second feature vector) of each target node and its node degree, the corresponding third product can be obtained. Still taking target node 1 as an example, its third product is [0.2×3, 0.5×3, 0.8×3], that is, [0.6, 1.5, 2.4].

[0159] Furthermore, the sum of multiple third products corresponding to all target nodes included in the target address cluster can be calculated to obtain the sum of features of the target address cluster. Taking target address cluster A as an example, if it has 3 target nodes, and the third products of each target node are [0.6, 1.5, 2.4], [0.3, 0.9, 1.2], [0.4, 1.2, 1.6] in sequence, then the multiple third products can be added to obtain the corresponding sum of features as [1.3, 3.6, 5.2].

[0160] In some embodiments, the degrees of multiple target nodes corresponding to all target nodes in the target address cluster can be added together to obtain the total target degree. Taking the target address cluster A as an example, if it has 3 target nodes and the degrees of each target node are 3, 4, and 5 respectively, then the total target degree corresponding to the target address cluster A is 12.

[0161] Further, the characteristic centroid of each target address cluster can be obtained through the ratio of the total characteristic to the total target degree. For example, if the total characteristic of the target address cluster A is [1.3, 3.6, 5.2] and the total target degree is 12, then the characteristic centroid is [0.1083, 0.3, 0.4333].

[0162] Through the above method, the characteristic centroid of each target address cluster can be obtained, so as to transform the complex network behavior characteristics into cluster fingerprints that can be efficiently compared, providing core support for large-scale attack homology analysis.

[0163] Step 105: Obtain the Internet Protocol address to be analyzed, and calculate the similarity between the target feature vector of the Internet Protocol address to be analyzed and the multiple characteristic centroids of multiple target address clusters, so as to obtain the network attack homology analysis result of the Internet Protocol address to be analyzed.

[0164] In some embodiments, in order to quickly identify newly emerging or unknown attack sources, the similarity between the target feature vector of the Internet Protocol address to be analyzed and the characteristic centroids of multiple determined target address clusters can be calculated to determine whether the Internet Protocol address to be analyzed belongs to a certain known attack group, and thus obtain its network attack homology analysis result. In this way, the response speed and accuracy to potential threats can be effectively improved, and the overall network security protection ability can be enhanced.

[0165] Among them, the Internet Protocol address to be analyzed can be a newly discovered or unclassified IP address that needs to perform homology analysis, and the attack group to which it belongs is unknown, and network attack homology analysis needs to be performed to determine it.

[0166] Among them, the target feature vector can be a set of features possessed by each Internet Protocol address to be analyzed.

[0167] Among them, the network attack homology analysis result can be the result obtained through the similarity calculation between the target feature vector of the Internet Protocol address to be analyzed and the characteristic centroids of each target address cluster, which characterizes which known attack group the Internet Protocol address to be analyzed is most likely to belong to, or indicates that it does not belong to any known group, thereby providing a basis for subsequent security policy formulation.

[0168] In some embodiments, after obtaining the Internet Protocol address to be analyzed, its target feature vector can be extracted to obtain multiple eigenvalues. Next, this target feature vector is compared one by one with the multiple feature centroids of multiple target address clusters pre-calculated. Specifically, the comparison can be performed by calculating the similarity between the target feature vector and each feature centroid.

[0169] Furthermore, the similarity between each feature in the target feature vector and the feature in the corresponding feature centroid can be measured. The features for which the similarity is measured should belong to the same dimension, such as all belonging to the payload entropy value, etc.

[0170] Specifically, methods such as cosine similarity and Jaccard similarity can be used for similarity measurement. When calculating the similarity between the target feature vector and each feature centroid, the similarity between each feature in the target feature vector and the sub-feature centroid in the feature centroid of the corresponding dimension can be calculated. After obtaining the corresponding feature similarity, the calculated result is multiplied by the importance score of the feature centroid corresponding to this feature to obtain the target feature similarity corresponding to each target feature.

[0171] Exemplarily, similar to calculating the importance score of the first feature vector, the importance score of the feature centroid can also be calculated through an initial gradient boosting model. Specifically, the loss value is calculated by comparing the predicted importance score with the actual importance score to train the initial gradient boosting model. After the model training is completed, the gains corresponding to each feature dimension can be accumulated to obtain the importance score corresponding to the feature of each feature dimension; alternatively, the importance score of the feature centroid can also be set manually, and the embodiments of the present application do not make specific limitations on this.

[0172] Furthermore, the target feature similarities corresponding to the multiple target features included in the target feature vector can be added together to obtain the total feature similarity corresponding to each feature centroid and the Internet Protocol address to be analyzed.

[0173] Furthermore, after calculating multiple total feature similarities, each total feature similarity can be compared with a preset third dynamic threshold. When there is a total feature similarity greater than the third dynamic threshold, the feature centroid corresponding to this total feature similarity can be used as the target feature centroid corresponding to the Internet Protocol address to be analyzed, and the target address cluster corresponding to this target feature centroid can be used as the target homologous cluster of the Internet Protocol address to be analyzed, thereby obtaining the network attack homologous analysis result of the Internet Protocol address to be analyzed.

[0174] In some embodiments, when there are multiple overall feature similarities greater than the third dynamic threshold, the target address cluster corresponding to the feature centroid with the maximum overall feature similarity can be determined as the target homologous cluster of the Internet Protocol address to be analyzed. For example, if the target overall feature similarity between the feature centroid 1 of the target address cluster A and the Internet Protocol address to be analyzed is 0.9, the target overall feature similarity between the feature centroid 2 of the target address cluster B and the Internet Protocol address to be analyzed is 0.75, and the third dynamic threshold is 0.7, then the target address cluster A can be determined as the target homologous cluster of the Internet Protocol address to be analyzed.

[0175] This application obtains multiple preset Internet Protocol (IP) addresses and a preset weighted attack behavior graph. The preset weighted attack behavior graph includes multiple preset IP addresses and multiple connection weights, where each connection weight represents the degree of association between two preset IP addresses with a homologous relationship; divides the multiple preset IP addresses in the preset weighted attack behavior graph into multiple initial address clusters; for each initial address cluster, determines the connection weights associated with each preset IP address included, and calculates the corresponding modularity based on the connection weights. The modularity is used to characterize the homologous compatibility degree among the multiple preset IP addresses included in the corresponding initial address cluster; re-determines the attribution relationship between each preset IP address and the initial address cluster according to the modularity corresponding to each initial address cluster, obtains multiple target address clusters, and determines the characteristic centroid of each target address cluster; obtains the IP address to be analyzed, and calculates the similarity between the target feature vector of the IP address to be analyzed and the multiple characteristic centroids of the multiple target address clusters to obtain the network attack homologous analysis result of the IP address to be analyzed. In this way, the connection weights between multiple IP addresses in the preset weighted attack behavior graph pre-generated can accurately quantify the relevance of the attack behavior between any two IP addresses, so as to improve the interpretability of the relationship between each IP address, and further improve the accuracy of homologous analysis; moreover, by adopting the method of modularity iterative processing, the preset weighted attack behavior graph is divided into target address clusters, and the characteristic centroid of each target address cluster is calculated. In this way, a globally representative core feature vector can be accurately generated based on the homologous IP addresses included in the same address cluster, so that when there is an IP address to be analyzed that needs to perform homologous analysis subsequently, the IP address to be analyzed can be directly compared with the characteristic centroids of each target address cluster (including a large number of homologous IP addresses), without the need to compare and cluster with each of the large number of IP addresses one by one, and homologous determination is performed based on the centroid similarity, avoiding the low efficiency problem caused by pairwise comparison of the IP address to be analyzed with a large number of IP addresses. In this way, both the traceability of the attack pattern and the accuracy of homologous analysis are retained, and the efficiency of homologous analysis is improved. In summary, this application can improve the efficiency and accuracy of network attack homologous analysis, which is of great significance for enhancing network security protection capabilities.

[0176] In some embodiments, in order to improve the efficiency and accuracy of homology analysis, the homology cluster to which the Internet Protocol address to be analyzed belongs can be determined by calculating the similarity between the feature vector of the Internet Protocol address to be analyzed and each known feature centroid, so as to improve the interpretability and analysis efficiency of homology analysis. For example, the "calculating the similarity between the target feature vector of the Internet Protocol address to be analyzed and the multiple feature centroids of multiple target address clusters to obtain the network attack homology analysis result of the Internet Protocol address to be analyzed" in step 105 may include: (105.1) calculating the similarity between each target feature in the target feature vector corresponding to the Internet Protocol address to be analyzed and each sub-feature centroid in each feature centroid to obtain the corresponding feature similarity; (105.2) Obtain the feature centroid importance score corresponding to each sub-feature centroid, and obtain the target feature similarity corresponding to each target feature based on the product of the feature similarity and the feature centroid importance score; (105.3) Adding the target feature similarities corresponding to the target features to obtain the total feature similarity corresponding to the centroid of each feature; (105.4) obtaining a preset third dynamic threshold, and sequentially comparing a plurality of total feature similarities corresponding to a plurality of feature centroids with the third dynamic threshold to obtain a comparison result; (105.5) When the comparison result indicates that there is a target total feature similarity among multiple total feature similarities that is greater than a third dynamic threshold, the corresponding target address cluster is determined as the target homology cluster of the Internet Protocol address to be analyzed, and based on the target homology cluster, the network attack homology analysis result of the Internet Protocol address to be analyzed is obtained.

[0177] Among them, the target features can be the specific features in the target feature vector of the Internet Protocol address to be analyzed, such as the feature values ​​corresponding to feature dimensions such as attack time interval, protocol distribution, and load entropy value. Multiple target features can form the target feature vector corresponding to the Internet Protocol address to be analyzed.

[0178] The sub-feature centroid may be a specific feature contained in each feature centroid, which is used to characterize the feature value of the target homologous cluster in a specific dimension.

[0179] The feature similarity may be the similarity between each target feature of the Internet Protocol address to be analyzed and the corresponding sub-feature centroid, which may be calculated by methods such as cosine similarity or Jaccard similarity.

[0180] The feature centroid importance score may be a value reflecting the relative importance of each sub-feature centroid in the feature centroid to which it belongs, which may be calculated through historical data statistics, technical personnel evaluation, or an initial gradient boosting model.

[0181] Among them, the target feature similarity can be the result obtained by multiplying the feature similarity by the feature centroid importance score, representing the weighted similarity between a certain target feature of the Internet Protocol address to be analyzed and a certain sub-feature centroid.

[0182] Among them, the total feature similarity can be the value obtained by adding up the target feature similarities corresponding to multiple target features, representing the overall similarity between the Internet Protocol address to be analyzed and the corresponding feature centroid.

[0183] Among them, the third dynamic threshold can be the threshold used to determine whether the total feature similarity reaches the attribution standard. If there is a total feature similarity greater than this threshold, it is considered that the Internet Protocol address to be analyzed has a significant similarity with the corresponding feature centroid.

[0184] Among them, the comparison result can be the result obtained by comparing multiple total feature similarities with the second dynamic threshold, used to determine whether there is a situation where the target total feature similarity is greater than the second dynamic threshold.

[0185] Among them, the target homologous cluster can be when the comparison result indicates that there is a target total feature similarity greater than the second dynamic threshold, the target address cluster to which the corresponding feature centroid belongs is the target homologous cluster, that is, it indicates that the probability that the Internet Protocol address to be analyzed belongs to this target homologous cluster is the largest.

[0186] In some embodiments, the total feature similarity between the Internet Protocol address to be analyzed and the corresponding feature centroid can be calculated by the following formula: ; Among them, represents the Internet Protocol address to be analyzed, represents the feature centroid of the j-th target address cluster, represents the feature centroid importance score, represents the k-th target feature, represents the k-th sub-feature centroid, represents the feature similarity.

[0187] Exemplarily, if the target feature vector of the Internet Protocol address to be analyzed is [0.8, 1, 0.7], and the feature centroid 1 of the target address cluster A is [0.9, 0, 0.6], by calculating the similarity between each target feature in the target vector and the corresponding sub-feature centroid, the corresponding feature similarity can be obtained. Taking the feature of the first dimension as an example, the feature similarity between the target feature 0.8 and the sub-feature centroid 0.9 can be calculated. For example, if the cosine similarity is used for calculation, then the feature similarity can be obtained as 0.95. Thus, the feature similarity between the target feature and the sub-feature centroid under each feature dimension can be calculated.

[0188] In some embodiments, the calculation methods of the feature similarities under each feature dimension can be the same or different, and can be specifically set according to the actual situation.

[0189] Further, if in the feature centroid 1, the weight w1 of the first dimension is 0.5, the weight w2 of the second dimension is 0.3, and the weight w3 of the third dimension is 0.2, then, if through the above calculation method, the feature similarities of the three dimensions are 0.95, 0, and 0.9 respectively. Thus, through the product of each feature similarity and the corresponding feature centroid importance score, the target feature similarity corresponding to each target feature can be calculated. For example, the target feature similarity of the first dimension is 0.95×0.5 = 0.475.

[0190] Further, the multiple target feature similarities corresponding to the target feature vector can be added together to obtain the total feature similarity corresponding to the corresponding feature centroid and the Internet Protocol address to be analyzed: the first dimension is 0.95×0.5 = 0.475, the second dimension is 0×0.3 = 0, and the third dimension is 0.9×0.2 = 0.18. Adding the target feature similarities corresponding to the 3 dimensions, that is, 0.475 + 0 + 0.18, the final total feature similarity is obtained as 0.655.

[0191] Exemplarily, if the preset third dynamic threshold is 0.6, then the total feature similarity can be compared with the third dynamic threshold. For example, comparing 0.655 with 0.6, 0.655 is greater than 0.6, indicating that the Internet Protocol address to be analyzed is homologous to the target address cluster A corresponding to the feature centroid 1. At the same time, if the total feature similarity between the Internet Protocol address to be analyzed and other target address clusters is lower than the third dynamic threshold, it indicates that the homologous relationship between the Internet Protocol address to be analyzed and these target address clusters is non-homologous.

[0192] Through the above method, a comprehensive search of the entire database can be avoided, effectively simplifying the homologous analysis process, realizing efficient, accurate, and interpretable homologous attack analysis, and providing key technical support for real-time threat response.

[0193] Please refer to Figure 3 , Figure 3 which is the overall flowchart of the network attack homology analysis method provided by the embodiments of this application. Exemplarily, multiple preset Internet protocol addresses can be extracted from the attack logs, and the preset Internet protocol addresses can be a large number of IP addresses covering various attack groups. Then, according to the multiple preset Internet protocol addresses extracted, an attack behavior feature library can be constructed to facilitate subsequent operations such as the division of address clusters, etc.

[0194] Specifically, the attack behavior feature library can include spatio-temporal features, behavior fingerprints, and association graphs. Exemplarily, the spatio-temporal features can depict the attacker's behavior patterns from three dimensions: the network layer, geographical features, and time patterns. The behavior fingerprints can depict the attacker's technical means and behavior habits from three dimensions: attack vectors, payload features, and password features. The association graphs can depict the association relationships among attackers from three dimensions: honeypot linkage, service association, and domain name mapping.

[0195] Furthermore, the features in the attack behavior feature library can be encoded to facilitate obtaining the first feature vector. The ways of encoding the features can include the original values of numerical features, One-Hot encoding (i.e., one-hot encoding) of categorical features, vectorization and cosine similarity of text features, binary values of boolean features, and so on. And based on the data encoded from any two preset Internet protocol addresses, the relationship between the two preset Internet protocol addresses is modeled, and the connection weight between the two preset Internet protocol addresses is determined. Finally, with the preset Internet protocol address as the target node, the homology relationship between two preset Internet protocol addresses as the connection edge, and the connection weight as the edge weight, a preset weighted attack behavior graph can be constructed.

[0196] Furthermore, multiple preset Internet protocol addresses can be divided into multiple initial address clusters, and for each initial address cluster, the modularity is calculated according to the connection weights between the preset Internet protocol addresses in the preset weighted attack behavior graph to evaluate the homology compatibility degree among the preset Internet protocol addresses within the address cluster. Then, according to the magnitudes of the modularity corresponding to each initial address cluster, the attribution relationship between each preset Internet protocol address and the initial address cluster is re-determined to obtain multiple target address clusters.

[0197] Furthermore, for each target address cluster, the feature centroid representing the average features of the entire address cluster can be calculated. In this way, when an Internet Protocol address to be analyzed for homologous analysis appears, the similarity between the target feature vector corresponding to the Internet Protocol address to be analyzed and the feature centroids of multiple target address clusters can be calculated respectively, so as to obtain the result of network attack homologous analysis, avoiding the low efficiency problem caused by pairwise comparison of the Internet Protocol address to be analyzed with a large number of Internet Protocol addresses. In this way, the traceability of the attack pattern is retained, the accuracy of homologous analysis is ensured, and the efficiency of homologous analysis is improved.

[0198] In some embodiments, after determining that the Internet Protocol address to be analyzed belongs to a certain target address cluster, the Internet Protocol address to be analyzed can be updated to the corresponding target address cluster, and the feature centroid of the target address cluster can be updated to ensure the real-time performance and accuracy of homologous determination.

[0199] Please refer to Figure 4 , the embodiment of the present application further provides a network attack homologous analysis device, which can implement the above network attack homologous analysis method. The network attack homologous analysis device includes: An acquisition module 41, configured to acquire a plurality of preset Internet Protocol addresses and a preset weighted attack behavior graph, where the preset weighted attack behavior graph includes a plurality of preset Internet Protocol addresses and a plurality of connection weights, and each connection weight represents the association degree between two corresponding preset Internet Protocol addresses having a homologous relationship; A partitioning module 42, configured to partition the plurality of preset Internet Protocol addresses in the preset weighted attack behavior graph into a plurality of initial address clusters; A first calculation module 43, configured to, for each initial address cluster, determine the connection weights associated with each preset Internet Protocol address included, and calculate the corresponding modularity according to the connection weights, where the modularity is used to characterize the homologous compatibility degree among the plurality of preset Internet Protocol addresses included in the corresponding initial address cluster; A determination module 44, configured to re-determine the attribution relationship between each preset Internet Protocol address and the initial address cluster according to the modularity corresponding to each initial address cluster, obtain a plurality of target address clusters, and determine the feature centroid of each target address cluster; A second calculation module 45, configured to acquire the Internet Protocol address to be analyzed, and calculate the similarity between the target feature vector of the Internet Protocol address to be analyzed and the plurality of feature centroids of the plurality of target address clusters, so as to obtain the network attack homologous analysis result of the Internet Protocol address to be analyzed.

[0200] The specific implementation manner of the network attack homology analysis device is basically the same as the specific embodiments of the above-mentioned network attack homology analysis method, and will not be elaborated here. On the premise of meeting the requirements of the embodiments of the present application, other functional modules can also be set in the network attack homology analysis device to implement the network attack homology analysis method in the above embodiments.

[0201] An embodiment of the present application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned network attack homology analysis method is implemented. The computer device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0202] Please refer to Figure 5 , Figure 5 which schematically shows the hardware structure of a computer device in another embodiment. The computer device includes: A processor 51, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application; A memory 52, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 52 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 52, and the processor 51 is called to execute the network attack homology analysis method of the embodiments of the present application; An input / output interface 53, which is used to implement information input and output; A communication interface 54, which is used to implement communication interaction between this device and other devices, and can communicate through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); A bus 55, which transmits information between various components of the device (such as the processor 51, the memory 52, the input / output interface 53, and the communication interface 54); Among them, the processor 51, the memory 52, the input / output interface 53, and the communication interface 54 are communicatively connected to each other inside the device through the bus 55.

[0203] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned network attack homology analysis method is implemented.

[0204] As a non-transitory computer-readable storage medium, a memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0205] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0206] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine some steps, or different steps.

[0207] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0208] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0209] In the description of the present application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0210] It should be understood that in the present application, "at least one (item)" and "several" mean one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one)" or similar expressions below are any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0211] In several embodiments provided by the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be an indirect coupling or communication connection through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms.

[0212] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0213] In addition, in each embodiment of the present application, each functional unit may be integrated into a processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0214] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0215] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A network attack homology analysis method, characterized in that: The method comprises: Acquire a plurality of preset Internet Protocol addresses and a preset weighted attack behavior graph, wherein the preset weighted attack behavior graph includes a plurality of preset Internet Protocol addresses and a plurality of connection weights, each connection weight representing a correlation between two preset Internet Protocol addresses having a homology relationship; Dividing a plurality of preset Internet Protocol addresses in the preset weighted attack behavior graph into a plurality of initial address clusters; For each initial address cluster, determining a connection weight associated with each preset Internet Protocol address contained therein, and calculating a corresponding modularity according to the connection weight, wherein the modularity is used to characterize the homology compatibility between the multiple preset Internet Protocol addresses contained in the corresponding initial address cluster; Re-determining the attribution relationship between each preset Internet Protocol address and the initial address cluster according to the modularity corresponding to each initial address cluster, obtaining multiple target address clusters, and determining the characteristic centroid of each target address cluster; An Internet Protocol address to be analyzed is obtained, and similarity calculation is performed between a target feature vector of the Internet Protocol address to be analyzed and multiple feature centroids of multiple target address clusters to obtain a network attack homology analysis result of the Internet Protocol address to be analyzed.

2. The network attack homology analysis method according to claim 1, characterized in that: The step of obtaining a plurality of preset Internet Protocol addresses and a preset weighted attack behavior graph includes: Acquire the plurality of preset Internet Protocol addresses and homology relationships between the plurality of preset Internet Protocol addresses, and determine a first feature vector and a relationship label between any two preset Internet Protocol addresses, wherein the first feature vector includes a plurality of first feature values ​​between the any two preset Internet Protocol addresses in a plurality of feature interaction dimensions; Inputting each first eigenvector and the corresponding relationship label into the initial gradient boosting model in turn, and performing decision tree split gain accumulation through the initial gradient boosting model to obtain a feature importance score corresponding to each first eigenvalue; Determining, based on the plurality of feature importance scores, connection weights of two preset Internet Protocol addresses corresponding to the first feature vector; Based on the multiple preset Internet Protocol addresses, the relationship labels between any two preset Internet Protocol addresses, and the corresponding connection weights, the preset weighted attack behavior graph corresponding to the multiple preset Internet Protocol addresses is constructed.

3. The network attack homology analysis method according to claim 2 is characterized in that: The step of sequentially inputting each first eigenvector and the corresponding relationship label into the initial gradient boosting model, and performing decision tree split gain accumulation through the initial gradient boosting model to obtain a feature importance score corresponding to each first eigenvalue includes: Inputting each first eigenvector and the corresponding relationship label into the initial gradient boosting model in turn, and predicting each first eigenvector through the initial gradient boosting model to obtain the corresponding prediction label; Determining a first target loss based on a difference between the predicted label and the relationship label; Based on the first target loss, splitting gains are calculated for the multiple first eigenvalues ​​included in each of the first eigenvectors, and the intermediate splitting feature with the largest splitting gain is determined from the multiple first eigenvalues ​​to perform node splitting, and the second target loss is calculated; Based on the difference between the second target loss and the first target loss, splitting gains are calculated for multiple first eigenvalues ​​included in the first eigenvector, and an intermediate splitting feature with the largest splitting gain is determined from the multiple first eigenvalues ​​to perform node splitting, and an updated second target loss is calculated; Repeat the step of calculating the splitting gain of multiple first eigenvalues ​​contained in the first eigenvector based on the difference between the updated second target loss and the first target loss, and determining the intermediate splitting feature with the largest splitting gain from the multiple first eigenvalues ​​for node splitting, and calculating the updated second target loss until a preset number of training times is reached, and obtaining the feature importance score corresponding to each first eigenvalue according to the multiple splitting gains iterated for each eigenvalue.

4. The network attack homology analysis method according to claim 2, characterized in that: The determining, based on the plurality of feature importance scores, connection weights of two preset Internet Protocol addresses corresponding to the first feature vector includes: Normalize each feature importance score to obtain the corresponding target feature importance score; Obtaining a sub-feature weight of each first eigenvalue according to the product of each first eigenvalue and the corresponding target feature importance score; Based on the sum of the multiple sub-feature weights corresponding to the multiple first feature values, the connection weights of the two preset Internet Protocol addresses corresponding to the first feature vector are obtained.

5. The network attack homology analysis method according to claim 2, characterized in that: Before constructing the preset weighted attack behavior graph corresponding to the plurality of preset Internet Protocol addresses based on the plurality of preset Internet Protocol addresses, the relationship labels between any two preset Internet Protocol addresses, and the corresponding connection weights, the method further includes: Obtaining multiple historical connection weights between any two preset Internet Protocol addresses, and score means and score standard deviations of the multiple historical connection weights; Obtaining a preset adjustment parameter, and obtaining a first product according to the product of the adjustment parameter and the scoring standard deviation; Determining a first dynamic threshold value according to a difference between the score mean and the first product; Comparing the connection weight between the arbitrary two preset Internet Protocol addresses with the first dynamic threshold to obtain a comparison result; Based on the comparison result, updating the relationship label between the arbitrary two preset Internet Protocol addresses to obtain a target relationship label; Then, based on the multiple preset Internet Protocol addresses, the relationship labels between any two preset Internet Protocol addresses, and the corresponding connection weights, constructing the preset weighted attack behavior graph corresponding to the multiple preset Internet Protocol addresses includes: Based on the multiple preset Internet Protocol addresses, the target relationship labels between any two preset Internet Protocol addresses, and the corresponding connection weights, the preset weighted attack behavior graph corresponding to the multiple preset Internet Protocol addresses is constructed.

6. The network attack homology analysis method according to claim 1, characterized in that: The calculating the corresponding modularity according to the connection weights includes: Obtaining a graph connection weight according to the sum of connection weights corresponding to all connection edges in the preset weighted attack behavior graph; Obtaining the node degree corresponding to the target node corresponding to each preset Internet Protocol address in the initial address cluster; According to the product of the node degrees between any two target nodes, the second product is obtained; Obtaining a first ratio based on a ratio of the second product to the graph connection weight; Obtain the difference between the connection weights corresponding to any two target nodes in the initial address cluster and the corresponding first ratio to obtain a first difference; Based on the graph connection weights and a plurality of first differences between a plurality of target nodes included in the initial address cluster, a modularity corresponding to the initial address cluster is obtained.

7. The network attack homology analysis method according to claim 1, characterized in that: The method of re-determining the attribution relationship between each preset Internet Protocol address and the initial address cluster according to the modularity corresponding to each initial address cluster to obtain multiple target address clusters includes: According to the modularity corresponding to each initial address cluster, determining an intermediate address cluster whose modularity is less than a preset second dynamic threshold; Obtaining a preset attack behavior feature library, and based on the attack behavior feature library, re-determining an associated Internet Protocol address having an attack feature relationship with each preset Internet Protocol address in each intermediate address cluster, and migrating each preset Internet Protocol address to the initial address cluster corresponding to the associated Internet Protocol address; Repeat the steps of re-determining the associated Internet Protocol address having the attack feature relationship with each preset Internet Protocol address in each intermediate address cluster based on the attack behavior feature library, and migrating each preset Internet Protocol address to the initial address cluster corresponding to the associated Internet Protocol address, until the modularity corresponding to each intermediate address cluster is greater than the second dynamic threshold, thereby obtaining multiple target address clusters.

8. The network attack homology analysis method according to claim 1, characterized in that: Determining the characteristic centroid of each target address cluster includes: For each target address cluster, obtaining a second feature vector of a target node corresponding to each preset Internet Protocol address in the target address cluster and a node degree of the target node; Obtaining a third product according to the product of the second eigenvector and the corresponding node degree; Obtaining a characteristic sum according to the sum of multiple third products corresponding to the multiple preset Internet Protocol addresses included in each target address cluster; Obtaining the sum of multiple node degrees of multiple target nodes corresponding to each target address cluster to obtain a target degree sum; Based on the ratio of the feature sum to the target degree sum, the feature centroid of each target address cluster is obtained.

9. The network attack homology analysis method according to claim 1, characterized in that: The similarity calculation of the target feature vector of the Internet Protocol address to be analyzed and the multiple feature centroids of the multiple target address clusters to obtain the network attack homology analysis result of the Internet Protocol address to be analyzed includes: For each target feature in the target feature vector corresponding to the Internet Protocol address to be analyzed, similarity calculation is performed with each sub-feature centroid in each feature centroid to obtain corresponding feature similarity; Obtaining a feature centroid importance score corresponding to each sub-feature centroid, and obtaining a target feature similarity corresponding to each target feature based on the product of the feature similarity and the feature centroid importance score; Adding multiple target feature similarities corresponding to multiple target features to obtain a total feature similarity corresponding to each feature centroid; Obtaining a preset third dynamic threshold, and sequentially comparing a plurality of total feature similarities corresponding to the plurality of feature centroids with the third dynamic threshold to obtain a comparison result; When the comparison result indicates that there is a target total feature similarity among the multiple total feature similarities that is greater than the third dynamic threshold, the corresponding target address cluster is determined as the target homology cluster of the Internet Protocol address to be analyzed, and based on the target homology cluster, the network attack homology analysis result of the Internet Protocol address to be analyzed is obtained.

10. A network attack homology analysis device, characterized in that: The device comprises: An acquisition module, used to acquire a plurality of preset Internet Protocol addresses and a preset weighted attack behavior graph, wherein the preset weighted attack behavior graph includes a plurality of preset Internet Protocol addresses and a plurality of connection weights, each connection weight representing a correlation between two preset Internet Protocol addresses having a homology relationship; A division module, used for dividing a plurality of preset Internet Protocol addresses in the preset weighted attack behavior graph into a plurality of initial address clusters; A first calculation module is used to determine, for each initial address cluster, a connection weight associated with each preset Internet Protocol address contained therein, and calculate a corresponding modularity according to the connection weight, wherein the modularity is used to characterize the homology compatibility between the multiple preset Internet Protocol addresses contained in the corresponding initial address cluster; A determination module, configured to re-determine the attribution relationship between each preset Internet Protocol address and the initial address cluster according to the modularity corresponding to each initial address cluster, obtain multiple target address clusters, and determine the characteristic centroid of each target address cluster; The second calculation module is used to obtain the Internet Protocol address to be analyzed, and perform similarity calculation on the target feature vector of the Internet Protocol address to be analyzed and multiple feature centroids of multiple target address clusters to obtain the network attack homology analysis result of the Internet Protocol address to be analyzed.

11. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the network attack homology analysis method according to any one of claims 1 to 9 when executing the computer program.

12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the network attack homology analysis method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • User risk assessment method and device

    CN117009879A

  • Human-cluster interaction method and system based on augmented reality

    CN117075725A

  • Intelligent management method and system for network attack blacklist

    CN119696906A

  • Graph-based analysis of security incidents

    US20230275912A1

Cited By

  • Homologous analysis method, apparatus and device for internet protocol address, and storage medium

    CN121792242A