Attack gang clustering method and device based on unsupervised clustering algorithm
The clustering of attack gangs through unsupervised clustering algorithms solves the problem of insufficient correlation between existing security products when identifying attack gangs, realizes in-depth analysis and effective defense of the attack ecosystem, and improves identification accuracy and defense efficiency.
Patent Information
- Application Number
- CN202510499909.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-25
AI Technical Summary
When facing organized and professional attack gangs, existing security products are difficult to deeply analyze the correlation and attack paths between attackers, resulting in the inability to effectively identify the true identity, organizational structure, attack motivation and tooling of the attack gang, unable to predict and prevent future attack behaviors, and lack an in-depth understanding of the patterns and trends of the entire attack ecosystem.
The attack gang clustering method based on the unsupervised clustering algorithm is adopted. By collecting multiple alarm information, field association and feature variable extraction are performed, graph feature matrix is constructed, and heterogeneous graph variational autoencoder and density space DBSCAN clustering algorithm is used to iteratively select and cluster unvisible data points to form an attack gang cluster.
It improves the accuracy of identification of attack gangs, explores the patterns and trends of potential attack gangs, reduces network security costs, improves defense capabilities and response speed, and achieves effective identification and defense of the attack ecosystem.
Smart Images

Figure CN120378156A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and particularly to a method and device for clustering attack groups based on an unsupervised clustering algorithm. Background Art
[0002] When existing security products are designed, they often focus on integrating and analyzing information from various network data sources, including but not limited to log files, network traffic records, intrusion detection system (IDS) alerts, and user behavior data, etc. Through their powerful data processing capabilities and complex algorithms, they can quickly identify individual attack events or abnormal behaviors, which is very effective for timely response and defense against isolated security threats.
[0003] However, when facing organized and professional attack groups, the limitations of these products gradually emerge. In order to achieve their goals, attack groups usually adopt highly concealed, variable, and complex attack strategies. They may use various technical means such as botnets, jump servers, and proxy servers to hide their true identities and locations, and even use completely different attack IP addresses at different attack stages to avoid detection and tracking by traditional security products.
[0004] When existing security products handle such attacks, they often only see individual isolated attack events and do not consider integrating and analyzing information from different data sources that seemingly have nothing to do with each other to discover the correlations and attack paths among the attackers hidden behind them. Therefore, existing security products are difficult to accurately identify the true identities, organizational structures, attack motives, and the tools and techniques used by attack groups, and thus cannot effectively predict and prevent future attack behaviors. At the same time, due to the lack of in-depth understanding of the patterns and trends of the entire attack ecosystem, when formulating defense strategies, they often have to rely on experience and intuition, which undoubtedly increases the difficulty and uncertainty of the defense work. Summary of the Invention
[0005] The present invention provides a method and device for clustering attack groups based on an unsupervised clustering algorithm to solve problems such as that existing security products do not deeply analyze the correlations between these data and attackers, resulting in insufficient insights in the dimension of attack groups and being unable to effectively identify the patterns and trends of the entire attack ecosystem.
[0006] An embodiment of the first aspect of the present invention provides a clustering method for attack groups based on an unsupervised clustering algorithm, including the following steps: collecting a plurality of alarm messages, and associating fields of the plurality of alarm messages to obtain an alarm log; extracting a plurality of feature variables from the alarm log, and constructing a graph feature matrix according to the plurality of feature variables; selecting any unvisited starting data point in the graph feature matrix as the first point in a new spatial clustering cluster, and clustering neighboring data points whose distance from the first point is less than or equal to a distance metric threshold into the current spatial clustering cluster, iteratively executing the selection and clustering process until all data points in the graph feature matrix are visited, to obtain a plurality of attack group clusters.
[0007] Optionally, the collecting a plurality of alarm messages, and associating fields of the plurality of alarm messages to obtain an alarm log includes:
[0008] Collecting the plurality of alarm messages based on honeypots and security devices;
[0009] Associating fields of the plurality of alarm messages to obtain the alarm log, where the alarm log includes identity information of the attacker, information of the attacked target, attack means characteristics, and the moment when the attack occurs.
[0010] Optionally, the extracting a plurality of feature variables from the alarm log, and constructing a graph feature matrix according to the plurality of feature vectors includes:
[0011] Extracting a plurality of feature variables from the alarm log;
[0012] Splitting each feature variable into nodes formed by concatenating source IP address and destination IP address strings, and forming an initial graph feature matrix with the feature vectors of each node;
[0013] Using a heterogeneous graph variational autoencoder to embed the initial graph feature matrix to obtain new feature vectors of each node and edge after learning;
[0014] Forming the graph feature matrix with the feature vectors of each node and the new feature vectors of each node and edge after learning.
[0015] Optionally, the selecting any unvisited starting data point in the graph feature matrix as the first point in a new spatial clustering cluster, and clustering neighboring data points whose distance from the first point is less than or equal to a distance metric threshold into the current spatial clustering cluster, iteratively executing the selection and clustering process until all data points in the graph feature matrix are visited, to obtain a plurality of attack group clusters includes:
[0016] Based on the density-based spatial clustering of applications with noise (DBSCAN) algorithm, select any unvisited starting data point in the graph feature matrix and determine whether there are neighboring data points of the current starting data point that are greater than or equal to the core point quantity threshold;
[0017] If there are neighboring data points that are greater than or equal to the core point quantity threshold, then take the current starting data point as the first point in the new spatial clustering cluster, and use the first point as the center of a circle, and cluster the neighboring data points whose distance from the first point is less than or equal to the distance metric threshold into the new spatial clustering cluster;
[0018] If there are no neighboring data points that are greater than or equal to the core point quantity threshold, then reselect any unvisited starting data point for access, and iteratively execute this process until all data points in the graph feature matrix have been visited, obtaining the multiple attack gang clusters.
[0019] An embodiment of the second aspect of the present invention provides an attack gang clustering device based on an unsupervised clustering algorithm, including: a collection module, configured to collect multiple alarm messages and perform field association on the multiple alarm messages to obtain an alarm log; a construction module, configured to extract multiple feature variables from the alarm log and construct a graph feature matrix according to the multiple feature variables; a clustering module, configured to select any unvisited starting data point in the graph feature matrix as the first point in the new spatial clustering cluster, and cluster the neighboring data points whose distance from the first point is less than or equal to the distance metric threshold into the current spatial clustering cluster, and iteratively execute the selection and clustering process until all data points in the graph feature matrix have been visited, obtaining multiple attack gang clusters.
[0020] Optionally, the collection module includes:
[0021] A collection unit, configured to collect the multiple alarm messages based on a honeypot and security devices;
[0022] An association unit, configured to perform field association on the multiple alarm messages to obtain the alarm log, where the alarm log includes the identity information of the attacker, the information of the attacked target, the characteristics of the attack means, and the moment when the attack occurred.
[0023] Optionally, the construction module includes:
[0024] An extraction unit, configured to extract multiple feature variables from the alarm log;
[0025] A first construction unit, configured to split each feature variable into nodes formed by splicing the source IP address and the destination IP address strings, and construct an initial graph feature matrix with the feature vectors of each node;
[0026] An embedding unit, configured to use a heterogeneous graph variational autoencoder to embed the initial graph feature matrix to obtain new feature vectors of each node and edge after learning;
[0027] A second construction unit, configured to construct the graph feature matrix from the feature vectors of each node and the new feature vectors of each node and edge after learning.
[0028] Optionally, the clustering module includes:
[0029] An access unit, configured to, based on the DBSCAN clustering algorithm, select any unvisited starting data point in the graph feature matrix for access, and determine whether there are neighboring data points of the current starting data point that are greater than or equal to the core point quantity threshold;
[0030] A clustering unit, configured to, if there are neighboring data points that are greater than or equal to the core point quantity threshold, use the current starting data point as the first point in a new spatial clustering cluster, and take the first point as the center of a circle, and cluster neighboring data points whose distance from the first point is less than or equal to the distance metric threshold into the new spatial clustering cluster;
[0031] A re-access unit, configured to, if there are no neighboring data points that are greater than or equal to the core point quantity threshold, re-select any unvisited starting data point for access, and iteratively execute the selection and clustering processes until all data points in the graph feature matrix have been accessed, to obtain the multiple attack gang clusters.
[0032] An embodiment of the third aspect of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the attack gang clustering method based on an unsupervised clustering algorithm as described in the above embodiment.
[0033] An embodiment of the fourth aspect of the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the above attack gang clustering method based on an unsupervised clustering algorithm.
[0034] The attack gang clustering method and device based on an unsupervised clustering algorithm proposed by the embodiments of the present invention collect alarm logs such as abnormal file access, network connection, and program execution with the help of honeypots, security devices, etc., extract alarm feature vectors, and use an unsupervised clustering method to analyze the feature data in each alarm log to classify attackers into gangs, so as to improve the accuracy of attack gang identification, discover potential attack gangs, effectively identify the patterns and trends of the entire attack ecosystem, enhance network security defense capabilities, reduce network security costs, and support rapid response and decision-making.
[0035] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, in which:
[0037] Figure 1 is a flowchart of a method for clustering attack groups based on an unsupervised clustering algorithm provided by an embodiment of the present invention;
[0038] Figure 2 is a schematic diagram of an IP address encoding algorithm provided by an embodiment of the present invention;
[0039] Figure 3 is a block diagram of a device for clustering attack groups based on an unsupervised clustering algorithm provided by an embodiment of the present invention;
[0040] Figure 4 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0042] The method and device for clustering attack groups based on an unsupervised clustering algorithm according to an embodiment of the present invention will be described below with reference to the accompanying drawings. Regarding the existing security products mentioned in the above background technology, they usually focus on integrating and analyzing network data sources. However, they have deficiencies in dealing with attack groups. Attack groups may use a variety of different attack IPs, but existing products usually do not consider in-depth analysis of the association between this data and attackers. This leads to insufficient insight in the dimension of attack groups and the problem of being unable to effectively identify the patterns and trends of the entire attack ecosystem. The present invention provides a method for clustering attack groups based on an unsupervised clustering algorithm. In this method, alarm feature vectors are extracted from a large number of company alarms, and the alarm information is clustered based on an unsupervised machine learning algorithm, so as to achieve the purpose of classifying individuals or groups, which helps the company better understand the intentions and methods of attackers.
[0043] Specifically, Figure 1 is a schematic flowchart of a method for clustering attack groups based on an unsupervised clustering algorithm provided by an embodiment of the present invention.
[0044] As Figure 1 shown, the clustering method for attack groups based on unsupervised clustering algorithms includes the following steps:
[0045] In step S101, collect multiple alarm messages and perform field association on the multiple alarm messages to obtain an alarm log.
[0046] In some embodiments, collecting multiple alarm messages and performing field association on the multiple alarm messages to obtain an alarm log includes:
[0047] Based on honeypots and security devices, collect multiple alarm messages;
[0048] Perform field association on the multiple alarm messages to obtain an alarm log, where the alarm log includes the identity information of the attacker, the information of the attacked target, the characteristics of the attack means, and the time when the attack occurred.
[0049] In the actual execution process, with the help of honeypots, security devices, etc., collect alarm messages such as abnormal file access, network connection, program execution, etc., and then, with the help of the existing intelligence database, associate the same fields to enrich the alarm content, and finally obtain the alarm log. Among them, the alarm log contains the identity information of the attacker (IP address, whether the IP address is in the blacklist, Cookie information, operating system version, MAC address), the information of the attacked target (IP address, port number), the characteristics of the attack means (vulnerability number exploited, service information attacked, scanning tool used), and the time when the attack occurred.
[0050] In step S102, extract multiple feature variables from the alarm log and construct a graph feature matrix according to the multiple feature variables.
[0051] In some embodiments, extracting multiple feature variables from the alarm log and constructing a graph feature matrix according to the multiple feature vectors includes:
[0052] Extract multiple feature variables from the alarm log;
[0053] Split each feature variable into nodes formed by concatenating the source IP address and the destination IP address strings, and form an initial graph feature matrix with the feature vectors of each node;
[0054] Use a heterogeneous graph variational autoencoder to embed the initial graph feature matrix to obtain new feature vectors of each node and edge after learning;
[0055] Form a graph feature matrix with the feature vectors of each node and the new feature vectors of each node and edge after learning.
[0056] In the actual execution process, after obtaining the various features of the alarm log, the alarm log needs to be encoded into a feature vector. It can be found that except for the moment when the attack occurs, which is a continuous variable, the other features are discrete variables. For discrete variables, if the range of possible values is limited and small, such as the protocol type, then the different possible values are numbered starting from 0, or one-hot encoding is used, etc., to form the first n - 1 dimensions of a feature vector of an alarm log. For the time information, the embodiment of the present invention uses a timestamp for encoding to form the nth dimension of a feature vector of an alarm log.
[0057] If the possible range of values of the discrete feature is very large, one-hot encoding may not be able to repeatedly learn and represent the feature. Therefore, the embodiment of the present invention proposes a method for feature embedding based on a heterogeneous graph variational autoencoder (H-VGAE) as follows.
[0058] Since the number of alarms may be huge, it is impossible to take each alarm as a node. Therefore, the embodiment of the present invention splits an alarm into a node v obtained by concatenating the source IP address and the destination IP address strings. Each node has a feature vector x, and all the feature vectors can form a feature matrix of the graph:
[0059]
[0060] As Figure 2 shown, the feature vector is composed of the embedding vectors of the source IP address (src_ip) and the destination IP address (dst_ip) concatenated:
[0061] x = [ip_encoding(src_ip), ip_encoding(dst_ip)]
[0062] Next, the co-occurrence edge is defined. If two alarms have the same certain type of discrete attribute, such as the same Cookie information, then there is a co-occurrence edge between the corresponding nodes of these two alarms. There may be multiple different types of co-occurrence edges between one node, such as Cookie information, MAC address, etc. The number of types is denoted as K, and the number of different values of the kth type is denoted as N k , and the feature vector on the kth co-occurrence edge is composed of the product of the One-Hot encoding and the learnable embedding matrix.
[0063] After completing the graph construction, the heterogeneous graph variational autoencoder (H-VGAE) is used to embed the graph to obtain the learned feature vectors of each node and edge. Finally, the feature of an original alarm is jointly composed of the feature vectors of its source node, destination node, and the embedding vectors of each type of edge.
[0064] It should be noted that the Heterogeneous Graph Variational Autoencoder (H-VGAE) proposed in the embodiments of the present invention is proposed based on the VGAE (Variational Graph Autoencoder), which is a graph data representation learning method integrating the characteristics of graph neural networks and variational autoencoders. This method uses variational inference to extract the low-dimensional representation of graph structure data, and its core lies in minimizing the KL (Kullback-Leibler) divergence between the variational distribution q(z) and the posterior distribution p(Z|X). Among them, Z represents the latent variable, X is the observed data, and the KL divergence can be calculated by the following formula:
[0065] KL(q(Z)||p(Z|X)) = E q(Z) [logq(Z) - logp(Z|X)]
[0066] Since it is difficult to directly minimize the KL divergence, VGAE indirectly achieves this goal by maximizing the Evidence Lower Bound (ELBO). The definition of ELBO is as follows:
[0067] ELBO(q) = E q ( z )[logp(X,Z) - logq(Z)]
[0068] Thus, the KL divergence can be decomposed into:
[0069] KL(q(Z)||p(Z|X)) = logp(X) - ELBO(q)
[0070] The VGAE model assumes that the latent variable Z follows a Gaussian distribution:
[0071]
[0072] Its encoder maps the feature vector of each node to the latent variables μ and δ through two layers of GCN networks 2 :
[0073] H = GCN1(X,A)
[0074] μ = GCN2(X,H)
[0075] δ = GCN3(X,H)
[0076] To better perform feature representation learning, the embodiment of the present invention replaces the original GCN network of VGAE with a GAT network. GAT (Graph Attention Network) uses an attention mechanism to specify the importance of different neighbor nodes to the central node. In traditional graph neural networks, the contributions of neighbor nodes to the central node are equal or based on predefined weights, such as the distance or connection strength between nodes; while GAT dynamically adjusts the importance of neighbor nodes through learned attention coefficients, enabling the model to better capture complex relationships in the graph structure.
[0077] The calculation process of GAT is defined as follows:
[0078]
[0079] where, α ij is the attention coefficient of node v i to node v j
[0080] To stabilize the learning process and improve performance, GAT adopts a multi-head attention mechanism, that is, multiple attention mechanisms are applied in parallel and then their results are combined. The calculation process is as follows:
[0081]
[0082] where, T is the number of attention heads, is the attention coefficient of the t-th head. Since the embodiment of the present invention has heterogeneous edge attributes, it is necessary to linearly transform the edge attributes during the message propagation process of the graph neural network and then concatenate them with the attributes of the nodes
[0083] The decoder of H-VGAE is responsible for reconstructing the graph structure from the latent space to maximize the probability of p(A|Z):
[0084]
[0085] where:
[0086]
[0087] z i is generated through the reparameterization process, enabling the gradient to be backpropagated from the decoder to the encoder:
[0088] z = μ + σ ⊙ ∈
[0089] where, ∈ is sampled from the standard normal distribution
[0090] The loss function of H-VGAE is defined as:
[0091]
[0092] Since the above feature vectors may still have a relatively high dimension, next, methods such as PCA (Principal Component Analysis) are used to transform the original data into a set of linearly independent representations in each dimension through linear transformation and other means without losing important information, so as to extract the main feature components of the alarm logs and obtain a low-dimensional feature representation.
[0093] In step S103, any unvisited starting data point in the graph feature matrix is selected as the first point in the new space clustering cluster, and the neighboring data points whose distance from the first point is less than or equal to the distance metric threshold are clustered into the current space clustering cluster. The selection and clustering process are iteratively executed until all data points in the graph feature matrix are visited, and multiple attack gang clusters are obtained.
[0094] In some embodiments, any unvisited starting data point in the graph feature matrix is selected as the first point in the new space clustering cluster, and the neighboring data points whose distance from the first point is less than or equal to the distance metric threshold are clustered into the current space clustering cluster. The selection and clustering process are iteratively executed until all data points in the graph feature matrix are visited, and multiple attack gang clusters are obtained, including:
[0095] Based on the density-based spatial DBSCAN clustering algorithm, any unvisited starting data point in the graph feature matrix is selected for access, and it is judged whether there are neighboring data points of the current starting data point that are greater than or equal to the core point quantity threshold;
[0096] If there are neighboring data points that are greater than or equal to the core point quantity threshold, the current starting data point is used as the first point in the new space clustering cluster, and with the first point as the center, the neighboring data points whose distance from the first point is less than or equal to the distance metric threshold are clustered into the new space clustering cluster;
[0097] If there are no neighboring data points that are greater than or equal to the core point quantity threshold, any unvisited starting data point is reselected for access, and this process is iteratively executed until all data points in the graph feature matrix are visited, and multiple attack gang clusters are obtained.
[0098] In the actual execution process, since one attack may contain several alarm logs, and different attacks may have overlapping time windows, it is not possible to simply use the session window size setting or the time interval threshold between two consecutive logs as the splitting method for two different attacks. Therefore, the embodiment of the present invention selects a method based on alarm log distance measurement and unsupervised clustering to aggregate the alarm logs belonging to the same attack together and can achieve incremental aggregation.
[0099] After obtaining the alarm log feature vectors, it is necessary to define the distance between two feature vectors. For the first n - 1 dimensions, if the values of a certain dimension are the same, the distance of this dimension is 1; otherwise, depending on the semantics expressed by the feature, it is mapped to a number between 0 and 1. If this dimension represents an important attack feature, weights can be considered to be introduced so that the distance between two attack samples with the same value or semantically close values in this dimension is relatively closer. For the last dimension of timestamp information, the L2 norm is used to calculate the distance metric, and then appropriate weights are introduced to adjust its average value to be of the same order of magnitude as the distance of the first n - 1 dimensions and generate appropriate time aggregation and splitting semantics. Finally, the distances of each dimension are added together or combined in other ways (such as using an autoencoder based on a neural network for feature representation learning) to obtain the distance between two feature vectors.
[0100] After obtaining the distance between two nodes, an unsupervised clustering algorithm is used to perform attack partitioning. Since in the actual scenario, the number of attacks is unknown information, it is not possible to use hard clustering algorithms represented by the K-Means algorithm. The DBSCAN clustering algorithm is a density-based spatial clustering algorithm that does not require specifying the number of clusters, so it meets the requirements. This algorithm divides the regions with sufficient density into clusters and discovers clusters of any shape in a spatial database with noise. It defines a cluster as the largest set of density-connected points. The basic steps of the algorithm are as follows.
[0101] Step 1: Start from an arbitrary starting data point that has not been visited.
[0102] Step 2: Mark this point as visited. If there are at least minPoints points in the neighborhood of this point, the current data point becomes the first point in the new cluster.
[0103] Step 3: All the points adjacent to the first point in the new cluster belong to the same cluster.
[0104] Step 4: Repeat the process of Step 2 and Step 3 until all the points in this cluster are determined, that is, all the points near the cluster have been visited.
[0105] Step 5: Jump to Step 1 until all points are marked as visited, complete the clustering of attack groups, and obtain multiple attack group clusters.
[0106] Through the above clustering process, for several alarm log sequences input to the algorithm, they can be divided into several subsequences, and each subsequence represents an attack that has occurred. For newly generated alarm logs, the DBSCAN algorithm can perform efficient incremental clustering, thereby realizing the update of subsequences, supplementing existing attack processes, or constructing new attack processes.
[0107] In summary, according to the attack group clustering method based on the unsupervised clustering algorithm proposed in the embodiments of the present invention, the following
[0108] Advantages are as follows:
[0109] (1) It can group attackers into different groups according to the similarity of attack behavior data, avoiding the subjectivity and errors of manual classification, improving the accuracy of recognition, and thus more effectively tracking and combating criminal activities;
[0110] (2) By clustering and analyzing large-scale data sets, potential attack groups can be mined from them, these potential threats can be discovered in a timely manner, and corresponding defense measures can be taken;
[0111] (3) Through in-depth analysis of attack groups, laws such as attack means, target selection, and activity time of attack groups can be discovered, effectively identifying the patterns and trends of the entire attack ecosystem, and thus formulating more effective defense strategies;
[0112] (4) The identification and classification of attack groups can be achieved in an automated manner, greatly reducing the network security cost. At the same time, this method can also improve the efficiency and accuracy of network security defense, making network security defense more efficient and intelligent.
[0113] Next, a description is given of an attack group clustering device based on an unsupervised clustering algorithm proposed in the embodiments of the present invention with reference to the accompanying drawings.
[0114] Figure 3 It is a block diagram of an attack group clustering device based on an unsupervised clustering algorithm according to an embodiment of the present invention.
[0115] As Figure 3 shown, the attack group clustering device 30 based on the unsupervised clustering algorithm includes: a collection module 301, a construction module 302, and a clustering module 303.
[0116] Among them, the acquisition module 301 is used to acquire multiple alarm messages and perform field association on the multiple alarm messages to obtain an alarm log. The construction module 302 is used to extract multiple feature variables from the alarm log and construct a graph feature matrix according to the multiple feature variables. The clustering module 303 is used to select any unvisited starting data point in the graph feature matrix as the first point in the new spatial clustering cluster, and cluster the neighboring data points whose distance from the first point is less than or equal to the distance metric threshold into the current spatial clustering cluster, and iteratively execute the selection and clustering process until all data points in the graph feature matrix are visited, obtaining multiple attack gang clusters.
[0117] In some embodiments, the acquisition module 301 includes:
[0118] An acquisition unit, configured to acquire multiple alarm messages based on a honeypot and security devices;
[0119] An association unit, which performs field association on the multiple alarm messages to obtain an alarm log, where the alarm log includes the identity information of the attacker, the information of the attacked target, the characteristics of the attack means, and the time when the attack occurred.
[0120] In some embodiments, the construction module 302 includes:
[0121] An extraction unit, configured to extract multiple feature variables from the alarm log;
[0122] A first construction unit, configured to split each feature variable into nodes formed by concatenating the source IP address and the destination IP address strings, and construct an initial graph feature matrix with the feature vectors of each node;
[0123] An embedding unit, configured to perform embedding on the initial graph feature matrix by using a heterogeneous graph variational autoencoder to obtain new feature vectors of each node and edge after learning;
[0124] A second construction unit, configured to construct a graph feature matrix with the feature vectors of each node and the new feature vectors of each node and edge after learning.
[0125] In some embodiments, the clustering module 303 includes:
[0126] An access unit, configured to select any unvisited starting data point in the graph feature matrix for access based on the DBSCAN clustering algorithm, and determine whether there are neighboring data points greater than or equal to the core point quantity threshold for the current starting data point;
[0127] A clustering unit, which is configured to, if there are neighboring data points greater than or equal to the core point quantity threshold, use the current starting data point as the first point in a new spatial clustering cluster, and with the first point as the center, cluster neighboring data points whose distance from the first point is less than or equal to the distance metric threshold into the new spatial clustering cluster;
[0128] A re - visit unit, which is configured to, if there are no neighboring data points greater than or equal to the core point quantity threshold, re - select any unvisited starting data point for access, and iteratively execute the selection and clustering process until all data points in the graph feature matrix have been visited, obtaining multiple attack gang clusters.
[0129] It should be noted that the foregoing explanation of the embodiment of the attack gang clustering method based on the unsupervised clustering algorithm also applies to the attack gang clustering device based on the unsupervised clustering algorithm in this embodiment, and will not be elaborated here.
[0130] The attack gang clustering device based on the unsupervised clustering algorithm proposed according to the embodiment of the present invention has the following beneficial effects:
[0131] (1) It can group attackers into different gangs according to the similarity of attack behavior data, avoiding the subjectivity and errors of manual classification, improving the accuracy of identification, and thus more effectively tracking and combating criminal activities;
[0132] (2) By clustering and analyzing large - scale data sets, potential attack gangs can be mined from them, potential threats can be discovered in a timely manner, and corresponding defense measures can be taken;
[0133] (3) Through in - depth analysis of attack gangs, rules such as attack means, target selection, and activity time of attack gangs can be discovered, effectively identifying the patterns and trends of the entire attack ecosystem, and thus formulating more effective defense strategies;
[0134] (4) The identification and classification of attack gangs can be achieved in an automated manner, greatly reducing the network security cost. At the same time, this method can also improve the efficiency and accuracy of network security defense, making network security defense more efficient and intelligent.
[0135] Figure 4 The structural schematic diagram of the electronic device provided by the embodiment of the present invention. The electronic device may include:
[0136] A memory 401, a processor 402, and a computer program stored on the memory 401 and executable on the processor 402.
[0137] When the processor 402 executes the program, it implements the attack gang clustering method based on the unsupervised clustering algorithm provided in the above - mentioned embodiment.
[0138] Further, the electronic device further includes:
[0139] A communication interface 403 for communication between the memory 401 and the processor 402.
[0140] A memory 401 for storing a computer program that can run on the processor 402.
[0141] The memory 401 may include a high-speed RAM memory and may also include a non-volatile memory, such as at least one disk memory.
[0142] If the memory 401, the processor 402, and the communication interface 403 are implemented independently, the communication interface 403, the memory 401, and the processor 402 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0143] Optionally, in a specific implementation, if the memory 401, the processor 402, and the communication interface 403 are integrated on a single chip, the memory 401, the processor 402, and the communication interface 403 can communicate with each other through an internal interface.
[0144] The processor 402 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0145] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the clustering method of attack groups based on an unsupervised clustering algorithm as described above is implemented.
[0146] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0147] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0148] Any process or method description shown in a flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or N executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0149] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or N wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0150] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0151] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0152] In addition, each functional unit in various embodiments of the present invention may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0153] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. An attack gang clustering method based on an unsupervised clustering algorithm, characterized in that Including the following steps: Collect multiple alarm messages, and perform field association on the multiple alarm messages to obtain an alarm log; Extract multiple feature variables from the alarm log, and construct a graph feature matrix according to the multiple feature variables; Select any unvisited starting data point in the graph feature matrix as the first point in the new spatial clustering cluster, and cluster the neighboring data points whose distance from the first point is less than or equal to the distance metric threshold into the current spatial clustering cluster. Iteratively execute the selection and clustering process until all data points in the graph feature matrix are visited, obtaining multiple attack gang clusters.
2. The attack gang clustering method based on the unsupervised clustering algorithm according to claim 1, wherein The step of collecting multiple alarm messages and performing field association on the multiple alarm messages to obtain an alarm log includes: Collect the multiple alarm messages based on honeypots and security devices; Perform field association on the multiple alarm messages to obtain the alarm log, where the alarm log includes the identity information of the attacker, the information of the attacked target, the characteristics of the attack means, and the time when the attack occurred.
3. The attack gang clustering method based on the unsupervised clustering algorithm according to claim 1, characterized in that The step of extracting multiple feature variables from the alarm log and constructing a graph feature matrix according to the multiple feature vectors includes: Extract multiple feature variables from the alarm log; Split each feature variable into nodes formed by concatenating the source IP address and the destination IP address strings, and form an initial graph feature matrix with the feature vectors of each node; Use a heterogeneous graph variational autoencoder to embed the initial graph feature matrix to obtain new feature vectors of each node and edge after learning; Form the graph feature matrix with the feature vectors of each node and the new feature vectors of each node and edge after learning.
4. The attack gang clustering method based on the unsupervised clustering algorithm according to claim 1, characterized in that, The step of selecting any unvisited starting data point in the graph feature matrix as the first point in the new spatial clustering cluster, and clustering the neighboring data points whose distance from the first point is less than or equal to the distance metric threshold into the current spatial clustering cluster. Iteratively execute the selection and clustering process until all data points in the graph feature matrix are visited, obtaining multiple attack gang clusters, includes: Based on the density-based spatial clustering of applications with noise (DBSCAN) clustering algorithm, select any unvisited starting data point in the graph feature matrix for access, and determine whether there are neighboring data points with a number greater than or equal to the core point quantity threshold for the current starting data point; If there are neighboring data points with a number greater than or equal to the core point quantity threshold, use the current starting data point as the first point in the new spatial clustering cluster, and with the first point as the center, cluster the neighboring data points whose distance from the first point is less than or equal to the distance metric threshold into the new spatial clustering cluster; If there are no neighboring data points with a number greater than or equal to the core point quantity threshold, re-select any unvisited starting data point for access, and iteratively execute this process until all data points in the graph feature matrix are visited, obtaining the multiple attack gang clusters.
5. An attack gang clustering device based on an unsupervised clustering algorithm, characterized in that, Including: A collection module, configured to collect multiple alarm messages, and perform field association on the multiple alarm messages to obtain an alarm log; A construction module, configured to extract multiple feature variables from the alarm log and construct a graph feature matrix according to the multiple feature variables; A clustering module, configured to select any unvisited starting data point in the graph feature matrix as the first point in a new spatial clustering cluster, and cluster neighboring data points whose distance from the first point is less than or equal to a distance metric threshold into the current spatial clustering cluster, and iteratively execute the selection and clustering processes until all data points in the graph feature matrix are visited, obtaining multiple attack gang clusters.
6. The attack gang clustering device based on the unsupervised clustering algorithm according to claim 5, characterized in that The acquisition module includes: An acquisition unit, configured to acquire the multiple alarm messages based on honeypots and security devices; An association unit, configured to perform field association on the multiple alarm messages to obtain the alarm log, where the alarm log includes the identity information of the attacker, the information of the attacked target, the attack means characteristics, and the time when the attack occurred.
7. The attack gang clustering device based on the unsupervised clustering algorithm according to claim 5, characterized in that The construction module includes: An extraction unit, configured to extract multiple feature variables from the alarm log; A first construction unit, configured to split each feature variable into nodes formed by concatenating the source IP address and the destination IP address strings, and construct an initial graph feature matrix with the feature vectors of each node; An embedding unit, configured to perform embedding on the initial graph feature matrix by using a heterogeneous graph variational autoencoder to obtain new feature vectors of each node and edge after learning; A second construction unit, configured to construct the graph feature matrix with the feature vectors of each node and the new feature vectors of each node and edge after learning.
8. The attack gang clustering device based on the unsupervised clustering algorithm according to claim 5, characterized in that, The clustering module includes: An access unit, configured to, based on the DBSCAN clustering algorithm, select any unvisited starting data point in the graph feature matrix for access, and determine whether there are neighboring data points whose number is greater than or equal to a core point quantity threshold for the current starting data point; A clustering unit, configured to, if there are neighboring data points whose number is greater than or equal to the core point quantity threshold, use the current starting data point as the first point in a new spatial clustering cluster, and with the first point as the center, cluster neighboring data points whose distance from the first point is less than or equal to the distance metric threshold into the new spatial clustering cluster; A re-access unit, configured to, if there are no neighboring data points whose number is greater than or equal to the core point quantity threshold, re-select any unvisited starting data point for access, and iteratively execute the selection and clustering processes until all data points in the graph feature matrix are visited, obtaining the multiple attack gang clusters.
9. An electronic device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the attack gang clustering method based on an unsupervised clustering algorithm according to any one of claims 1-4.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to be used to implement the attack gang clustering method based on an unsupervised clustering algorithm according to any one of claims 1-4.