An attack group similarity processing method and device

By constructing and fusing static and dynamic similarity matrices of attack groups, the problem of inaccurate attack group similarity determination in existing technologies is solved, achieving more efficient network intelligence analysis.

CN118487798BActive Publication Date: 2025-11-21QI AN XIN TECHNOLOGY GROUP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410527522.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-28
Publication Date
2025-11-21
Estimated Expiration
2044-04-28

AI Technical Summary

Technical Problem

Existing technologies are not very effective in identifying and analyzing connections and similarities between attack groups, resulting in poor accuracy and efficiency in determining the similarity of attack groups.

Method used

By acquiring static and dynamic information about the attacking groups, static and dynamic similarity matrices are constructed respectively, and feature fusion processing is performed to generate a fused similarity matrix to determine the similarity between the attacking groups.

Benefits of technology

It improved the accuracy and efficiency of attack group similarity determination and enhanced network intelligence analysis capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118487798B_ABST
    Figure CN118487798B_ABST
Patent Text Reader

Abstract

The application provides an attack group similarity processing method and device. The method comprises the following steps: obtaining static information of at least two attack groups to be processed and dynamic information of the attack groups; obtaining a static similarity matrix between each two attack groups based on the static information of each attack group; obtaining a dynamic similarity matrix between each two attack groups based on the dynamic information of each attack group; performing feature fusion processing on the static similarity matrix and the dynamic similarity matrix to obtain a fused similarity matrix; and determining the similarity between the attack groups based on a similarity score in the fused similarity matrix. The attack group similarity processing method provided by the application can effectively improve the accuracy and efficiency of attack group similarity determination, thereby effectively improving the network intelligence analysis capability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of network security, in particular to an attack gang similarity processing method and device. In addition, an electronic device and a processor readable storage medium are also disclosed. BACKGROUND

[0002] In recent years, the network security situation is becoming increasingly severe, and the network security field is facing the challenge of organization and scale. With the continuous development of attacks in a high-level and sustainable direction, more and more attack gangs are driven by interests to use controlled attack resources to implement fine-grained attacks on targets, which are highly threatening. In the face of complex and variable attack gangs, the existing technology mostly identifies attack gangs by analyzing the attack behaviors of different attackers, lacks a method to identify and analyze the connection and similarity between different attack gangs, and the effect is not good, resulting in poor accuracy and efficiency of actual attack gang similarity determination. SUMMARY

[0003] Therefore, the present application provides an attack gang similarity processing method and device to solve the defects of poor effect of the attack gang similarity processing scheme in the prior art, resulting in poor accuracy and efficiency of actual attack gang similarity determination.

[0004] In a first aspect, the present application provides an attack gang similarity processing method, comprising:

[0005] Obtaining static information of at least two attack gangs to be processed and dynamic information of the attack gangs;

[0006] Based on the static information of each attack gang, a static similarity matrix between each two gangs in the attack gangs is obtained;

[0007] Based on the dynamic information of each attack gang, a dynamic similarity matrix between each two gangs in the attack gangs is obtained;

[0008] Performing feature fusion processing on the static similarity matrix and the dynamic similarity matrix to obtain a fusion similarity matrix, and determining the similarity between the attack gangs based on the similarity score in the fusion similarity matrix.

[0009] Further, the static similarity matrix between each two gangs in the attack gangs is obtained based on the static information of each attack gang, specifically comprising:

[0010] Performing static feature extraction on the static information of each attack gang to obtain the static features of the static information of each attack gang;

[0011] The static features are input into a preset static clustering model to obtain the cluster information corresponding to each attack group in the attack group output by the static clustering model; wherein, the static clustering model is trained based on the static features of the samples and the static clustering labels corresponding to the static features of the samples, and the static clustering model is used to cluster each attack group according to the static features of the static information of each attack group to obtain the cluster information corresponding to each attack group; the cluster information includes the cluster identifier, the name of the attack group, and the distance between each attack group and the cluster center;

[0012] Based on whether the cluster identifiers in the cluster information corresponding to each attack group are the same, multiple attack groups belonging to the same cluster are identified.

[0013] The distances of multiple attack groups belonging to the same cluster to the cluster center are compared with a preset distance threshold. At least two attack groups belonging to the same cluster whose distances to the cluster center are less than or equal to the distance threshold are identified as similar attack groups, and a first similarity score is determined between each pair of similar attack groups.

[0014] The first similarity score is used as the matrix content of the static similarity matrix, and the name of the attacking group in the cluster information is used as the row and column of the static similarity matrix to obtain the static similarity matrix between each pair of attacking groups.

[0015] Furthermore, obtaining the dynamic similarity matrix between each pair of attack groups based on the dynamic information of each attack group specifically includes:

[0016] Dynamic features are extracted from the dynamic information of each attack group to obtain the dynamic features of the dynamic information of each attack group.

[0017] The second similarity score between each pair of attack groups is calculated based on the dynamic features of the dynamic information of each attack group, and the dynamic similarity matrix between each pair of attack groups is obtained based on the second similarity score.

[0018] Furthermore, the step of extracting dynamic features from the dynamic information of each attack group to obtain the dynamic features of the dynamic information of each attack group specifically includes:

[0019] Statistical analysis is performed on the text features in the dynamic information of each attack group to obtain the target dynamic features of the dynamic information of each attack group.

[0020] The number of identical samples shared by the attack groups is determined, and the number of identical samples shared by the attack groups is compared and analyzed with a preset number threshold to obtain the sample association characteristics of the dynamic information of the attack groups.

[0021] The dynamic features of the target and the associated features of the sample are fused to obtain the dynamic features of the attack group's dynamic information.

[0022] Furthermore, the step of calculating a second similarity score between pairs of groups based on the dynamic features of the dynamic information of each attack group, and obtaining a dynamic similarity matrix between pairs of groups within the attack group based on the second similarity score, specifically includes:

[0023] The dynamic features are input into a preset similarity prediction model to obtain the second similarity score between each pair of groups in the attack group;

[0024] The second similarity score is used as the matrix content corresponding to each row and column of the dynamic similarity matrix, and the name of the attacking group is used as the row and column of the dynamic similarity matrix to obtain the dynamic similarity matrix between each pair of groups in the attacking group.

[0025] Furthermore, the step of performing feature fusion processing on the static similarity matrix and the dynamic similarity matrix to obtain a fused similarity matrix specifically includes:

[0026] Determine the static weights corresponding to the static similarity matrix and the dynamic weights corresponding to the dynamic similarity matrix;

[0027] Based on the first similarity score and its static weight in the static similarity matrix between any two groups in the attack group, and the second similarity score and its dynamic weight in the dynamic similarity matrix between any two groups, the fusion similarity between any two groups is calculated to obtain the similarity score between any two groups in the attack group. The similarity score between any two groups in the attack group is then filled into a preset initial fusion similarity matrix to obtain the fusion similarity matrix between the attack groups.

[0028] Secondly, the present invention also provides an attack group similarity processing device, comprising:

[0029] The information acquisition unit is used to acquire static information and dynamic information of at least two attack groups to be processed.

[0030] A static similarity processing unit is used to obtain a static similarity matrix between each pair of attack groups based on the static information of each attack group.

[0031] A dynamic similarity processing unit is used to obtain a dynamic similarity matrix between each pair of attack groups based on the dynamic information of each attack group.

[0032] The similarity processing unit is used to perform feature fusion processing on the static similarity matrix and the dynamic similarity matrix to obtain a fused similarity matrix, and to determine the similarity between the attack groups based on the similarity scores in the fused similarity matrix.

[0033] Furthermore, the static similarity processing unit is specifically used for:

[0034] Static features are extracted from the static information of each attack group to obtain the static features of the static information of each attack group.

[0035] The static features are input into a preset static clustering model to obtain the cluster information corresponding to each attack group in the attack group output by the static clustering model; wherein, the static clustering model is trained based on the static features of the samples and the static clustering labels corresponding to the static features of the samples, and the static clustering model is used to cluster each attack group according to the static features of the static information of each attack group to obtain the cluster information corresponding to each attack group; the cluster information includes the cluster identifier, the name of the attack group, and the distance between each attack group and the cluster center;

[0036] Based on whether the cluster identifiers in the cluster information corresponding to each attack group are the same, multiple attack groups belonging to the same cluster are identified.

[0037] The distances of multiple attack groups belonging to the same cluster to the cluster center are compared with a preset distance threshold. At least two attack groups belonging to the same cluster whose distances to the cluster center are less than or equal to the distance threshold are identified as similar attack groups, and a first similarity score is determined between each pair of similar attack groups.

[0038] The first similarity score is used as the matrix content of the static similarity matrix, and the name of the attacking group in the cluster information is used as the row and column of the static similarity matrix to obtain the static similarity matrix between each pair of attacking groups.

[0039] Furthermore, the dynamic similarity processing unit is specifically used for:

[0040] Dynamic features are extracted from the dynamic information of each attack group to obtain the dynamic features of the dynamic information of each attack group.

[0041] The second similarity score between each pair of attack groups is calculated based on the dynamic features of the dynamic information of each attack group, and the dynamic similarity matrix between each pair of attack groups is obtained based on the second similarity score.

[0042] Furthermore, the step of extracting dynamic features from the dynamic information of each attack group to obtain the dynamic features of the dynamic information of each attack group specifically includes:

[0043] Statistical analysis is performed on the text features in the dynamic information of each attack group to obtain the target dynamic features of the dynamic information of each attack group.

[0044] The number of identical samples shared by the attack groups is determined, and the number of identical samples shared by the attack groups is compared and analyzed with a preset number threshold to obtain the sample association characteristics of the dynamic information of the attack groups.

[0045] The dynamic features of the target and the associated features of the sample are fused to obtain the dynamic features of the attack group's dynamic information.

[0046] Furthermore, the step of calculating a second similarity score between pairs of groups based on the dynamic features of the dynamic information of each attack group, and obtaining a dynamic similarity matrix between pairs of groups within the attack group based on the second similarity score, specifically includes:

[0047] The dynamic features are input into a preset similarity prediction model to obtain the second similarity score between each pair of groups in the attack group;

[0048] The second similarity score is used as the matrix content corresponding to each row and column of the dynamic similarity matrix, and the name of the attacking group is used as the row and column of the dynamic similarity matrix to obtain the dynamic similarity matrix between each pair of groups in the attacking group.

[0049] Furthermore, the similarity processing unit is specifically used for:

[0050] Determine the static weights corresponding to the static similarity matrix and the dynamic weights corresponding to the dynamic similarity matrix;

[0051] Based on the first similarity score and its static weight in the static similarity matrix between any two groups in the attack group, and the second similarity score and its dynamic weight in the dynamic similarity matrix between any two groups, the fusion similarity between any two groups is calculated to obtain the similarity score between any two groups in the attack group. The similarity score between any two groups in the attack group is then filled into a preset initial fusion similarity matrix to obtain the fusion similarity matrix between the attack groups.

[0052] Thirdly, the present invention also provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the attack group similarity processing method as described in any of the above.

[0053] Fourthly, the present invention also provides a processor-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the attack group similarity processing method as described in any of the preceding claims.

[0054] The attack group similarity processing method provided by this invention obtains static and dynamic information of at least two attack groups to be processed. Based on the static information of each attack group, a static similarity matrix between each pair of attack groups is obtained, and based on the dynamic information of each attack group, a dynamic similarity matrix between each pair of attack groups is obtained. Then, the static and dynamic similarity matrices are fused to obtain a fused similarity matrix. The similarity between the attack groups is determined based on the similarity scores in the fused similarity matrix. This method can effectively improve the accuracy and efficiency of attack group similarity determination, thereby effectively enhancing network intelligence analysis capabilities. Attached Figure Description

[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0056] Figure 1 This is a flowchart illustrating the attack group similarity processing method provided in an embodiment of the present invention;

[0057] Figure 2 This is a complete schematic diagram of the attack group similarity processing method provided in the embodiments of the present invention;

[0058] Figure 3 This is a schematic diagram of the static similarity processing procedure provided in an embodiment of the present invention;

[0059] Figure 4 This is a schematic diagram of the dynamic similarity processing process provided in an embodiment of the present invention;

[0060] Figure 5 This is a schematic diagram of the fusion similarity matrix provided in an embodiment of the present invention;

[0061] Figure 6 This is a schematic diagram of the attack group similarity processing device provided in an embodiment of the present invention;

[0062] Figure 7 This is a schematic diagram of the physical structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0064] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar users and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0065] This invention provides a method for processing attack group similarity. First, the input attack group data needs to be standardized. This data includes two main categories: static information and dynamic information. Static information typically comes from threat intelligence data and includes at least one or more categories of data such as attack group location, attacker type, attack frequency, affected area, responsible area, and affected industry. Dynamic information typically includes one or more categories of data such as the attack group's ATT&CK information (including tactics, techniques, and sub-techniques) and sample relationships between attack groups. Then, based on the input static information of all attack groups, static features are extracted, a machine learning static clustering model is trained, and static similarity information between attack groups is output. A static similarity matrix is ​​constructed based on this static similarity information. Based on the input dynamic information of all or individual attack groups, dynamic features are extracted, the dynamic similarity between attack groups is compared, and dynamic similarity information between attack groups is output. A dynamic similarity matrix is ​​constructed based on this dynamic similarity information.

[0066] The following is a detailed description of embodiments of the attack group similarity processing method described in this invention. For example... Figure 1 The diagram shown is a flowchart illustrating the attack group similarity processing method provided in an embodiment of the present invention. The specific process includes the following steps:

[0067] Step 101: Obtain static information and dynamic information of at least two attack groups to be processed.

[0068] In this embodiment of the invention, static and dynamic information of the attack group can be obtained through a preset multi-source input module. Specifically, the information typically originates from threat intelligence data, including at least one or more types of data such as the attack group's geographical location, attacker type, attack frequency, affected area, responsible area, and affected industry. The dynamic information typically includes one or more types of data such as the attack group's ATT&CK information (including tactics, techniques, and sub-techniques) and sample relationships between attack groups.

[0069] Step 102: Based on the static information of each attack group, obtain the static similarity matrix between each pair of attack groups.

[0070] In this embodiment of the invention, firstly, a static similarity processing module is used to extract static features from the static information of each attack group to obtain the static features of the static information of each attack group. Then, the static features are input into a preset static clustering model to obtain the cluster information corresponding to each attack group in the attack group output by the static clustering model. The static clustering model is trained based on the static features of the samples and the corresponding static clustering labels. This model is used to cluster each attack group according to the static features of its static information to obtain cluster information for each attack group. The cluster information includes a cluster identifier, the name of the attack group, and the distance between each attack group and the cluster center. Multiple attack groups belonging to the same cluster are identified based on whether their cluster identifiers are identical. The distances of each attack group to the cluster center are compared with a preset distance threshold. At least two attack groups whose distances to the cluster center are less than or equal to the distance threshold are identified as similar attack groups, and a first similarity score is determined between each pair of similar attack groups. The first similarity score is used as the matrix content of the static similarity matrix, and the names of the attack groups in the cluster information are used as the rows and columns of the static similarity matrix to obtain the static similarity matrix between each pair of attack groups.

[0071] Specifically, static features are extracted based on the input static information. This static information typically originates from threat intelligence data, including at least one or more categories of data such as the attack group's geographic location, attacker type, attack frequency, affected area, responsible area, and affected industry. Based on these static features, and combined with machine learning clustering algorithms, a static clustering model can be generated. It is understood that the machine learning clustering algorithm can be chosen based on the characteristics of the features, such as K-Means++ or DBSCAN, without specific limitations here. Based on the generated static clustering model, cluster information corresponding to each attack group is output. This cluster information includes at least a cluster identifier (i.e., cluster ID) and the distance of the attack group from the cluster center. If the distance of the attack group from the cluster center is less than or equal to a preset first distance threshold, the final output information includes at least the attack group name, cluster ID, and distance of the attack group from the cluster center. Attack groups with the same cluster ID are considered similar groups, and a static similarity matrix of the attack groups is output. Specifically, static features are extracted based on the static information. Static information typically originates from threat intelligence data, including at least one or more categories of data such as attack group region, attacker type, attack frequency, affected area, responsible area, and affected industry. A feasible implementation is to use bag-of-words model, TF-IDF, etc. to statistically analyze text features: (1) Attack group region features: can be divided into region A, region B, region C, etc., and the region features of the attack group are statistically analyzed. (2) Attacker type: can be divided into regional background organizations, network scope organizations, etc., and the attacker type features of the attack group are statistically analyzed. (3) Affected area: usually refers to the area frequently attacked by the attack group, and the affected area features are statistically analyzed. (4) Responsible area: usually refers to the area behind the attack group, and the responsible area features are statistically analyzed. (5) Affected industry: such as pharmaceuticals, energy, etc., and the affected industry features are statistically analyzed. Based on static features, combined with machine learning clustering algorithms, a static clustering model of static information is generated. Considering that it is time-consuming to regenerate the static clustering model each time, a periodic training method is usually used to generate a new static clustering model. Understandably, machine learning clustering algorithms can be chosen based on the characteristics of the features, such as K-Means++ or DBSCAN, without specific limitations here. The input to this static clustering model is the static features of the entire dataset, and the output is cluster information corresponding to each attack group, including at least the cluster ID and the distance of the attack group from the cluster center. Further, the cluster information output by the static clustering model is processed to obtain a static similarity matrix of the attack groups. One feasible implementation is that if the distance of each attack group from the cluster center is less than or equal to a preset first distance threshold, then the final output cluster information includes at least the name of the attack group, the cluster ID, and the distance of the attack group from the cluster center. Understandably, attack groups with the same cluster ID are considered similar groups, and a static similarity matrix of the attack groups is output.The static similarity matrix of the attacking group can be a two-dimensional matrix, with the row and column names being the names of the attacking group, and the matrix content being the first similarity score, which can take values ​​of 0 (i.e., not similar) or 1 (i.e., similar).

[0072] Step 103: Based on the dynamic information of each attack group, obtain the dynamic similarity matrix between each pair of attack groups.

[0073] In embodiments of the present invention, such as Figure 3 As shown, firstly, a dynamic similarity processing module is needed to extract dynamic features from the dynamic information of each attack group to obtain the dynamic features of the dynamic information of each attack group. Then, based on the dynamic features of the dynamic information of each attack group, a second similarity score is calculated between each pair of groups, and a dynamic similarity matrix between each pair of groups in the attack group is obtained based on the second similarity score.

[0074] The specific implementation process of extracting dynamic features from the dynamic information of each attack group to obtain the dynamic features of the dynamic information of each attack group includes: statistically analyzing the text features in the dynamic information of each attack group to obtain the target dynamic features of the dynamic information of each attack group; determining the number of common samples shared by the attack groups and comparing the number of common samples shared by the attack groups with a preset threshold to obtain the sample association features of the dynamic information of the attack groups; and performing feature fusion on the target dynamic features and the sample association features to obtain the dynamic features of the dynamic information of the attack groups (i.e., Figure 4 The process involves: 1) identifying APT (Advanced Persistent Threat) group characteristics. 2) calculating a second similarity score between each pair of groups based on the dynamic features of each attack group's dynamic information. 3) obtaining a dynamic similarity matrix between each pair of groups within the attack group based on the second similarity score. The specific implementation process includes: inputting the dynamic features into a preset similarity prediction model to obtain the second similarity score between the attack groups; 4) using the second similarity score as the matrix content corresponding to each row and column of the dynamic similarity matrix, and using the names of the attack groups as the rows and columns of the dynamic similarity matrix to obtain the dynamic similarity matrix between each pair of groups within the attack group.

[0075] Specifically, such as Figure 4As shown, dynamic features can be extracted based on the dynamic information of the attacking group through the dynamic similarity evaluation module. Among them, dynamic information usually includes one or more types of data such as the attacking group's ATT&CK information (including tactics, techniques, and sub-techniques) and sample relationships between attacking groups. Based on the extracted dynamic features, the dynamic similarity evaluation module uses a similarity algorithm to evaluate the similarity between attacking groups and obtains a second similarity score. Based on the second similarity score, a dynamic similarity matrix of similar groups is output. Dynamic features are extracted from the dynamic information of the attacking group using the attacking group dynamic feature extraction module. Among them, dynamic information usually includes one or more types of data such as the attacking group's ATT&CK information (including tactics, techniques, and sub-techniques) and sample relationships between attacking groups. A feasible embodiment is as follows: (1) ATT&CK features: Based on the tactics, techniques, and sub-techniques commonly used by the attacking group, the bag-of-words model is used to statistically analyze text features to obtain the attacking group's ATT&CK features (i.e., target dynamic features). (2) Sample association features: Based on the number of common samples shared by attacking groups. This section typically recommends incorporating a knowledge graph, which illustrates the relationships between samples from attack groups. If two attack groups share the same samples, the number of identical samples is used to further determine the features. For example, a pre-defined three-level threshold mechanism can be designed: if the number of identical samples is less than the first-level threshold, the feature value is 0.3; if it is between the second and third-level thresholds, the feature value is 0.6; and if it exceeds the third-level threshold, the feature value is 1. If multiple features are extracted, feature fusion needs to be considered. Common fusion methods are concatenation fusion and multinomial fusion. Concatenation fusion directly concatenates the ATT&CK features and sample association features; multinomial fusion introduces the concept of weights, i.e., processing the fused features, for example: (weight 1 * ATT&CK feature, weight 2 * sample association feature).

[0076] The dynamic similarity evaluation module, after obtaining the feature representations of the attack groups, compares the similarity between each pair of attack groups to obtain a dynamic similarity matrix. Understandably, the similarity comparison algorithm can choose common distance calculation methods such as cosine distance or Jaccard distance. Furthermore, the dynamic similarity matrix is ​​typically a two-dimensional matrix, with rows and columns representing the names of the attack groups, and each row and column corresponding to a second similarity score between the two attack groups. The second similarity score can range from [0,1].

[0077] Step 104: Perform feature fusion processing on the static similarity matrix and the dynamic similarity matrix to obtain a fused similarity matrix, and determine the similarity between the attack groups based on the similarity scores in the fused similarity matrix.

[0078] In this embodiment of the invention, it is first necessary to determine the static weights corresponding to the static similarity matrix and the dynamic weights corresponding to the dynamic similarity matrix. Then, based on the first similarity score in the static similarity matrix and its static weights, and the second similarity score in the dynamic similarity matrix and its dynamic weights, a similarity score is obtained. A fused similarity matrix is ​​then obtained based on the similarity score. Specifically, as shown... Figure 2 As shown, the result processing module of this invention is used to fuse the static similarity matrix and dynamic similarity matrix of similar groups to output the final similarity matrix. It can be understood that the similarity matrix is ​​only one form of the final output result; its core is to obtain the similarity between attacking groups. (See below) Figure 5 The diagram illustrates an embodiment of static and dynamic similarity matrices. For ease of understanding, assume that the first similarity score between attack group 1 and attack group n in the static similarity matrix is ​​1, and the second similarity score between attack group 1 and attack group n in the dynamic similarity matrix is ​​0.3. Then, using feature fusion, we can consider weighting and adding the similarity scores of the two groups to obtain: the similarity score (i.e., the similarity score) between attack group 1 and attack group 3 = 1 * static weight + 0.3 * dynamic weight. Where, static weight + dynamic weight = 1. This outputs the final similarity matrix. In this embodiment, a multi-source approach is used to compare the similarities between attack groups to obtain the final similarity score, which helps to uncover the connections between attack groups and empowers intelligence analysis. Considering the infrequent changes in static information during the static similarity processing, machine learning clustering is used, which helps to uncover the connections between all attack groups and improve the accuracy of similarity determination. In the dynamic similarity processing module, considering the frequent changes in dynamic information, a similarity algorithm is used for real-time calculation. This helps to update the relationships between attack groups in a timely manner and improve intelligence analysis capabilities. This enables accurate identification of connections and common characteristics between different attack groups, playing a crucial role in threat intelligence expansion and malware analysis.

[0079] The attack group similarity processing method described in this invention obtains static and dynamic information of at least two attack groups to be processed. Based on the static information of each attack group, a static similarity matrix is ​​obtained between each pair of attack groups. Based on the dynamic information of each attack group, a dynamic similarity matrix is ​​obtained between each pair of attack groups. Then, the static and dynamic similarity matrices are fused to obtain a fused similarity matrix. The similarity between the attack groups is determined based on the similarity scores in the fused similarity matrix. This method can effectively improve the accuracy and efficiency of attack group similarity determination, thereby effectively enhancing network intelligence analysis capabilities.

[0080] Corresponding to the attack group similarity processing method provided above, this invention also provides an attack group similarity processing device. Since the embodiments of this device are similar to the above method embodiments, the description is relatively simple. For relevant details, please refer to the description in the above method embodiment section. The embodiments of the attack group similarity processing device described below are merely illustrative. Please refer to... Figure 6 The diagram shown is a structural schematic of an attack group similarity processing device provided in an embodiment of the present invention. The attack group similarity processing device of the present invention specifically includes the following parts:

[0081] The information acquisition unit 601 is used to acquire static information and dynamic information of at least two attack groups to be processed.

[0082] The static similarity processing unit 602 is used to obtain a static similarity matrix between each pair of attack groups based on the static information of each attack group.

[0083] The dynamic similarity processing unit 603 is used to obtain a dynamic similarity matrix between each pair of attack groups based on the dynamic information of each attack group.

[0084] The similarity processing unit 604 is used to perform feature fusion processing on the static similarity matrix and the dynamic similarity matrix to obtain a fused similarity matrix, and to determine the similarity between the attack groups based on the similarity scores in the fused similarity matrix.

[0085] Furthermore, the static similarity processing unit is specifically used for:

[0086] Static features are extracted from the static information of each attack group to obtain the static features of the static information of each attack group.

[0087] The static features are input into a preset static clustering model to obtain the cluster information corresponding to each attack group in the attack group output by the static clustering model; wherein, the static clustering model is trained based on the static features of the samples and the static clustering labels corresponding to the static features of the samples, and the static clustering model is used to cluster each attack group according to the static features of the static information of each attack group to obtain the cluster information corresponding to each attack group; the cluster information includes the cluster identifier, the name of the attack group, and the distance between each attack group and the cluster center;

[0088] Based on whether the cluster identifiers in the cluster information corresponding to each attack group are the same, multiple attack groups belonging to the same cluster are identified.

[0089] The distances of multiple attack groups belonging to the same cluster to the cluster center are compared with a preset distance threshold. At least two attack groups belonging to the same cluster whose distances to the cluster center are less than or equal to the distance threshold are identified as similar attack groups, and a first similarity score is determined between each pair of similar attack groups.

[0090] The first similarity score is used as the matrix content of the static similarity matrix, and the name of the attacking group in the cluster information is used as the row and column of the static similarity matrix to obtain the static similarity matrix between each pair of attacking groups.

[0091] Furthermore, the dynamic similarity processing unit is specifically used for:

[0092] Dynamic features are extracted from the dynamic information of each attack group to obtain the dynamic features of the dynamic information of each attack group.

[0093] The second similarity score between each pair of attack groups is calculated based on the dynamic features of the dynamic information of each attack group, and the dynamic similarity matrix between each pair of attack groups is obtained based on the second similarity score.

[0094] Furthermore, the step of extracting dynamic features from the dynamic information of each attack group to obtain the dynamic features of the dynamic information of each attack group specifically includes:

[0095] Statistical analysis is performed on the text features in the dynamic information of each attack group to obtain the target dynamic features of the dynamic information of each attack group.

[0096] The number of identical samples shared by the attack groups is determined, and the number of identical samples shared by the attack groups is compared and analyzed with a preset number threshold to obtain the sample association characteristics of the dynamic information of the attack groups.

[0097] The dynamic features of the target and the associated features of the sample are fused to obtain the dynamic features of the attack group's dynamic information.

[0098] Furthermore, the step of calculating a second similarity score between pairs of groups based on the dynamic features of the dynamic information of each attack group, and obtaining a dynamic similarity matrix between pairs of groups within the attack group based on the second similarity score, specifically includes:

[0099] The dynamic features are input into a preset similarity prediction model to obtain the second similarity score between each pair of groups in the attack group;

[0100] The second similarity score is used as the matrix content corresponding to each row and column of the dynamic similarity matrix, and the name of the attacking group is used as the row and column of the dynamic similarity matrix to obtain the dynamic similarity matrix between each pair of groups in the attacking group.

[0101] Furthermore, the similarity processing unit is specifically used for:

[0102] Determine the static weights corresponding to the static similarity matrix and the dynamic weights corresponding to the dynamic similarity matrix;

[0103] Based on the first similarity score and its static weight in the static similarity matrix between any two groups in the attack group, and the second similarity score and its dynamic weight in the dynamic similarity matrix between any two groups, the fusion similarity between any two groups is calculated to obtain the similarity score between any two groups in the attack group. The similarity score between any two groups in the attack group is then filled into a preset initial fusion similarity matrix to obtain the fusion similarity matrix between the attack groups.

[0104] The attack group similarity processing device described in this embodiment of the invention obtains static information and dynamic information of at least two attack groups to be processed. Based on the static information of each attack group, it obtains a static similarity matrix between each pair of attack groups, and based on the dynamic information of each attack group, it obtains a dynamic similarity matrix between each pair of attack groups. Then, it performs feature fusion processing on the static similarity matrix and the dynamic similarity matrix to obtain a fused similarity matrix, and determines the similarity between the attack groups based on the similarity scores in the fused similarity matrix. This device can effectively improve the accuracy and efficiency of attack group similarity determination, thereby effectively enhancing network intelligence analysis capabilities.

[0105] Corresponding to the attack group similarity processing method provided above, this invention also provides an electronic device. Since the embodiment of this electronic device is similar to the method embodiment described above, the description is relatively simple. For relevant details, please refer to the description in the method embodiment section above. The electronic device described below is merely illustrative. Figure 7The diagram shows a physical structure of an electronic device disclosed in an embodiment of the present invention. The electronic device may include a processor 701, a memory 702, and a communication bus 703. The processor 701 and the memory 702 communicate with each other via the communication bus 703 and communicate with the outside world via a communication interface 704. The processor 701 can call logical instructions in the memory 702 to execute an attack group similarity processing method. This method includes: acquiring static information and dynamic information of at least two attack groups to be processed; obtaining a static similarity matrix between each pair of attack groups based on the static information of each attack group; obtaining a dynamic similarity matrix between each pair of attack groups based on the dynamic information of each attack group; performing feature fusion processing on the static similarity matrix and the dynamic similarity matrix to obtain a fused similarity matrix; and determining the similarity between the attack groups based on the similarity scores in the fused similarity matrix.

[0106] Furthermore, the logical instructions in the aforementioned memory 702 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as memory chips, USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0107] On the other hand, embodiments of the present invention also provide a computer program product, the computer program product including a computer program stored on a processor-readable storage medium, the computer program including program instructions, and when the program instructions are executed by a computer, the computer is able to execute the attack group similarity processing method provided in the above-described method embodiments. The method includes: acquiring static information and dynamic information of at least two attack groups to be processed; obtaining a static similarity matrix between each pair of attack groups based on the static information of each attack group; obtaining a dynamic similarity matrix between each pair of attack groups based on the dynamic information of each attack group; performing feature fusion processing on the static similarity matrix and the dynamic similarity matrix to obtain a fused similarity matrix; and determining the similarity between the attack groups based on the similarity scores in the fused similarity matrix.

[0108] In another aspect, embodiments of the present invention also provide a processor-readable storage medium storing a computer program, which, when executed by a processor, implements the attack group similarity processing method provided in the above embodiments. The method includes: acquiring static information and dynamic information of at least two attack groups to be processed; obtaining a static similarity matrix between each pair of attack groups based on the static information of each attack group; obtaining a dynamic similarity matrix between each pair of attack groups based on the dynamic information of each attack group; performing feature fusion processing on the static similarity matrix and the dynamic similarity matrix to obtain a fused similarity matrix; and determining the similarity between the attack groups based on the similarity scores in the fused similarity matrix.

[0109] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).

[0110] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0111] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for processing the similarity of attack groups, characterized in that, include: Obtain static and dynamic information of at least two attack groups to be processed; Based on the static information of each attack group, a static similarity matrix between each pair of attack groups is obtained. Based on the dynamic information of each attack group, a dynamic similarity matrix between each pair of attack groups is obtained. The static similarity matrix and the dynamic similarity matrix are subjected to feature fusion processing to obtain a fused similarity matrix. The similarity between the attack groups is determined based on the similarity scores in the fused similarity matrix. The step of obtaining a static similarity matrix between each pair of attack groups based on the static information of each attack group specifically includes: Static features are extracted from the static information of each attack group to obtain the static features of the static information of each attack group. The static features are input into a preset static clustering model to obtain the cluster information corresponding to each attack group in the attack group output by the static clustering model; wherein, the static clustering model is trained based on the static features of the samples and the static clustering labels corresponding to the static features of the samples, and the static clustering model is used to cluster each attack group according to the static features of the static information of each attack group to obtain the cluster information corresponding to each attack group; the cluster information includes the cluster identifier, the name of the attack group, and the distance between each attack group and the cluster center; Based on whether the cluster identifiers in the cluster information corresponding to each attack group are the same, multiple attack groups belonging to the same cluster are identified. The distances of multiple attack groups belonging to the same cluster to the cluster center are compared with a preset distance threshold. At least two attack groups belonging to the same cluster whose distances to the cluster center are less than or equal to the distance threshold are identified as similar attack groups, and a first similarity score is determined between each pair of similar attack groups. The first similarity score is used as the matrix content of the static similarity matrix, and the name of the attacking group in the cluster information is used as the row and column of the static similarity matrix to obtain the static similarity matrix between each pair of attacking groups.

2. The method for processing the similarity of attack groups according to claim 1, characterized in that, The step of obtaining a dynamic similarity matrix between each pair of attack groups based on the dynamic information of each attack group specifically includes: Dynamic features are extracted from the dynamic information of each attack group to obtain the dynamic features of the dynamic information of each attack group. The second similarity score between each pair of attack groups is calculated based on the dynamic features of the dynamic information of each attack group, and the dynamic similarity matrix between each pair of attack groups is obtained based on the second similarity score.

3. The method for processing the similarity of attack groups according to claim 2, characterized in that, The step of extracting dynamic features from the dynamic information of each attack group to obtain the dynamic features of the dynamic information of each attack group specifically includes: Statistical analysis is performed on the text features in the dynamic information of each attack group to obtain the target dynamic features of the dynamic information of each attack group. The number of identical samples shared by the attack groups is determined, and the number of identical samples shared by the attack groups is compared and analyzed with a preset number threshold to obtain the sample association characteristics of the dynamic information of the attack groups. The dynamic features of the target and the associated features of the sample are fused to obtain the dynamic features of the attack group's dynamic information.

4. The method for processing the similarity of attack groups according to claim 2, characterized in that, The calculation of a second similarity score between pairs of attack groups based on the dynamic features of the dynamic information of each attack group, and the obtaining of a dynamic similarity matrix between pairs of attack groups based on the second similarity score, specifically includes: The dynamic features are input into a preset similarity prediction model to obtain the second similarity score between each pair of groups in the attack group; The second similarity score is used as the matrix content corresponding to each row and column of the dynamic similarity matrix, and the name of the attacking group is used as the row and column of the dynamic similarity matrix to obtain the dynamic similarity matrix between each pair of groups in the attacking group.

5. The method for processing the similarity of attack groups according to claim 1, characterized in that, The step of performing feature fusion processing on the static similarity matrix and the dynamic similarity matrix to obtain a fused similarity matrix specifically includes: Determine the static weights corresponding to the static similarity matrix and the dynamic weights corresponding to the dynamic similarity matrix; Based on the first similarity score and its static weight in the static similarity matrix between any two groups in the attack group, and the second similarity score and its dynamic weight in the dynamic similarity matrix between any two groups, the fusion similarity between any two groups is calculated to obtain the similarity score between any two groups in the attack group. The similarity score between any two groups in the attack group is then filled into a preset initial fusion similarity matrix to obtain the fusion similarity matrix between the attack groups.

6. A device for processing the similarity of attack groups, characterized in that, include: The information acquisition unit is used to acquire static information and dynamic information of at least two attack groups to be processed. A static similarity processing unit is used to obtain a static similarity matrix between each pair of attack groups based on the static information of each attack group. A dynamic similarity processing unit is used to obtain a dynamic similarity matrix between each pair of attack groups based on the dynamic information of each attack group. The similarity processing unit is used to perform feature fusion processing on the static similarity matrix and the dynamic similarity matrix to obtain a fused similarity matrix, and to determine the similarity between the attack groups based on the similarity scores in the fused similarity matrix; the static similarity processing unit is specifically used for: Static features are extracted from the static information of each attack group to obtain the static features of the static information of each attack group. The static features are input into a preset static clustering model to obtain the cluster information corresponding to each attack group in the attack group output by the static clustering model; wherein, the static clustering model is trained based on the static features of the samples and the static clustering labels corresponding to the static features of the samples, and the static clustering model is used to cluster each attack group according to the static features of the static information of each attack group to obtain the cluster information corresponding to each attack group; the cluster information includes the cluster identifier, the name of the attack group, and the distance between each attack group and the cluster center; Based on whether the cluster identifiers in the cluster information corresponding to each attack group are the same, multiple attack groups belonging to the same cluster are identified. The distances of multiple attack groups belonging to the same cluster to the cluster center are compared with a preset distance threshold. At least two attack groups belonging to the same cluster whose distances to the cluster center are less than or equal to the distance threshold are identified as similar attack groups, and a first similarity score is determined between each pair of similar attack groups. The first similarity score is used as the matrix content of the static similarity matrix, and the name of the attacking group in the cluster information is used as the row and column of the static similarity matrix to obtain the static similarity matrix between each pair of attacking groups.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the attack group similarity processing method as described in any one of claims 1 to 5.

8. A processor-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the attack group similarity processing method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • AI detection method and device for deeply tracking gang attack behaviors

    CN110912861A

  • Traffic flow prediction method based on static and dynamic multi-graph fusion

    CN116052421A