Attack homologous analysis method and device and storage medium
By building an attack source map and a behavior similarity map, analyzing the behavior similarity relationship between the attack source IPs, the problem of low efficiency of attack homologous analysis under massive alarm data is solved, and rapid attack tracing is achieved.
Patent Information
- Application Number
- CN202311815273.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-27
AI Technical Summary
Faced with massive alarm data, how to conduct attack homologous analysis more efficiently has become an urgent problem in the industry.
By building an attack source map for each attack source IP based on security monitoring data, analyzing the behavioral similarity relationship between each attack source IP, and building a behavioral similarity map for attack homologous analysis.
It improves the efficiency of attack homologous analysis, so that the attack source can be quickly determined when an attack alarm is monitored, and the attack source can be traced through map query operations.
Smart Images

Figure CN120223335A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of security technologies, and in particular, to a method, device, and storage medium for analyzing attack homology. Background Art
[0002] Homology analysis of network attacks is a necessary step in tracing the origin of network attacks. Due to the concealment and anonymity of network attacks, the attack source does not directly expose its identity. Therefore, attack homology analysis is to analyze whether different network attacks are organized and whether they originate from the same attack source or attack organization. Through attack homology analysis, it is helpful to more quickly determine the source of network attacks and thus take targeted protection measures in a timely manner.
[0003] Facing a large amount of alarm data, how to perform attack homology analysis more efficiently has become an urgent problem in the industry. Summary of the Invention
[0004] Multiple aspects of this application provide a method, device, and storage medium for analyzing attack homology to improve the efficiency of attack homology analysis.
[0005] An embodiment of this application provides a method for analyzing attack homology, including:
[0006] Based on security monitoring data, construct an attack source map for each discovered attack source IP to describe the attack behavior;
[0007] According to the attack source maps corresponding to each attack source IP, analyze the behavioral similarity relationship between the attack source IPs;
[0008] Based on the behavioral similarity relationship, construct a behavioral similarity map between the attack source IPs;
[0009] Use the behavioral similarity map to perform attack homology analysis.
[0010] Further, according to the attack source maps corresponding to each attack source IP, analyzing the behavioral similarity relationship between the attack source IPs includes:
[0011] Respectively query, from the attack source maps corresponding to each attack source IP, the behavioral link starting from the attack source IP, where the behavioral link includes behavioral entities and the behavioral relationships between the behavioral entities;
[0012] Based on the behavioral links corresponding to each attack source IP, calculate the behavioral similarity between the attack source IPs to characterize the behavioral similarity relationship.
[0013] Further, based on the behavior chains corresponding to each attacking source IP, calculate the behavior similarity between each attacking source IP, including:
[0014] Convert the behavior chains corresponding to each attacking source IP into behavior description texts respectively;
[0015] Calculate the similarity between the behavior description texts corresponding to each attacking source IP as the behavior similarity.
[0016] Further, converting the behavior chains corresponding to each attacking source IP into behavior description texts respectively includes:
[0017] For the target behavior chain corresponding to the target attacking source IP, obtain the attribute fields corresponding to each behavior entity and behavior relationship included in the target behavior chain in the security monitoring data;
[0018] Concatenate the obtained attribute fields in the connection order of the behavior entities in the target behavior chain to obtain the behavior description text corresponding to the target behavior chain.
[0019] Further, based on the behavior chains corresponding to each attacking source IP, calculate the behavior similarity between each attacking source IP, including:
[0020] Save the single behavior chain corresponding to a single attacking source IP as an independent subgraph;
[0021] Calculate the representation vectors for each independent subgraph associated with each attacking source IP to obtain the representation vector groups corresponding to each attacking source IP respectively;
[0022] Calculate the similarity between the representation vector groups corresponding to each attacking source IP as the behavior similarity.
[0023] Further, based on the security monitoring data, construct an attacking source graph for each discovered attacking source IP to describe the attacking behavior, including:
[0024] Perform data connection on various data streams included in the security monitoring data to obtain several behavior record data;
[0025] Extract the behavior entities and the behavior relationships between the behavior entities from the behavior record data, where the attacking source IP is one of the behavior entities;
[0026] Based on the extracted behavior entities and the behavior relationships between the behavior entities, perform graph construction to obtain the attacking source graphs corresponding to each attacking source IP respectively.
[0027] Further, the method further includes:
[0028] Transfer the extracted behavior entities and the behavior relationships between the behavior entities to a data warehouse for data traceability.
[0029] Further, the method further includes:
[0030] Perform data stitching on the attack organization disclosure data and the security monitoring data;
[0031] If there is an attack source IP that can be associated with a known attack organization, add an attack organization entity associated with the attack source IP to the attack source graph corresponding to the attack source IP.
[0032] Further, the method further includes:
[0033] Cluster the respective attack source IPs based on the behavior similarity graph to obtain a plurality of attack source IP clusters;
[0034] If there is an attack source IP in the target attack source IP cluster that has been associated with a known attack organization, add the attack organization entity corresponding to the known attack organization to the attack source graphs corresponding to the attack source IPs in the target attack source IP cluster that have not been associated with an attack organization.
[0035] Further, the method further includes:
[0036] If there is no attack source IP in the target attack source IP cluster that has been associated with a known attack organization, configure a custom attack organization identifier for the target attack source IP;
[0037] Add the attack organization entity corresponding to the custom attack organization identifier to the attack source graphs corresponding to the attack source IPs included in the target attack source IP cluster.
[0038] Further, the method further includes:
[0039] In response to an attack organization traceability instruction for a specified attack, determine the attack source IP corresponding to the specified attack;
[0040] Query the attack organization entity from the attack source graph corresponding to the determined attack source IP to determine the attack organization to which the specified attack belongs.
[0041] Further, the security monitoring data includes one or more data streams such as alarm data and network log data.
[0042] An embodiment of the present application further provides a computing device, including a memory, a processor, and a communication component;
[0043] The memory is used to store one or more computer instructions;
[0044] The processor is coupled to the memory and the communication component and is configured to execute the one or more computer instructions for performing the foregoing attack homology analysis method.
[0045] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to execute the foregoing attack homology analysis method.
[0046] In the embodiment of the present application, the attack source IP can be discovered from a large amount of security monitoring data, and an attack source map can be constructed for each discovered attack source IP based on the security monitoring data, so as to mine the attack behaviors corresponding to each attack source IP through the attack source map. On this basis, according to the attack source maps corresponding to the respective attack source IPs, the behavioral similarity relationships between the respective attack source IPs can be analyzed, and then a behavioral similarity map can be constructed, and the behavioral similarity map can be used as the basis for attack homology analysis. In this way, when an attack alert is detected, the attack source IP corresponding to the current attack can be determined first, and then, by querying the behavioral similarity map, similar attack source IPs can be found for the current attack, and then the attack homology analysis of the current attack can be realized. Accordingly, the attack behaviors of different attack source IPs can be mined through the attack source map, and the behavioral similarity map can be constructed more efficiently based on the attack source map, and then the attack homology analysis can be converted into a map query operation, which can effectively improve the efficiency of attack homology analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The exemplary embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:
[0048] Figure 1 is a schematic flowchart of an attack homology analysis method provided by an exemplary embodiment of the present application;
[0049] Figure 2 is a schematic logical diagram of an attack homology analysis method provided by an exemplary embodiment of the present application;
[0050] Figure 3 is a schematic logical diagram of a preferred implementation manner of constructing an attack source map provided by an exemplary embodiment of the present application;
[0051] Figure 4 is a schematic diagram of an exemplary attack source map provided by an exemplary embodiment of the present application;
[0052] Figure 5 is a schematic flowchart of a clustering scheme provided by an exemplary embodiment of the present application;
[0053] Figure 6 A structural schematic diagram of a computing device provided for another exemplary embodiment of the present application. Detailed implementation manners
[0054] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0055] As introduced in the background art, attack homology analysis is an important means for attack traceability. With the increasing complexity and concealment of network attacks, in the network security scenario, the demand for attack homology analysis is stronger. If the attack source can be determined in time after an attack is discovered, targeted emergency response and protection measures can be taken. Currently, according to the different types of clues, attack homology analysis can be mainly divided into malicious sample homology analysis and attack behavior homology analysis. Among them, there are relatively few research results on attack behavior homology analysis and most of them are limited to the laboratory environment. Moreover, most of these research results require a large amount of feature extraction work, which makes these research results often overwhelmed by a large amount of alarm data when applied to the actual network security scenario, resulting in the inability to correctly perceive attack behaviors and thus unable to complete attack homology analysis.
[0056] Therefore, the embodiments of the present application propose an attack homology analysis method to achieve attack behavior homology analysis. In the face of a large amount of data, the attack behavior can still be clearly determined, so as to accurately complete attack homology analysis.
[0057] The following will detail the technical solutions provided by the embodiments of the present application with reference to the drawings.
[0058] Figure 1 A flowchart of an attack homology analysis method provided for an exemplary embodiment of the present application, Figure 2 A logic diagram of an attack homology analysis method provided for an exemplary embodiment of the present application. This method can be executed by an attack homology analysis system, which can be implemented as software, hardware, or a combination of software and hardware, and can be integrated in a computing device. Refer to Figure 1 , this method may include:
[0059] Step 100: Based on security monitoring data, construct an attack source map for each discovered attack source IP to describe the attack behavior;
[0060] Step 101: Analyze the behavioral similarity relationships among the respective attacking source IPs according to the attacking source graphs corresponding to each attacking source IP.
[0061] Step 102: Based on the behavioral similarity relationships, construct a behavioral similarity graph among the respective attacking source IPs.
[0062] Step 103: Use the behavioral similarity graph to perform attacking source homology analysis.
[0063] In the field of network security, many security service nodes for attack detection are deployed in the network. For example, cloud firewalls, web application firewalls, etc., and the list is not exhaustive here. In this embodiment, the data generated by these security service nodes is described as security monitoring data. In this way, the security monitoring data in this embodiment may include, but is not limited to, alarm data, network log data, etc. Moreover, in actual applications, this data is generated in real time. Therefore, the security monitoring data can be accessed into the attacking source homology analysis system provided in this embodiment in the form of a data stream as the input in Step 100.
[0064] Reference Figure 2 , the attacking source homology analysis system provided in this embodiment may include an attacking source graph construction module, a behavioral similarity analysis module, and a discovery module. Among them, homology can be understood as the same attacking source. Homology analysis can be understood as various analyses related to attacks from the same source.
[0065] In Step 100, the attacking source graph construction module can respectively construct an attacking source graph for describing the attacking behavior for each discovered attacking source IP based on the security monitoring data.
[0066] Based on Step 100, this embodiment innovatively proposes to use the attacking source IP as the core and construct an attacking source graph for it to describe the attacking behavior corresponding to the attacking source IP. Among them, the attacking source IP is the IP that initiates the attacking behavior. The attacking behavior is the behavioral trace left by the attacking source during the attack implementation process. In this embodiment, it is first proposed to use the attacking source IP as the analysis perspective. In this way, although the attacking source may be unknown, the inventors found during the research process that when the same attacking source initiates attacks using different attacking source IPs, these attacking source IPs have similarities in attacking behavior. Therefore, this embodiment proposes to discover the common attacking source behind these attacking source IPs by analyzing the attacking behavior similarities among the attacking source IPs. This enables this embodiment to not only achieve attacking source homology analysis but also efficiently and accurately discover hidden attacking sources. In this embodiment, it is further proposed to construct an attacking source graph for the attacking source IP to graphically represent the relevant attacking behaviors, so that the attacking behaviors corresponding to each attacking source IP are clearer, more accurate, and comprehensive enough.
[0067] In this embodiment, there is no limitation on the implementation method for constructing the attack source graph from security monitoring data. It is only required to accurately extract the required behavioral entities and the behavioral relationships between the behavioral entities from the security monitoring data for each attack source IP and generate a graph. Figure 3 It is a logical schematic diagram of a preferred implementation method for constructing an attack source graph provided by an exemplary embodiment of the present application. Refer to Figure 3 , in this preferred implementation method:
[0068] Multiple data streams included in the security monitoring data can be connected to obtain several pieces of behavior record data; behavioral entities and the behavioral relationships between the behavioral entities are extracted from the behavior record data, where the attack source IP is one of the behavioral entities; based on the extracted behavioral entities and the behavioral relationships between the behavioral entities, graph construction is performed to obtain the attack source graph corresponding to each attack source IP.
[0069] Refer to Figure 3 , in this preferred implementation method, the attack source graph construction module can be docked to multiple data streams included in the security monitoring data with the help of a streaming execution engine such as flink, and data connection (join) is performed on these data streams to generate several pieces of behavior record data. The inventor found during the research process that the alarm data usually records the description information of various abnormal events, including but not limited to the identifiers of the process or file where the abnormal event occurs, the rule name used for the abnormal event, the occurrence frequency of the abnormal event, and the type of the resulting alarm message, etc., which are not enumerated here. And the network log data usually records various operations occurring in the network, including but not limited to the source IP (which can be used as the attack source IP in this embodiment) and destination IP of the network access, the network access path, the identifier of the process or file accessed, etc., which are not enumerated here. Based on this, through data connection, the relevant data in different data streams can be connected together to form behavior record data. For example, there are the same fields - process and file, etc. in the above-mentioned alarm data and network log data, so data connection can be performed through the same fields. It can be seen that in this preferred implementation method, the behavior record data will contain multi-dimensional data.
[0070] On this basis, the preferred implementation proposes that the behavior entities and the behavior relationships between the behavior entities can be extracted from the behavior record data. Among them, the behavior entity can be understood as the subject related to the attack behavior. For example, the object of a certain behavior action or the initiator who initiates the next behavior action under the trigger of the previous behavior action, etc. The behavior relationship can be used to describe the behavior actions that occur between two behavior entities. In this preferred implementation, the types of behavior entities and the types of behavior relationships concerned in the attack source graph can be preset in advance.
[0071] Several exemplary behavior entities may include:
[0072] Attacker source IP: process network connection, DNS resolution, http request...;
[0073] login_event: system login event;
[0074] process: process. Derived from process startup, process snapshots, and other network logs related to process behaviors;
[0075] vuln: vulnerability / baseline / weakness. Derived from vulnerability / baseline scan result data;
[0076] file: file. Derived from alert data, or network logs such as process startup and process file operations;
[0077] alert: various alert messages.
[0078] Several exemplary behavior relationships may include:
[0079] *—trigger_alert: trigger an alert;
[0080] process_spawn_process: a process starts a subprocess;
[0081] process_open_process: a process opens a process;
[0082] process_ptrace_process: a process makes a system call;
[0083] process_connect_ip: a process connects to the external network;
[0084] ip_login_loginevent: the attacker source IP logs in to generate a login event;
[0085] loginevent_create_process: a login event starts a process;
[0086] process_exec_file: Process startup file;
[0087] process_load_file: Process loading file;
[0088] process_exec_scriptfile: Process executing script file;
[0089] process_related_ip: IP related in the process;
[0090] process_related_host: Domain name related in the process;
[0091] file_has_sandboxtag: Sandbox tag marked in the file;
[0092] file_has_md5: File MD5;
[0093] ip_related_vuln: Vulnerability associated with the attacking source IP.
[0094] It should be understood that the above - shown behavior entities and behavior relationships are only exemplary, and this embodiment is not limited thereto. Moreover, for a single attacking source IP, only some of the foregoing behavior entities and behavior relationships may be involved in its corresponding attacking source graph. During the process of constructing the attacking source graph, the graph can be constructed according to the actually extracted behavior entities and behavior relationships, and it is not necessary to involve all the above - mentioned behavior entities and behavior relationships.
[0095] Reference Figure 3 , after completing data splicing, according to the types of the foregoing - set behavior entities and behavior relationships, behavior entities and behavior relationships can be extracted from the behavior record data generated after splicing. Among them, a relatively special behavior entity is the attacking source IP, which is usually the source IP that initiates network behavior mentioned in the previous text. In this way, in this preferred implementation, the attacking source IP can be discovered during the process of extracting behavior entities and the behavior relationships between behavior entities from the behavior record data.
[0096] Continue to refer to Figure 3 , the extracted behavior entities and behavior relationships can be connected again to a streaming execution engine such as Flink in the form of logs respectively, and it will send the behavior entities and behavior relationships to a queue structure such as Swift. In this way, the behavior entities and behavior relationships can be classified and stored in the queue respectively, such as Figure 3 the behavior entity data queue and the behavior relationship queue shown. On this basis, graph construction operations can be performed on the behavior entity data queue and the behavior relationship queue to generate the attacking source graphs corresponding to different attacking source IPs.
[0097] Figure 4 A schematic diagram of an exemplary attack source map provided for an exemplary embodiment of the present application. Refer to Figure 4 , where the points (circular or oval) correspond to behavioral entities, and the edges correspond to behavioral relationships. It can be seen that the attack source map includes behavioral entities - attack source IP, process, and alarm message, and these three behavioral entities are connected by edges. No more explanations will be given here for other behavioral entities and behavioral relationships in the attack source map. It should be understood that in the attack source map, the attack source IP can be directly or indirectly associated with other behavioral entities through edges, and the attack behavior of the attack source IP can be reflected through the behavioral entities and behavioral relationships.
[0098] In addition, refer to Figure 3 , in this preferred implementation, considering the real-time nature of security monitoring data, it is also proposed to transfer the previously extracted behavioral entities and behavioral relationships to a data warehouse such as ODPS through data synchronization attacks such as dataX, so as to trace historical data.
[0099] It should be understood that in this preferred implementation, the attack source map can be constructed in a full-volume plus incremental manner. That is, when initially constructing, the existing full-volume security monitoring data can be used as the basis for map construction to initially generate the attack source map. After that, as more real-time security monitoring data is generated, the operations of extracting behavioral entities and behavioral relationships can be performed based on the incremental security monitoring data, and the incremental behavioral entities and behavioral relationships can be supplemented to the attack source map. In this way, the attack source map can be continuously updated following the dynamic changes of security monitoring data to maintain the accuracy of the attack source map.
[0100] Continue to refer to Figure 1 and Figure 2 , in step 101, the behavioral similarity relationships between each attack source IP can be analyzed according to the respective attack source maps corresponding to each attack source IP. That is, in this embodiment, the behavioral similarity relationships between each attack source IP can be determined by comparing and analyzing the attack source maps. Among them, the behavioral similarity relationship can be understood as the degree of behavioral similarity between different attack source IPs. As mentioned later, the behavioral similarity degree between attack source IPs can be measured by behavioral similarity.
[0101] After step 101, the behavioral similarity relationships between each attack source IP can be obtained. On this basis, refer to Figure 1 and Figure 2The discovery module can construct a behavior similarity map between attack source IPs based on the behavior similarity relationship. In step 101, the attack source IP can be used as an entity, and the behavior similarity between the attack source IPs can be used as the entity relationship to construct a behavior similarity map.
[0102] It can be understood that the behavior similarity graph constructed in this embodiment can present the behavior similarity relationship between the attack source IPs, and the attack source IPs with similar attack behaviors will be associated together.
[0103] Continue to refer Figure 1 and Figure 2 In step 103, the behavior similarity map can be used to perform attack homology analysis. As mentioned above, the attack source IPs with similar attack behaviors in the behavior similarity map will be associated together. In this way, if the attack source IP corresponding to a single attack can be determined, a similar attack source IP can be found from the behavior similarity map as the attack source IP with the same source as the attack. Furthermore, the attack source behind can be further mined based on these homologous attack source IPs, thereby tracing the attack. In actual applications, an attack source may use multiple attack source IPs to launch an attack. Therefore, here, the attack source behind can be understood as the attack source to which the attack source IP belongs.
[0104] In step 103, how to accurately determine the attack source IP for a single attack becomes a problem that needs to be solved. In this regard, it is proposed in this embodiment that after monitoring the alarm information corresponding to a single attack, the attack source IP associated with the alarm information can be determined by querying each attack source map, and used as the attack source IP corresponding to this attack. Referring to the above, in addition to the special behavior entity of the attack source IP, the attack source map constructed in this embodiment also includes another special behavior entity-the alarm message. Of course, referring to Figure 4 , the attack source IP and the alarm information may be directly associated or indirectly associated through other behavior entities. In any case, there is at least one path between the alarm information and the attack source IP. In this way, in this embodiment, the attack source map constructed in step 101 and the behavior similarity map constructed in step 102 can be used together as the basis for attack homology analysis to more efficiently support attack homology analysis.
[0105] The following takes a single attack occurring in the field of network security as an example to illustrate step 103 as follows.
[0106] In the case of detecting the alarm information corresponding to this attack, the attack source graph can be queried first to determine the attack source IP associated with the alarm information as the attack source IP corresponding to this attack. After that, the behavior similarity graph can be queried to determine the attack source IP similar to the attack source IP corresponding to this attack as the result of the attack homology analysis.
[0107] It can be seen that in this embodiment, the attack homology analysis work can be converted into graph query work. Therefore, the efficiency of the attack homology analysis can be effectively improved.
[0108] In summary, in this embodiment, the attack source IP can be discovered from a large amount of security monitoring data, and an attack source graph can be constructed for each discovered attack source IP based on the security monitoring data to mine the attack behavior corresponding to each attack source IP through the attack source graph. On this basis, the behavior similarity relationship between each attack source IP can be analyzed according to the attack source graph corresponding to each attack source IP, and then a behavior similarity graph can be constructed, and the behavior similarity graph can be used as the basis for the attack homology analysis. In this way, when an attack alarm is detected, the attack source IP corresponding to this attack can be determined first, and then, by querying the behavior similarity graph, a similar attack source IP can be found for this attack, and then the attack homology analysis of this attack can be realized. Accordingly, the attack behavior of different attack source IPs can be mined through the attack source graph, and the behavior similarity graph can be constructed more efficiently based on the attack source graph, and then the attack homology analysis can be converted into a graph query operation, which can effectively improve the efficiency of the attack homology analysis.
[0109] In the above or following embodiments, various implementation manners can be adopted to analyze the behavior similarity relationship between each attack source IP according to the attack source graph corresponding to each attack source IP.
[0110] To support more efficient use of the attack source graph constructed in this embodiment, refer to Figure 4 , in this embodiment, graph construction tools such as iGraph can be used to construct a graph index for the attack source graph. In this way, through the graph index, the query operation of the attack source graph in this embodiment can be efficiently performed. In addition, to better support the query operation of the attack source graph, refer to Figure 4 , in this embodiment, in the data warehouse, the entities and entity relationships ( Figure 4 described as point and edge data in
[0111] In this way, in this embodiment, query operations on the attack source graph can be supported.
[0112] On this basis, in an exemplary implementation:
[0113] Behavior chains starting from the attack source IP can be respectively queried from the attack source graphs corresponding to each attack source IP. The behavior chains include behavior entities and the behavior relationships between the behavior entities;
[0114] Based on the behavior chains corresponding to each attack source IP respectively, calculate the behavior similarity between each attack source IP to characterize the behavior similarity relationship.
[0115] First of all, it is worth noting that, as mentioned above, the attack source graph in this embodiment is dynamically updated with the increment of security monitoring data. Therefore, the behavior similarity between each attack source IP can be periodically analyzed in this embodiment to determine whether the behavior similarity graph needs to be updated. Among them, the width of the analysis period can be 1 hour or 1 day, etc., which is not limited here. Through periodic analysis, the accuracy of the behavior similarity graph can be maintained.
[0116] Within a single analysis period, the aforementioned behavior chain query operation can be respectively executed on the attack source graphs corresponding to each attack source IP. Among them, one or more behavior chains may be queried from the attack source graph corresponding to a single attack source IP. In practical applications, all the behavior chains in the attack source graph corresponding to a single attack source IP can be queried to improve the accuracy of the analysis. Refer to Figure 4 , in the attack source graph, the entity relationships between behavior entities are intricate. In this implementation, the attack source IP can be used as the starting point of the behavior chain, and the behavior chain can be queried from the attack source graph for the attack source IP. For example, Figure 4 One behavior chain in can be attack source IP - process - alert information, and another behavior chain can be attack source IP - system login event - alert information. No more enumeration of behavior chains is provided here.
[0117] In this way, in this implementation, the behavior chains corresponding to each attack source IP can be queried. On this basis, based on the behavior chains corresponding to each attack source IP respectively, the behavior similarity between each attack source IP can be calculated to characterize the behavior similarity relationship. That is to say, the similarity degree between the behavior chains corresponding to each attack source IP can be compared to obtain the behavior similarity between each attack source IP.
[0118] In this implementation, multiple solutions can be adopted to compare the similarity degree between the behavior chains corresponding to each attack source IP. Two exemplary solutions are provided below:
[0119] In one solution: the behavior links corresponding to each attack source IP can be converted into behavior description texts respectively; and the similarity between the behavior description texts corresponding to each attack source IP is calculated as the behavior similarity.
[0120] In this solution, the behavior links contained in the graph are converted into behavior description texts, so that the problem of similarity comparison between behavior links can be converted into a text similarity evaluation problem. On this basis, the similarity between behavior description texts can be evaluated based on various existing text similarity evaluation methods, and the similarity generated can be used as the behavior similarity between attack source IPs.
[0121] The following takes the target behavior link corresponding to the target attack source IP as an example to illustrate the process of converting the behavior link into the behavior description text: for the target behavior link corresponding to the target attack source IP, obtain the attribute fields corresponding to each behavior entity and behavior relationship contained in the target behavior link in the security monitoring data; according to the connection order of the behavior entities in the target behavior link, splice the obtained attribute fields to obtain the behavior description text corresponding to the target behavior link. Among them, the behavior entities and behavior relationships in the attack source map are all derived from the security monitoring data. Therefore, each behavior entity and behavior relationship in the target behavior link is associated with a corresponding attribute field in the security monitoring data. In this exemplary conversion process, these attribute fields can be found and spliced in sequence to obtain the behavior description text of the target behavior link. Of course, in the process of splicing the attribute fields, in order to obtain a better splicing effect, more splicing rules can be added, for example, converting some special characters into separators, deduplicating, filtering and keyword extraction of field content, etc., which will not be explained in more detail here. At this point, the behavior description text can achieve the same description effect of the attack behavior as the behavior link. Therefore, by evaluating the similarity of the behavior description text between the attack source IPs, it is equivalent to evaluating the similarity of the attack behaviors between the attack source IPs. The similarity obtained can be used as the behavior similarity between the attack source IPs.
[0122] In another solution: a single behavior link corresponding to a single attack source IP can be transferred to an independent subgraph; a representation vector is calculated for each independent subgraph associated with each attack source IP to obtain a representation vector group corresponding to each attack source IP; and the similarity between the representation vector groups corresponding to each attack source IP is calculated as the behavior similarity.
[0123] In this solution, the essence of the independent sub-graph is still a graph. Thus, it is equivalent to splitting the attack source graph corresponding to the attack source IP into multiple small graphs. On this basis, the similarity evaluation of the respective independent sub-graphs can be performed between two attack source IPs. Under a single attack source IP, one or more independent sub-graphs will be generated. Regarding the sub-graph similarity evaluation, there are already many evaluation methods. In this solution, the method of using the representation vector is adopted for similarity evaluation. That is, according to the preset vector rules, the representation vector graph embedding is calculated for each independent sub-graph respectively. In this way, a single attack source IP will obtain a set of representation vectors, which is described as a representation vector group. Then, the similarity between the respective representation vector groups corresponding to each attack source IP can be calculated. It should be understood that by performing the similarity evaluation of the representation vector groups between the attack source IPs, it is equivalent to performing the similarity evaluation of the attack behaviors between the attack source IPs. The similarity obtained accordingly can be used as the behavior similarity between the attack source IPs.
[0124] It should be understood that the above two solutions are only exemplary. In this implementation manner, other solutions can also be adopted to compare the similarity degree between the respective behavior links corresponding to each attack source IP, and no more examples will be given here.
[0125] In addition, in this embodiment, the implementation manner of analyzing the behavior similarity relationship by querying the behavior link is only exemplary, and this embodiment is not limited thereto. In addition, in this embodiment, other implementation manners can also be adopted to analyze the behavior similarity relationship between each attack source IP according to the attack source graph corresponding to each attack source IP. For example, directly performing the similarity evaluation on the whole attack source graph corresponding to each attack source IP, etc., and no further elaboration will be given here.
[0126] Accordingly, in this embodiment, the already constructed attack source graph can be used as the basis for analysis. By querying the behavior link, the attack behaviors of each attack source IP can be disassembled, making the attack behaviors concrete. Furthermore, performing the similarity analysis on the behavior link is equivalent to performing the similarity analysis on the attack behaviors between the attack source IPs. In this way, through the concrete analysis basis, the accurate similarity analysis of the abstract and concealed attack behaviors can be realized, and the behavior similarity graph can be constructed more efficiently and accurately.
[0127] In the above or following embodiments, it is further proposed to introduce data streams with more dimensions. For example, data disclosed by attack organizations, etc. The data disclosed by attack organizations can be the relevant data of known attack organizations disclosed in public reports, blogs, websites, etc. The data disclosed by attack organizations can include data such as the identifiers of known attack organizations, attack preferences, attack paths, and the attack source IPs used, etc., and no more examples will be given here.
[0128] Based on this, in this embodiment, it is proposed that the attack organization disclosure data and security monitoring data can be spliced; if there is an attack source IP that can be associated with a known attack organization, an attack organization entity associated with the attack source IP is added to the attack source graph corresponding to the attack source IP.
[0129] That is to say, based on the attack organization disclosure data, an attack organization can be introduced as an entity into the attack source graph. Through data splicing, the direct or indirect association relationship between the attack source IP and the attack organization can be mined. Furthermore, for the attack source IP that can be associated with a known attack organization, the corresponding attack organization entity can be added to its attack source graph. Refer to Figure 4 , in this exemplary attack source graph, the attack source IP is associated with an attack organization entity (APT in the figure). Among them, APT refers to Advanced Persistent Threat, and its full English name is Advanced Persistent Thteat. Compared with ordinary attack organizations, the frequency and scope of attacks launched by APT attack organizations have increased significantly, and the advanced attack techniques and ideas of these attack organizations have also played a great role in promoting ordinary network attack sources. Therefore, it is of great security significance to detect the attacks launched by such attack organizations in a timely manner.
[0130] In this way, in this embodiment, the attack source graph corresponding to some attack source IPs will include the entity of the attack organization. Based on this, in this embodiment, in response to an attack organization tracing instruction for a specified attack, the attack source IP corresponding to the specified attack can be determined; from the attack source graph corresponding to the determined attack source IP, the attack organization entity is queried to determine the attack organization to which the specified attack belongs.
[0131] That is to say, in this embodiment, based on the attack source graph, in addition to supporting attack homology analysis, it can further support timely determining the attack organization to which the monitored attack belongs, so that more targeted and effective protection measures can be taken.
[0132] During the research process, the inventor found that according to the existing attack organization disclosure data, many attack source IPs cannot be associated with a consistent attack organization, resulting in a lack of basis for tracing the attack organization for these attack source IPs. For this reason, in this embodiment, it is proposed that Figure 2 the discovery module shown in
[0133] In an exemplary discovery solution: Based on the behavior similarity graph, each attack source IP can be clustered to obtain multiple attack source IP clusters; if there is an attack source IP in the target attack source IP cluster that is already associated with a known attack organization, the attack organization entity corresponding to the known attack organization is added to the attack source graph corresponding to the attack source IPs in the target attack source IP cluster that are not associated with an attack organization.
[0134] In this exemplary direction solution, it is proposed to cluster the attack source IPs based on the behavior similarity graph. As described in the previous text, the behavior similarity graph describes the behavior similarity relationships between each attack source IP. Therefore, through clustering, the attack source IPs with a high enough behavior similarity can be clustered into the same attack source IP cluster. Since the attack source IPs in the same attack source IP cluster have similarity in attack behavior, it is considered that the attack source IPs in the same attack source IP cluster are likely to belong to the same attack organization. Accordingly, in this exemplary discovery solution, the known attack organizations associated with the attack source IPs in the same attack source IP cluster can be passed on to the attack source IPs in the attack source IP cluster. It should be understood that this inheritance can be to associate all the attack source IPs in the same attack source IP cluster with the same attack organization. Of course, the attack organizations associated with the attack source IPs in the same attack source IP cluster may not be exactly the same. For example, if there are two known attack organizations in the same attack source IP cluster, the known attack organization with the most associated times in the attack source IP cluster can be selected, and it is ensured that the attack source graph corresponding to each attack source IP in the attack source IP cluster contains the entity corresponding to the known attack organization; while the other unselected known attack organization can maintain its existing association status and does not need to be added to the attack source graphs of more attack source IPs.
[0135] Figure 5 It is a schematic flowchart of a clustering solution provided by an exemplary embodiment of the present application. Refer to Figure 5 , in this exemplary clustering solution, the first-round clustering can be first performed based on the connection relationship between each attack source IP in the behavior similarity graph to generate at least one group of attack source IPs; refer to Figure 5 , if there is a group with the number of attack source IPs exceeding the set value, then within these groups, the second-round clustering is performed based on the structural relationship of the relevant attack source IPs in the behavior similarity graph to split such groups into at least one attack source IP cluster. It should be understood that the groups that exceed the foregoing set value are automatically used as the foregoing attack source IP clusters. Preferably, refer to Figure 5, in this exemplary clustering scheme, the k-core clustering algorithm or other clustering algorithms can be used in the first round of clustering. In the first round of clustering, unconnected attack source IPs can be divided into different groups to avoid directly using the clustering algorithm in the second round and dividing unconnected attack source IPs into the same group. In the second round of clustering, the louvain clustering algorithm or other clustering algorithms can be used. Parameters involved, such as the modularity shown in Figure 5 , by adjusting these parameters, the specifications of the clustered attack source IP clusters can be adjusted.
[0136] Through Figure 5 the two rounds of clustering shown, the clustered attack source IP clusters can be made more accurate and reasonable.
[0137] The inventor found during the research process that after the aforementioned clustering operation, there are no attack source IPs associated with known attack organizations in some of the attack source IP clusters. Taking such attack source IP clusters as an example, in this exemplary discovery scheme, a custom attack organization identifier can be configured for the target attack source IP; the attack organization entity corresponding to the custom attack organization identifier is added to the attack source map corresponding to the attack source IP included in the target attack source IP cluster.
[0138] In this way, in this embodiment, the problem of discovering unknown attack organizations can be converted into an attack homology analysis problem, that is, during the attack homology analysis process, attack source IP clusters not associated with any known attack organizations are discovered and custom attack organization identifiers are marked for them.
[0139] After the above attack organization discovery process, the attack organization entities included in the attack source map in this embodiment can be made more accurate and comprehensive, so as to better respond to the aforementioned attack organization traceability request and determine the attack organization to which the network attack belongs in a timely manner.
[0140] Referring to Figure 4 , in this embodiment, in addition to configuring attack organization entities in the attack source map, more entities related to attack organizations and related entity relationships can also be configured.
[0141] Several exemplary entities can be:
[0142] group: Custom attack organization. It is mainly the gang name and gang description calculated by the algorithm and verified;
[0143] apt: Known attack organization. APT organization information and related IOC information obtained through public reports, blogs, websites, etc.;
[0144] url: url extracted from the attack path;
[0145] host: The host, which is the host extracted from the attack path.
[0146] Several exemplary entity relationships can be:
[0147] apt_related_url: The URL associated with the attacking organization;
[0148] apt_related_ip: The IP associated with the attacking organization;
[0149] apt_related_host: The domain name associated with the attacking organization;
[0150] apt_related_hash: The file hash associated with the attacking organization;
[0151] apt_related_cve: The publicly disclosed vulnerability associated with the attacking organization;
[0152] group_belong_ip: The general term for a class of similar attacking source IPs.
[0153] In this way, in the attacking source graph, after finding the attacking organization, further understanding and analysis of the attacking organization can be carried out based on other entities and entity relationships related to the attacking organization.
[0154] Reference Figure 4 , in this embodiment, in addition to adding entities and entity relationships related to the attacking organization to the attacking source graph, more types of entities and entity relationships can be added to the attacking source graph to better support network security. Reference Figure 4 , an entity can also be added to the attacking source graph: incident: The general term for a class of alarm information. In this embodiment, the general term for the alarm can be generated through aggregation operations such as graph association, security rules, and aggregation of similar attacking behaviors. Correspondingly, an entity relationship corresponding to this can also be added to the attacking source graph: alert_belong_incident: The general term for the alarm information to which the alarm message belongs, to characterize which general term for the alarm the alarm information belongs to. In this way, in the case of monitoring the alarm information, the general term for the alarm to which the alarm information belongs can be determined by querying the attacking source graph, providing more valuable reference for security personnel. Continuing to refer to Figure 4 , by introducing vulnerability disclosure data CVE and threat intelligence data, and the type of data introduction method for the attacking organization's disclosure data, through data splicing, entities can be added to the attacking source graph:
[0155] cve: The identifier of the common vulnerability disclosure;
[0156] vlun: The common name of the vulnerabilities that are widely recognized and exposed in the vulnerabilities.
[0157] Accordingly, entity relationships can also be added to the attack source graph:
[0158] cve_belong_vuln: The specific vulnerability / baseline / weakness to which the CVE belongs.
[0159] It should be understood that the entities and entity relationships added to the attack source graph in this embodiment are exemplary, and this embodiment is not limited thereto. In this embodiment, more types of entities and entity relationships can also be added to the attack source graph to provide more support for network security work. No more examples are given here.
[0160] In addition, referring to Figure 4 , the above-introduced attack organization disclosure data, vulnerability disclosure data, threat intelligence data, etc. are basically static data. Therefore, in this embodiment, these static data can be transferred and stored in a data warehouse such as ODPS through data synchronization.
[0161] In summary, in this embodiment, by transforming the attack source graph, the following optimized technical effects can be obtained:
[0162] 1. It is possible to discover new behaviors belonging to known attack organizations and predict the subsequent attack methods of known attack organizations based on the discovered new behaviors.
[0163] 2. It is possible to separate organized behaviors and unorganized behaviors from a large number of attack alerts, and further analyze known attack organizations and potential attack organizations, providing a reference basis for subsequent traceability work.
[0164] 3. It is possible to provide homologous analysis data from multiple dimensions as a reference basis for network defense, thereby improving the pertinence of network defense work.
[0165] It should be noted that in some of the processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this article or in parallel. The operation numbers such as 101 and 102 are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in order or in parallel.
[0166] Figure 6 This is a schematic structural diagram of a computing device provided in another exemplary embodiment of the present application. As Figure 6 shown, the computing device includes: a memory 60, a processor 61, and a communication component 62.
[0167] A processor 61, coupled to a memory 60, is configured to execute a computer program in the memory 60 for:
[0168] Based on security monitoring data, construct an attack source map for each discovered attack source IP to describe the attack behavior;
[0169] According to the attack source maps corresponding to the respective attack source IPs, analyze the behavioral similarity relationships among the attack source IPs;
[0170] Based on the behavioral similarity relationships, construct a behavioral similarity map among the attack source IPs;
[0171] Utilize the behavioral similarity map for attack homology analysis.
[0172] In an alternative embodiment, when the processor 61 analyzes the behavioral similarity relationships among the attack source IPs according to the attack source maps corresponding to the respective attack source IPs, it is specifically configured to:
[0173] Query, from the attack source maps corresponding to the respective attack source IPs, the behavioral links starting from the attack source IPs, where the behavioral links include behavioral entities and the behavioral relationships between the behavioral entities;
[0174] Based on the behavioral links corresponding to the respective attack source IPs, calculate the behavioral similarity among the attack source IPs to characterize the behavioral similarity relationships.
[0175] In an alternative embodiment, when the processor 61 calculates the behavioral similarity among the attack source IPs based on the behavioral links corresponding to the respective attack source IPs, it is specifically configured to:
[0176] Convert the behavioral links corresponding to the respective attack source IPs into behavioral description texts respectively;
[0177] Calculate the similarity among the behavioral description texts corresponding to the respective attack source IPs as the behavioral similarity.
[0178] In an alternative embodiment, when the processor 61 converts the behavioral links corresponding to the respective attack source IPs into behavioral description texts respectively, it is specifically configured to:
[0179] For the target behavioral link corresponding to the target attack source IP, obtain the attribute fields corresponding to each behavioral entity and behavioral relationship included in the target behavioral link in the security monitoring data;
[0180] Concatenate the obtained attribute fields in the connection order of the behavioral entities in the target behavioral link to obtain the behavioral description text corresponding to the target behavioral link.
[0181] In an alternative embodiment, when calculating the behavior similarity between each attacking source IP based on the respective behavior links corresponding to each attacking source IP, the processor 61 can be specifically used for:
[0182] Transfer a single behavior link corresponding to a single attacking source IP into an independent subgraph;
[0183] Calculate representation vectors for each independent subgraph associated with each attacking source IP to obtain a group of representation vectors corresponding to each attacking source IP;
[0184] Calculate the similarity between the groups of representation vectors corresponding to each attacking source IP as the behavior similarity.
[0185] In an alternative embodiment, when constructing an attacking source map for each discovered attacking source IP based on security monitoring data to describe the attacking behavior, the processor 61 can be specifically used for:
[0186] Perform data connection on various data streams included in the security monitoring data to obtain several pieces of behavior record data;
[0187] Extract behavior entities and the behavior relationships between the behavior entities from the behavior record data, where the attacking source IP is one of the behavior entities;
[0188] Based on the extracted behavior entities and the behavior relationships between the behavior entities, perform map construction to obtain the attacking source maps corresponding to each attacking source IP.
[0189] In an alternative embodiment, the processor 61 can also be used for:
[0190] Transfer the extracted behavior entities and the behavior relationships between the behavior entities into a data warehouse for data traceability.
[0191] In an alternative embodiment, the processor 61 can also be used for:
[0192] Perform data splicing on the attack organization disclosure data and the security monitoring data;
[0193] If there is an attacking source IP that can be associated with a known attack organization, add an attack organization entity associated with the attacking source IP to the attacking source map corresponding to the attacking source IP.
[0194] In an alternative embodiment, the processor 61 can also be used for:
[0195] Cluster the attacking source IPs based on the behavior similarity map to obtain multiple attacking source IP clusters;
[0196] If there is an attack source IP in the target attack source IP cluster that is already associated with a known attack organization, add the attack organization entity corresponding to the known attack organization to the attack source map corresponding to the attack source IPs in the target attack source IP cluster that are not associated with other attack organizations.
[0197] In an alternative embodiment, the processor 61 may further be configured to:
[0198] If there is no attack source IP in the target attack source IP cluster that is already associated with a known attack organization, configure a custom attack organization identifier for the target attack source IP;
[0199] Add the attack organization entity corresponding to the custom attack organization identifier to the attack source map corresponding to the attack source IPs included in the target attack source IP cluster.
[0200] In an alternative embodiment, the processor 61 may further be configured to:
[0201] In response to an attack organization traceability instruction for a specified attack, determine the attack source IP corresponding to the specified attack;
[0202] Query the attack organization entity from the attack source map corresponding to the determined attack source IP to determine the attack organization to which the specified attack belongs.
[0203] In an alternative embodiment, the security monitoring data includes one or more data streams such as alarm data and network log data.
[0204] Further, as Figure 6 shown, the computing device further includes: a power supply component 63 and other components. Figure 6 Only some components are schematically shown, and it does not mean that the computing device only includes Figure 6 the components shown.
[0205] It should be noted that for the technical details in the above embodiments of the computing device, reference may be made to the relevant descriptions in the foregoing method embodiments. To save space, they are not repeated here, but this should not cause loss of the protection scope of this application.
[0206] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement the steps executable by the computing device in the above method embodiments.
[0207] The above Figure 6The memory therein is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks or optical discs.
[0208] The above-mentioned Figure 6 The communication component therein is configured to facilitate communication in a wired or wireless manner between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0209] The above-mentioned Figure 6 The power component therein provides power to various components of the device where the power component is located. The power component can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device where the power component is located.
[0210] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0211] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.
[0212] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or the functions specified in multiple blocks.
[0213] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or the functions specified in multiple blocks.
[0214] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or elements inherent to such a process, method, commodity, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity, or device including the said element.
[0215] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0216] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. An attack homology analysis method, characterized in that, Including: Based on security monitoring data, construct an attack source graph for each discovered attack source IP to describe the attack behavior; According to the attack source graphs corresponding to each attack source IP, analyze the behavior similarity relationship between the attack source IPs; Based on the behavior similarity relationship, construct a behavior similarity graph between the attack source IPs; Use the behavior similarity graph to perform attack homology analysis.
2. The method according to claim 1, wherein According to the attack source graphs corresponding to each attack source IP, analyze the behavior similarity relationship between the attack source IPs, including: From the attack source graphs corresponding to the attack source IPs, respectively query the behavior links starting from the attack source IP, where the behavior links contain behavior entities and the behavior relationships between the behavior entities; Based on the behavior links corresponding to each attack source IP, calculate the behavior similarity between the attack source IPs to characterize the behavior similarity relationship.
3. The method according to claim 2, wherein Based on the behavior links corresponding to each attack source IP, calculate the behavior similarity between the attack source IPs, including: Convert the behavior links corresponding to each attack source IP into behavior description texts respectively; Calculate the similarity between the behavior description texts corresponding to each attack source IP as the behavior similarity.
4. The method according to claim 3, wherein Convert the behavior links corresponding to each attack source IP into behavior description texts respectively, including: For the target behavior link corresponding to the target attack source IP, obtain the attribute fields corresponding to each behavior entity and behavior relationship included in the target behavior link in the security monitoring data; According to the connection order of the behavior entities in the target behavior link, splice the obtained attribute fields to obtain the behavior description text corresponding to the target behavior link.
5. The method according to claim 2, wherein Based on the behavior links corresponding to each attack source IP, calculate the behavior similarity between the attack source IPs, including: Convert a single behavior link corresponding to a single attack source IP into an independent sub-graph; Calculate representation vectors for each independent sub-graph associated with each attack source IP to obtain a representation vector group corresponding to each attack source IP; Calculate the similarity between the representation vector groups corresponding to each attack source IP as the behavior similarity.
6. The method according to claim 1, wherein Based on security monitoring data, construct an attack source graph for each discovered attack source IP to describe the attack behavior, including: Perform data connection on multiple data streams included in the security monitoring data to obtain several behavior record data; Extract behavior entities and the behavior relationships between the behavior entities from the behavior record data, where the attack source IP is one of the behavior entities; Based on the extracted behavior entities and the behavior relationships between the behavior entities, perform graph construction to obtain the attack source graphs corresponding to each attack source IP.
7. The method according to claim 6, characterized in that, Also including: Transfer the extracted behavior entities and the behavior relationships between the behavior entities to a data warehouse for data traceability.
8. The method according to claim 1, characterized in that, Also including: Perform data splicing on the attack organization disclosure data and the security monitoring data; If there is an attack source IP that can be associated with a known attack organization, an attack organization entity associated with the attack source IP is added to the attack source graph corresponding to the attack source IP.
9. The method according to claim 8, characterized in that, It further includes: Based on the behavior similarity graph, clustering the respective attack source IPs to obtain multiple attack source IP clusters; If there is an attack source IP in the target attack source IP cluster that has been associated with a known attack organization, the attack organization entity corresponding to the known attack organization is added to the attack source graphs corresponding to the attack source IPs in the target attack source IP cluster that have not been associated with an attack organization.
10. The method according to claim 9, characterized in that, Based on the behavior similarity graph, clustering the respective attack source IPs to obtain multiple attack source IP clusters, including: Based on the connectivity relationship between the respective attack source IPs in the behavior similarity graph, performing a first round of clustering to generate at least one group of attack source IPs; If there is a group in which the number of attack source IPs exceeds a set value, performing a second round of clustering within the group based on the structural relationship of the relevant attack source IPs in the behavior similarity graph to split the group into at least one attack source IP cluster.
11. The method according to claim 9, wherein It further includes: If there is no attack source IP in the target attack source IP cluster that has been associated with a known attack organization, configuring a custom attack organization identifier for the target attack source IP; Adding the attack organization entity corresponding to the custom attack organization identifier to the attack source graphs corresponding to the attack source IPs included in the target attack source IP cluster.
12. The method according to any one of claims 8-11, characterized in that, It further includes: In response to an attack organization traceability instruction for a specified attack, determining the attack source IP corresponding to the specified attack; Querying the attack organization entity from the attack source graph corresponding to the determined attack source IP to determine the attack organization to which the specified attack belongs.
13. A computing device, characterized in that, It includes a memory, a processor, and a communication component; The memory is used to store one or more computer instructions; The processor is coupled to the memory and the communication component and is used to execute the one or more computer instructions to execute the attack homology analysis method according to any one of claims 1-12.
14. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are executed by one or more processors, causing the one or more processors to execute the attack homology analysis method according to any one of claims 1-12.