A threat path tracing method, device, equipment and medium

By building an IP address association diagram based on security logs, the problem that the firewall cannot restore the network attack path is solved, and in-depth analysis and timely processing of threat traffic paths are achieved to ensure network security.

CN119109679BActive Publication Date: 2025-07-04CITIC TELECOM INTERNATIONAL CPC LIMITED +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411303783.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-18
Publication Date
2025-07-04
Estimated Expiration
2044-09-18

AI Technical Summary

Technical Problem

In the prior art, the firewall cannot effectively restore the network attack path after blocking threat traffic, resulting in the inability of professional and technical personnel to analyze the attack path of threat traffic in depth.

Method used

By obtaining the security log information of the network security system, generating an IP address collection, determining dangerous IP addresses and their associated addresses, building a list of associated IP addresses, determining frequent IP addresses and their associated relationships, and generating node association diagrams to characterize the traceability of threat traffic paths.

Benefits of technology

It realizes an intuitive understanding of the attack path of threatened traffic, helps network security personnel to handle it in a timely manner, and ensures the security of network assets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119109679B_ABST
    Figure CN119109679B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of network security technology, and discloses a threat path tracing method, device, equipment and medium. The method includes: obtaining security log information in a preset historical period, including source IP addresses and destination IP addresses corresponding to multiple threat traffic data; generating a first IP address set based on the source IP addresses and destination IP addresses corresponding to each threat traffic data, and determining multiple dangerous IP addresses and their respective associated IP addresses to generate an associated IP address list; determining frequent IP addresses in the associated IP address list and the association relationships between the respective frequent IP addresses; generating a node association graph based on the association relationships between the multiple frequent IP addresses, and the node association graph is used to represent the path tracing of threat traffic. The present invention generates a node association graph that can trace the path of threat traffic by performing correlation analysis on the IP addresses of threat traffic, so as to facilitate security personnel to analyze the attack path of threat traffic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and particularly relates to a threat path tracing method, device, equipment and medium. Background Art

[0002] With the development of network technology, the network is becoming increasingly important for daily life and industrial production. At the same time, the attack paths of network threat traffic are becoming more and more complex. In order to better prevent dangerous traffic, it has become increasingly important to trace the access paths of dangerous traffic.

[0003] In the prior art, after intercepting threat traffic through a firewall and obtaining the log information corresponding to the intercepted threat traffic data, only the address information is often simply recorded, and the restoration of the network attack path cannot be achieved, resulting in professional technicians being unable to deeply analyze the attack paths of threat traffic. Summary of the Invention

[0004] In view of this, the present invention provides a threat path tracing method, device, equipment and medium to solve the problem that the tracing of the network attack path cannot be achieved, resulting in technicians being unable to deeply analyze threat traffic.

[0005] In a first aspect, the present invention provides a threat path tracing method, and the method includes:

[0006] Obtain the security log information of the network security system in a preset historical period, where the security log information includes: the source IP addresses and destination IP addresses corresponding to multiple threat traffic data;

[0007] Generate a first IP address set based on the source IP addresses and destination IP addresses corresponding to each threat traffic data, and determine multiple dangerous IP addresses in the first IP address set;

[0008] According to the security log information, determine the associated IP addresses corresponding to each of the multiple dangerous IP addresses, and generate an associated IP address list;

[0009] Determine the frequent IP addresses in the associated IP address list and the association relationships between the frequent IP addresses;

[0010] Generate a node association graph based on the association relationships between the multiple frequent IP addresses, and the node association graph is used to represent the path tracing of threat traffic.

[0011] This method generates a first set of IP addresses based on the source IP addresses and destination IP addresses corresponding to the threat traffic data recorded in the security logs, determines the dangerous IP addresses among them, generates a list of associated IP addresses based on the associated IP addresses corresponding to each dangerous IP, and finally determines the frequently occurring IP addresses and the association relationships between these frequently occurring IP addresses, thereby finally generating a node association graph that can represent the traceability of the threat traffic path, so that the staff can more intuitively understand the attack path of the threat traffic, facilitating timely pre-processing by relevant network security personnel and ensuring the security of network assets.

[0012] In an alternative embodiment, determining the associated IP addresses corresponding to each of the multiple dangerous IP addresses according to the security log information and generating a list of associated IP addresses includes:

[0013] According to the security log information, determine the associated destination IP addresses when each dangerous IP address is the source IP address, and the associated source IP addresses when each dangerous IP address is the destination IP address;

[0014] Generate a corresponding first sub-table according to each dangerous IP address and the corresponding associated destination IP address, and generate a corresponding second sub-table according to each dangerous IP address and the corresponding associated source IP address;

[0015] Construct a list of associated IP addresses according to the first sub-table and the second sub-table corresponding to each dangerous IP address.

[0016] In this embodiment, by generating the first sub-table and the second sub-table corresponding to each dangerous IP address according to the IP addresses when each dangerous IP address is the destination IP address and the source IP address respectively, a list of associated IP addresses including multiple sub-tables can be generated to ensure the accuracy of the finally generated list of associated IP addresses, facilitating the subsequent accurate determination of the frequently occurring IP addresses and the corresponding association relationships.

[0017] In an alternative embodiment, determining the multiple dangerous IP addresses in the first set of IP addresses includes:

[0018] Determine the public information corresponding to each IP address in the first set of IP addresses;

[0019] Based on the public information corresponding to each IP address, determine the multiple dangerous IP addresses in the first set of IP addresses.

[0020] In this embodiment, by using the public information of each IP address in the first set of IP addresses to determine the dangerous IP addresses, the dangerous IP addresses with risks can be determined more accurately, ensuring the accuracy of the dangerous address judgment.

[0021] In an optional implementation, the determining the frequent IP addresses in the associated IP address list and the association relationship between the frequent IP addresses includes:

[0022] By using a preset association rule mining algorithm, the frequent IP addresses in the associated IP address list and the corresponding association confidences between the frequent IP addresses in different association directions are determined;

[0023] The association relationship between the frequent IP addresses is determined according to the association direction and the corresponding association confidence.

[0024] In this implementation, a preset association rule mining algorithm is used to determine frequent IP addresses and the association confidence between each frequent IP address in different association directions, thereby more accurately determining the association relationship between frequent IP addresses and ensuring the accuracy of the subsequently generated node association graph.

[0025] In an optional implementation, before generating a node association graph based on the association relationship between the multiple frequent IP addresses, the method further includes:

[0026] An association relationship whose association confidence is less than a preset value is defined as a low association relationship, and the low association relationship between the frequent IP addresses is released.

[0027] This implementation method eliminates associations with low association confidence levels, thereby ensuring that the node association graph determined subsequently can better characterize the path tracing of threat traffic.

[0028] In an optional implementation, generating a node association graph based on the association relationship between the multiple frequent IP addresses includes:

[0029] Generate a node corresponding to each frequent IP address, and generate a directed edge between the corresponding nodes based on the association direction and the corresponding association confidence between the frequent IP addresses;

[0030] A node association graph is generated based on the nodes corresponding to the frequent IP addresses and the directed edges between the nodes.

[0031] In this implementation, directed edges are generated through the corresponding association directions and association confidences between frequent IP addresses, thereby connecting the nodes corresponding to the frequent IP addresses, thereby generating a node association graph that accurately represents the association relationship and ensuring the display effect of the node association graph.

[0032] In an optional embodiment, the method further includes:

[0033] By means of a preset link analysis algorithm, determine the influence weights of the nodes corresponding to each frequent IP address in the node association graph; the influence weights are used to characterize the participation degree of the corresponding frequent IP address in the threat traffic behavior.

[0034] In this embodiment, by determining the influence weights of the nodes corresponding to each frequent IP address, relatively important nodes in the node association graph can be determined, so that the staff can understand the participation degree of each frequent IP address in the threat traffic behavior, facilitating corresponding processing in a timely manner and achieving effective network protection.

[0035] In a second aspect, the present invention provides a threat path tracing device, which includes:

[0036] A log information acquisition module, configured to acquire security log information of a network security system in a preset historical period, where the security log information includes source IP addresses and destination IP addresses corresponding to multiple threat traffic data;

[0037] A dangerous address determination module, configured to generate a first IP address set based on the source IP addresses and destination IP addresses corresponding to each threat traffic data, and determine multiple dangerous IP addresses in the first IP address set;

[0038] An address list generation module, configured to determine associated IP addresses corresponding to each of the multiple dangerous IP addresses according to the security log information, and generate an associated IP address list;

[0039] An association relationship determination module, configured to determine frequent IP addresses in the associated IP address list and the association relationships between the frequent IP addresses

[0040] An association graph generation module, configured to generate a node association graph based on the association relationships between the multiple frequent IP addresses, where the node association graph is used to represent the path tracing of threat traffic. In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other, and a computer instruction is stored in the memory. The processor executes the computer instruction to execute the threat path tracing method according to the first aspect or any corresponding embodiment thereof.

[0041] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer instruction is stored, and the computer instruction is used to cause a computer to execute the threat path tracing method according to the first aspect or any corresponding embodiment thereof. Description of the Drawings

[0042] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0043] Figure 1 is a schematic flowchart of a threat path tracing method according to an embodiment of the present invention;

[0044] Figure 2 is a schematic flowchart of another threat path tracing method according to an embodiment of the present invention;

[0045] Figure 3 is a schematic diagram of node in-links and out-links according to an embodiment of the present invention;

[0046] Figure 4 is a schematic flowchart of calculating the PR value based on the transition probability matrix according to an embodiment of the present invention;

[0047] Figure 5 is a structural block diagram of a threat path tracing device according to an embodiment of the present invention;

[0048] Figure 6 is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Specific Embodiments

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0050] With the development of network technology, the network is becoming more and more important for daily life and industrial production. At the same time, the attack paths of network threat traffic are becoming more and more complex. To better prevent dangerous traffic, it is becoming more and more important to trace the access paths of dangerous traffic.

[0051] In the prior art, after intercepting threat traffic through a firewall and obtaining the log information corresponding to the intercepted threat traffic data, only the address information is often simply recorded, and the restoration of the network attack path cannot be achieved, resulting in professional technicians being unable to deeply analyze the attack paths of threat traffic.

[0052] To this end, an embodiment of the present invention provides a threat path tracing method. By using the source IP address and destination IP address corresponding to the threat traffic data recorded in the security log, a first set of IP addresses is generated, and the dangerous IP addresses therein are determined. Then, an associated IP address list is generated based on the associated IP addresses corresponding to each dangerous IP. Finally, the frequent IP addresses and the association relationships between these frequent IP addresses are determined, and a node association graph that can represent the threat traffic path tracing is ultimately generated, so that the staff can more intuitively understand the attack path of the threat traffic, facilitating timely pre-processing by relevant network security personnel and ensuring the security of network assets.

[0053] According to an embodiment of the present invention, an embodiment of a threat path tracing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0054] In this embodiment, a threat path tracing method is provided, which can be used in the above-mentioned network security protection process. Figure 1 It is a flowchart of a threat path tracing method according to an embodiment of the present invention, as Figure 1 shown. The process includes the following steps:

[0055] Step S101, obtain the security log information of the network security system in a preset historical period. The security log information includes: the source IP address and destination IP address corresponding to multiple threat traffic data.

[0056] It can be understood that the network security system can intercept threat traffic data, which includes the source IP address corresponding to the sender of the traffic and the IP address corresponding to the receiver of the traffic data. Usually, the network security system stores the specific information corresponding to the intercepted traffic data in the security log. Specifically, it can be stored in the UTM security log. In the security log, information such as the source IP address, destination IP address, threat occurrence time, threat type, source port, destination port, threat level, and execution action corresponding to each threat traffic is recorded. Network security staff can understand the specific threat traffic information in the security log and can also call the traffic information therein to analyze the threat traffic, that is, obtain the security log information within a preset historical period. Specifically, it can be the security log information in the most recent week or the most recent month. Exemplarily, the specific format of the dangerous traffic data in the security log information can be: threat traffic data 1, source IP address A, destination IP address B, and so on. Each threat traffic data contains its corresponding source IP address and destination IP address.

[0057] Step S102: Generate a first IP address set based on the source IP addresses and destination IP addresses corresponding to each threat traffic data, and determine multiple dangerous IP addresses in the first IP address set.

[0058] After obtaining the security log information for a preset historical period, count the source IP addresses and destination IP addresses corresponding to each threat traffic data in the security log, and remove duplicates from the repeatedly occurring IP addresses after statistics to obtain an IP address set, that is, the first IP address set. For example, for threat traffic data 1, the source IP address is A and the destination IP address is B; for threat traffic data 2, the source IP address is A and the destination IP address is C; for threat traffic data 3, the source IP address is D and the destination IP address is A. Taking these three threat traffic data as an example, the IP addresses in the corresponding first IP address set are {A, B, C, D}. Among these IP addresses, multiple dangerous IP addresses can be determined through preset determination rules. For example, determine dangerous IP addresses based on the prefix of the IP address, or determine dangerous IP addresses according to the historical records used to record dangerous IP addresses in the network security system, or determine multiple dangerous IP addresses according to other preset rules. The specific determination method of dangerous IP addresses can be set according to the actual situation or the experience of the staff, and is not limited here.

[0059] Step S103: According to the security log information, determine the associated IP addresses corresponding to each of the multiple dangerous IP addresses, and generate an associated IP address list.

[0060] After determining the dangerous IP addresses, since the security log information records the source IP addresses and threat IP addresses of each threat traffic data, the associated IP addresses corresponding to each threat IP address can be determined according to the corresponding relationship between these source IP addresses and threat IP addresses, thereby generating an associated IP address list.

[0061] For example, the following multiple threat traffic data are recorded in the security log information: threat traffic data 1, source IP address A, destination IP address B; threat traffic data 2, source IP address A, destination IP address C; threat traffic data 3, source IP address D, destination IP address A. Assuming the dangerous IP address is A, its corresponding associated IP addresses are B, C, and D. Among them, B and C are the destination IP addresses corresponding to A when A is the source IP address, and D is the source IP address corresponding to A when A is the destination IP address. These IP addresses and A can be jointly formed into a sub-list. By analogy, obtain the sub-lists corresponding to each of the multiple IP addresses, and thus jointly form the associated IP address list from these sub-lists.

[0062] Specifically, the sub-lists in the associated IP address list can be divided according to whether the dangerous IP address is the source IP address or the destination IP address, and multiple sub-tables corresponding to each dangerous IP address are obtained. Taking A as an example, when A is the source IP address, the corresponding sub-table 1 is B and C, and when A is the source IP address, the corresponding sub-table 2 is D, and so on, which will not be elaborated here.

[0063] Step S104, determine the frequent IP addresses in the associated IP address list and the association relationships between the respective frequent IP addresses.

[0064] After determining the associated IP address list, since it includes multiple sub-tables corresponding to different dangerous IP addresses, the frequent IP addresses therein and the association relationships between the frequent IP addresses can be determined according to the occurrence frequencies of different IP addresses in different sub-tables and the association when the same sub-table appears simultaneously. For example, a preset association analysis algorithm can be used to determine the specific frequent IP addresses and the associations between the frequent IP addresses.

[0065] For example, assuming that the determined frequent IP addresses are A and B, the possibility of the occurrence of IP address B when IP address A appears, or the possibility of the occurrence of IP address B when IP address B appears, for each sub-table in the associated IP address list, can be determined through the association analysis method, which is the association between IP address A and IP address B.

[0066] Specifically, the Apriori algorithm can be used to determine the specific frequent IP addresses and the associations between the specific frequent IP addresses.

[0067] Step S105, generate a node association graph based on the association relationships between multiple frequent IP addresses, and the node association graph is used to represent the path tracing of threat traffic.

[0068] After determining the frequent IP addresses in the associated IP address list, these frequent IP addresses can be determined as nodes, and the frequent IP addresses with relatively strong association relationships between these frequent IP addresses are connected. Specifically, for two frequent IP addresses, the association relationship between the two has a directionality. Taking frequent IP address A and frequent IP address B as an example, there is an association of A relative to B and an association of B relative to A between them. Therefore, when connecting the frequent IP address A, a directed edge can be used for connection, and these directed edges can represent the traffic paths of threat traffic. The stronger the association relationship, the more likely the corresponding directed edge is the path of real threat traffic.

[0069] The threat path tracing method provided in this embodiment generates a first set of IP addresses based on the source IP address and destination IP address corresponding to the threat traffic data recorded in the security log, and determines the dangerous IP addresses among them. Then, an associated IP address list is generated based on the associated IP addresses corresponding to each dangerous IP, and finally, the frequent IP addresses and the association relationships between these frequent IP addresses are determined. Thus, a node association graph that can represent the threat traffic path tracing is finally generated, enabling the staff to more intuitively understand the attack path of the threat traffic, facilitating the relevant network security personnel to carry out timely pre-processing and ensuring the security of network assets.

[0070] According to an embodiment of the present invention, another embodiment of the threat path tracing method is provided, which can be used in the above network security protection process. Figure 2 It is a flowchart of another threat path tracing method according to an embodiment of the present invention, as Figure 2 shown. The process includes the following steps:

[0071] Step S201, obtain the security log information of the network security system in a preset historical period. The security log information includes: the source IP address and destination IP address corresponding to multiple threat traffic data. For the specific implementation, refer to Figure 1 Step S101 of the embodiment shown, which will not be elaborated here.

[0072] Step S202, based on the source IP address and destination IP address corresponding to each threat traffic data, generate a first set of IP addresses, and determine multiple dangerous IP addresses in the first set of IP addresses.

[0073] Specifically, in step S202, determining multiple dangerous IP addresses in the first set of IP addresses includes:

[0074] Determine the public information corresponding to each IP address in the first set of IP addresses;

[0075] Based on the public information corresponding to each IP address, determine multiple dangerous IP addresses in the first set of IP addresses.

[0076] It can be understood that after obtaining the first set of IP addresses, each IP address therein can be understood as a website address. For these website addresses, their corresponding public information can be found. The public information of the website addresses can include whether there are dangerous behaviors in their past access behaviors, or whether there are dangerous information on their websites, etc. The staff can label each IP address in the first set of IP addresses through the public information of these website addresses, and mark the dangerous IP addresses with a dangerous label for annotation.

[0077] Step S203: Determine the associated IP addresses corresponding to each of the multiple dangerous IP addresses according to the security log information, and generate a list of associated IP addresses.

[0078] Specifically, in the above step S203, it includes:

[0079] Step S2031: According to the security log information, determine the corresponding associated destination IP addresses when each dangerous IP address is the source IP address, and the corresponding associated source IP addresses when each dangerous IP address is the destination IP address.

[0080] Exemplarily, assume that the threat traffic data recorded in the security log information is as shown in Table 1 below:

[0081] Table 1

[0082] Threat traffic data Source IP address Destination IP address 1 A B 2 A C 3 D B 4 E A 5 C D 6 F C 7 G D

[0083] Through such security log information, the first set of IP addresses obtained is {A, B, C, E, F, G}. Assume that the dangerous IP addresses among them are A, C, and D. Then, taking A as an example, the associated destination IP addresses corresponding to it when it is the source IP address are B and C; when it is the destination IP address, the associated source IP address corresponding to it is E. By analogy, the determination processes of the associated source IP addresses and associated destination IP addresses of other dangerous IP addresses are not elaborated here.

[0084] Step S2032: Generate a corresponding first sub-table according to each dangerous IP address and the corresponding associated destination IP address, and generate a corresponding second sub-table according to each dangerous IP address and the corresponding associated source IP address.

[0085] Still taking the dangerous IP address A in the above example as an example, the associated destination IP addresses corresponding to it are B and C. Therefore, the IP addresses in the corresponding first sub-table are [A, B, C]. The associated destination IP address corresponding to A is E. Therefore, the IP addresses in the corresponding first sub-table are [A, E]. In some cases, after obtaining multiple sub-tables corresponding to multiple dangerous IP addresses, if there are exactly the same sub-tables among these sub-tables, deduplication processing can be performed to avoid the existence of exactly the same sub-table information, which may affect the determination of subsequent frequent IP addresses and association relationships. In some cases, if there is no situation where a certain dangerous IP address is the source IP address or the destination IP address, there is no need to determine the corresponding first sub-table or second sub-table.

[0086] Step S2033: Construct a list of associated IP addresses according to the first sub-table and the second sub-table corresponding to each dangerous IP address.

[0087] After obtaining the first sub-table and the second sub-table corresponding to each dangerous IP address, these sub-tables can be combined to obtain an associated IP address list. Taking the dangerous IP addresses mentioned in the above example as an example, the first sub-table corresponding to the dangerous IP address A is: [A, B, C], and the second sub-table is: [A, E]; the first sub-table corresponding to the dangerous IP address C is: [C, D], and the second sub-table is: [A, C, F]; for the dangerous IP address D, the corresponding first sub-table is: [D, B], so there is no first sub-table, and the second sub-table corresponding to the dangerous IP address is: [D, C, G].

[0088] Combining the first sub-tables and the second sub-tables of these different dangerous IP addresses together constitutes an associated IP address list. Exemplarily, the associated IP address list can be as shown in Table 2 below:

[0089] Table 2

[0090] IP address table IP address 1 A, B, C 2 A, E 3 C, D 4 A, C, F 5 D, B 6 D, C, G

[0091] Step S204, determine the frequent IP addresses in the associated IP address list and the association relationships between each frequent IP address.

[0092] Specifically, in the above step S204, it includes:

[0093] Step S2041, through a preset association rule mining algorithm, determine the frequent IP addresses in the associated IP address list and the association confidence corresponding to each frequent IP address in different association directions.

[0094] When determining the frequent IP addresses and specific association relationships, existing association rule mining algorithms can be used to associate the frequently occurring IP addresses and specific association relationships in the IP address list. Exemplarily, the Apriori algorithm can be used for association analysis.

[0095] During the process of determining the association relationship, it is first necessary to determine the frequent IP addresses. Suppose the specific information of the associated IP address list is as shown in Table 3 below:

[0096] Table 3

[0097] IP address table IP address 1 1.1,1.2,1.3,1.4,1.5 2 1.2,1.3,1.4,1.6 3 1.1,1.3,1.5,1.7 4 1.2,1.4,1.5,1.8 5 1.1,1.3,1.4,1.5,1.9

[0098] Among them, 1.1, 1.2, 1.3... in the IP address can be regarded as abbreviations of the IP address. For example, 1.1 can represent 192.168.1.1, which is only used for auxiliary understanding here. When determining frequent IP addresses, it is first necessary to determine the support of each IP address in each table and the preset support threshold. Support refers to an item set, that is, the frequency of a certain IP address appearing in all tables. For example, if the item set {1.1} appears 3 times in all address tables, its support is 3 / 5 = 0.6. Suppose we set the minimum support threshold to 0.6 (i.e., 60%), which means that any frequent item set needs to appear in at least 2 or more IP address tables.

[0099] First, it is necessary to construct frequent item sets. When determining frequent item sets through the Apriori algorithm, multiple rounds of construction of frequent item sets are required. First is the construction of frequent 1-item sets.

[0100] Calculate the support of each single IP address as follows:

[0101] 1.1: 3 / 5 = 0.6; 1.2: 3 / 5 = 0.6; 1.3: 4 / 5 = 0.8; 1.4: 4 / 5 = 0.8; 1.5: 4 / 5 = 0.8; 1.6: 1 / 5 = 0.2; 1.7: 1 / 5 = 0.2; 1.8: 1 / 5 = 0.2; 1.9: 1 / 5 = 0.2. Among them, the preset support threshold is 0.6, so the frequent 1-item sets are: {1.1, 1.2, 1.3, 1.4, 1.5}.

[0102] Next, it is necessary to determine the 2-item sets. Generate candidate 2-item sets from the frequent 1-item sets, that is, {1.1, 1.2}, {1.1, 1.3}, {1.1, 1.4}, {1.1, 1.5}, {1.2, 1.3}, {1.2, 1.4}, {1.2, 1.5}, {1.3, 1.4}, {1.3, 1.5}, {1.4, 1.5}, and calculate their support, that is, the frequency of the two corresponding elements in each item set appearing simultaneously.

[0103] Specifically, {1.1,1.2}:1 / 5 = 0.2; {1.1,1.3}:3 / 5 = 0.6; {1.1,1.4}:2 / 5 = 0.4; {1.1,1.5}:3 / 5 = 0.6; {1.2,1.3}:2 / 5 = 0.4; {1.2,1.4}:3 / 5 = 0.6; {1.2,1.5}:2 / 5 = 0.4; {1.3,1.4}:3 / 5 = 0.6; {1.3,1.5}:3 / 5 = 0.6; {1.4,1.5}:3 / 5 = 0.6. Similarly, setting the support threshold to 0.6, the frequent 2-itemsets are: {1.1,1.3}, {1.1,1.5}, {1.3,1.5}, {1.2,1.4}, {1.3,1.4}, {1.4,1.5}.

[0104] Next, the minimum confidence threshold needs to be defined. Suppose we choose the minimum confidence to be 0.8 (i.e., 80%). Association rules are generated from the frequent itemsets. For example, starting from the frequent 2-itemset {1.1,1.3}, its corresponding association relationship is {1.1}|{1.3}, where (1.1->1.3): 3 / 3 = 1, that is, the association confidence of 1.3 relative to 1.1 is 1. For the calculation method of the association confidence, it can be understood as the support of {1.1,1.3} divided by the support of {1.1}, that is, the frequency of the simultaneous appearance of 1.1 and 1.3 in the table divided by the frequency of the appearance of 1.1. And so on. The confidences between the IP addresses in each specific frequent 2-itemset are as follows:

[0105] {1.1,1.3}->{1.1}|{1.3}, confidence (1.1->1.3): 3 / 3 = 1, confidence (1.3->1.1): 3 / 4 = 0.75;

[0106] {1.1,1.5}->{1.1}|{1.5}, confidence (1.1->1.5): 3 / 3 = 1, confidence (1.5->1.1): 3 / 4 = 0.75;

[0107] {1.3,1.5}->{1.3}|{1.5}, confidence (1.3->1.5): 3 / 4 = 0.75, confidence (1.5->1.3): 3 / 4 = 0.75;

[0108] {1.2,1.4}->{1.2}|{1.4}, confidence (1.2->1.4): 3 / 3 = 1, confidence (1.4->1.2): 3 / 4 = 0.75;

[0109] {1.3,1.4}->{1.3}|{1.4}, confidence (1.3->1.4): 3 / 4 = 0.75, confidence (1.4->1.3): 3 / 4 = 0.75;

[0110] {1.4,1.5}->{1.4}|{1.5}, confidence level (1.4->1.5):3 / 4=0.75, confidence level (1.5->1.4):3 / 4=0.75.

[0111] The IP addresses involved in these frequent 2-item sets are frequent IP addresses.

[0112] Step S2042: determining the association relationship between the frequent IP addresses according to the association direction and the corresponding association confidence.

[0113] According to the association directions between these frequent IP addresses and the association confidences under each association direction, the association relationships between the frequent IP addresses are determined, which can also be regarded as specific association rules. These association relationships include specific association directions and specific association confidences. Next, a node association graph can be generated based on the specific association relationships.

[0114] Specifically, before generating a node association graph based on the association relationship between the multiple frequent IP addresses, the method further includes:

[0115] The association relationship with an association confidence less than a preset value is defined as a low association relationship, and the low association relationship between each frequent IP address is removed.

[0116] Taking the IP addresses determined in the above example as an example, assuming that the preset value is 0.8, the association relationship with an association confidence less than 0.8 is determined as a low association relationship, as follows:

[0117] {1.1,1.3}->{1.1}|{1.3}, confidence (1.3->1.1): 3 / 4 = 0.75;

[0118] {1.1,1.5}->{1.1}|{1.5}, confidence (1.5->1.1): 3 / 4 = 0.75;

[0119] {1.3,1.5}->{1.3}|{1.5}, confidence (1.3->1.5):3 / 4=0.75, confidence (1.5->1.3):3 / 4=0.75;

[0120] {1.2,1.4}->{1.2}|{1.4}, confidence (1.4->1.2): 3 / 4 = 0.75;

[0121] {1.3,1.4}->{1.3}|{1.4}, confidence (1.3->1.4):3 / 4=0.75, confidence (1.4->1.3):3 / 4=0.75;

[0122] {1.4, 1.5} -> {1.4}|{1.5}, confidence(1.4 -> 1.5): 3 / 4 = 0.75, confidence(1.5 -> 1.4): 3 / 4 = 0.75.

[0123] The corresponding association relationships involved above are low association relationships. After deleting them, the remaining association relationships are as follows:

[0124] {1.1, 1.3} -> {1.1}|{1.3}, confidence(1.1 -> 1.3): 3 / 3 = 1;

[0125] {1.1, 1.5} -> {1.1}|{1.5}, confidence(1.1 -> 1.5): 3 / 3 = 1;

[0126] {1.2, 1.4} -> {1.2}|{1.4}, confidence(1.2 -> 1.4): 3 / 3 = 1.

[0127] Step S205, generate a node association graph based on the association relationships between multiple frequent IP addresses. The node association graph is used to represent the path tracing of threat traffic.

[0128] Specifically, in the above step S205, generating a node association graph based on the association relationships between multiple frequent IP addresses includes:

[0129] Generate nodes corresponding to each frequent IP address, and generate directed edges between the corresponding nodes based on the association direction and the corresponding association confidence between each frequent IP address.

[0130] Generate a node association graph based on the nodes corresponding to each frequent IP address and the directed edges between the nodes.

[0131] Taking the above example, the frequent IP addresses are: 1.1, 1.3, 1.4, 1.5. Therefore, create nodes corresponding to these four frequent IP addresses, and generate directed edges corresponding to these nodes according to the specific association direction. For {1.1, 1.3} -> {1.1}|{1.3} (confidence 1.0), that is, the confidence of the directed edge where the node corresponding to 1.1 points to the node corresponding to 1.3 is 1.0. And so on, generate directed edges between each node to obtain the node association graph. The confidence of the directed edges between the nodes in the node association graph can represent that when the traffic data corresponding to the IP address of a certain node is threat traffic data, the possibility that the previous node corresponding to the directed edge is the source of the threat traffic data. If the corresponding confidence is greater, the IP address corresponding to the previous node is more likely to be the source of the threat traffic data. The staff can analyze the source of the threat traffic data corresponding to each node according to the edge relationship between the specific associated nodes.

[0132] Step S206, determine the influence weights of the nodes corresponding to each frequent IP address in the node association graph through a preset link analysis algorithm; the influence weights are used to characterize the participation degree of the corresponding frequent IP address in the threat traffic behavior.

[0133] After determining the specific node association graph, the influence weights of each node therein can be determined through a preset link analysis algorithm. Exemplarily, PageRank can be used to determine the influence weights of specific nodes. Specifically, the main basis for determining the influence weights is the number of in-links and out-links corresponding to each node, that is, for each node, the number of directed edges pointing to it and the number of directed edges it points to other nodes. The specific calculation method will not be elaborated here.

[0134] After determining the influence weights of each node, it can further assist the staff in making the following judgments:

[0135] Determine key nodes: The PageRank algorithm can help determine the most important nodes in the traceability graph. In security traceability, these nodes may be the key IP addresses in the attack path or the core participants in the attack event. By calculating the PageRank value of the nodes, the most influential and critical nodes can be identified, and these nodes may need to be prioritized or further investigated.

[0136] Discover hidden associations: The traceability graph can show direct IP address association relationships, while the PageRank algorithm can discover important associations hidden in complex network structures. For example, even if there is no direct communication record between some IP addresses, they may be indirectly connected through common intermediate nodes or common patterns. PageRank can help reveal these indirect and potential association relationships.

[0137] Measure the influence of nodes:

[0138] Security traceability not only focuses on the connection relationships between nodes, but also needs to consider the influence and importance of nodes. The PageRank algorithm calculates the importance of nodes through iteration and assigns weights based on the connection situation of nodes and the importance of the connected nodes. This method helps to quantify the influence of nodes in the attack path or network, so as to optimize security analysis and response strategies.

[0139] Explore the propagation and expansion of the attack path: The PageRank algorithm is not only applicable to a single traceability graph, but can also analyze the propagation and expansion of attacks in complex attack paths. By applying PageRank to the entire attack network, the main propagation paths of the attack and the nodes with the greatest influence can be identified, helping the security team understand the global impact and spread mode of the attack.

[0140] Support for decision-making and response: The analysis results of the traceability graph combined with the PageRank algorithm can provide deeper insights for the security team, supporting more effective decision-making and response measures. For example, resources can be preferentially allocated or defensive measures can be taken based on the nodes ranked by PageRank to minimize the impact and spread of security incidents.

[0141] The graph relationship of the IP traceability path has been constructed according to the association rules and the PageRank algorithm, and the association relationship of risky IPs can be visually seen. After selecting the IP of the specified target, all related IPs and the paths between the IPs can be obtained by using the constructed traceability path graph relationship. These paths can help users more intuitively understand the association relationship between dangerous IPs and other IPs, so as to conduct more in-depth security analysis.

[0142] On the IP details page for viewing the specified target, users can obtain threat event statistics and threat event details related to the IP, which helps users comprehensively understand the security threat situation involved in a specific IP and helps users take corresponding security protection measures.

[0143] The threat path traceability method provided by the embodiment of the present invention generates a first IP address set through the source IP address and destination IP address corresponding to the threat traffic data recorded in the security log, determines the dangerous IP addresses among them, and thus generates an associated IP address list according to the associated IP addresses corresponding to each dangerous IP, and finally determines the frequent IP addresses and the association relationship between these frequent IP addresses, so as to finally generate a node association graph that can represent the threat traffic path traceability, so that the staff can more intuitively understand the attack path of the threat traffic, facilitating timely pre-processing by relevant network security personnel and ensuring the security of network assets.

[0144] To assist in understanding the specific analysis process of the association rules in the embodiment of the present invention, the following content is used to specifically explain the association rule mining algorithm. First, the obtained relevant IP list is modeled for association rules; the Apriori algorithm (association rule mining algorithm) is used, and its basic information is as follows.

[0145] 1. Basic concepts: Association rule mining can enable us to discover the relationships between items (item and item) in a dataset. It has many application scenarios in our lives. "Basket analysis" is a common scenario, which can discover the association relationships between products from consumer transaction records, and then bring more sales through product bundling or related recommendations.

[0146] Take an example of supermarket shopping. The following Table 4 is a list of products purchased by several customers:

[0147] Table 4

[0148] Order number Purchased goods 1 Milk, bread, diapers 2 Cola, bread, diapers, beer 3 Milk, diapers, beer, eggs 4 Bread, milk, diapers, beer 5 Bread, milk, diapers, cola

[0149] 2. Support:

[0150] Support is a percentage, which refers to the ratio between the number of occurrences of a certain product combination and the total number of occurrences.

[0151] In this example, we can see that "milk" appears 4 times. Then the support of "milk" in these 5 orders is 4 / 5 = 0.8.

[0152] Similarly, "milk + bread" appears 3 times. Then the support of "milk + bread" in these 5 orders is 3 / 5 = 0.6.

[0153] 3. Confidence:

[0154] It refers to the probability of purchasing product B when you have purchased product A.

[0155] Confidence (milk → beer) = 2 / 4 = 0.5, which represents the probability of purchasing beer if you have purchased milk.

[0156] Confidence (beer → milk) = 2 / 3 = 0.67, which represents the probability of purchasing milk if you have purchased beer.

[0157] Confidence is a conditional concept, that is, given that A has occurred, what is the probability that B will occur.

[0158] 4. Lift:

[0159] When making product recommendations, we mainly consider lift, because lift represents the degree to which the occurrence of product A increases the probability of the occurrence of product B.

[0160] Lift (A → B) = Confidence (A → B) / Support (B).

[0161] This formula is used to measure whether the occurrence of A will increase the probability of the occurrence of B.

[0162] So there are three possibilities for lift:

[0163] Lift (A → B) > 1: It represents an increase;

[0164] Lift (A → B) = 1: It represents neither an increase nor a decrease;

[0165] Lift (A → B) < 1: It represents a decrease.

[0166] Regarding the lift, it is not set in the embodiments of this method. However, in combination with the actual situation, it can be determined whether to set it according to the specific application scenario, and there is no limitation here.

[0167] 5. Frequent itemset:

[0168] If the support of itemset X is greater than or equal to the specified minimum support threshold, then the itemset X is called a frequent itemset.

[0169] In the process of executing the Apriori algorithm, first, we represent the products in the above case with IDs. The product IDs of milk, bread, diapers, cola, beer, and eggs are set to 1-6 respectively. The above data table can be changed to Table 5:

[0170] Table 5

[0171] Order number Purchased goods 1 4、2、3、5 2 4、2、3、5 3 1、3、5、6 4 2、1、3、5 5 2、1、3、4

[0172] The Apriori algorithm is actually a process of finding frequent itemsets. A frequent itemset is an itemset whose support is greater than or equal to the minimum support threshold. Therefore, items with a support less than the minimum support are non-frequent itemsets, and itemsets with a support greater than or equal to the minimum support are frequent itemsets.

[0173] Suppose I randomly specify the minimum support as 50%, that is, 0.5. First, we calculate the support of individual products, that is, we obtain the support of K = 1 items, as shown in Table 6:

[0174]

[0175]

[0176] Since the minimum support is 0.5, you can see that products 4 and 6 do not meet the minimum support and do not belong to the frequent itemset. Therefore, after screening, the frequent itemset of products becomes:

[0177] Item set of goods Support degree 1 3 / 5 2 4 / 5 3 1 5 4 / 5

[0178] On this basis, we combine the products in pairs to obtain the support of K = 2 items, as shown in Table 7:

[0179] Table 7

[0180] Item set of goods Support degree 1,2 2 / 5 1,3 3 / 5 1,5 2 / 5 2,3 4 / 5 2,5 3 / 5 3,5 4 / 5

[0181] Then, screen out the product combinations with a support less than the minimum support, and Table 8 can be obtained:

[0182] Table 8

[0183]

[0184]

[0185] Recursive process of Apriori algorithm:

[0186] 1. When K = 1, calculate the support of the K-item set;

[0187] 2. Filter out the item sets with support less than the minimum support;

[0188] 3. If the item set is empty, the result of the corresponding (K - 1)-item set is the final result; otherwise, K = K + 1, and repeat steps 1 - 3.

[0189] When determining frequent item sets through the Apriori algorithm, it can be determined that it recurs multiple times to determine the support of 3-item sets or even 4-item sets. In the embodiment of this method, since only the support of 2-item sets needs to be determined, the subsequent calculation process of the support of other k-item sets will not be elaborated here. In some cases, in order to determine a more accurate association relationship, for the embodiment of this method, the way of using higher item sets can also be used to determine association rules. However, since there are often many IP addresses, when using higher item sets, it will consume a large amount of computing resources and computing time. Therefore, more preferably, the way of using 2-item sets is used to determine association rules.

[0190] After determining the 2-item sets corresponding to k equal to 2, specific association rules need to be determined. For example, in the above example, the item sets of goods are {1, 3}, {2, 3}, {2, 5}, and {3, 5}. Among them, for {1, 3}, there are association rules 1 -> 3 and 3 -> 1. The association rule 1 -> 3 means that when item 1 appears, the possibility that item 3 appears simultaneously. The specific calculation method of confidence is: the support corresponding to the simultaneous appearance of item 1 and 3 divided by the support of item 1, that is, (3 / 5) / (3 / 5), which is 1. On the contrary, 3 -> 1 means that when item 3 appears, the possibility that item 1 appears simultaneously. The specific calculation method of confidence is: the support corresponding to the simultaneous appearance of item 1 and 3 divided by the support of item 3, that is, (3 / 5) / 1, which is 3 / 5. And so on, determine the association rules between different goods. After determining the confidence corresponding to each association rule, screening can be performed through a preset threshold to obtain relatively strong association rules.

[0191] To assist in understanding the step S6 in the above method embodiment, that is, through the preset link analysis algorithm, the following example can be used for auxiliary understanding.

[0192] The preset link analysis algorithm can be a PageRank algorithm, which is usually used to rank web pages, calculate the importance of websites, and optimize search engine search results. The PR value is a factor that represents its importance. The central idea is that in a node graph, the more in-links a node receives from other nodes, the more important the node is. When a high-quality node points to (out-links) a node, it means that the pointed node is important, and the node is the frequent IP address mentioned in the above method embodiment. It mainly analyzes the number of outlinks and inlinks corresponding to each node to determine the specific PR value, that is, the impact weight. For example, Figure 3 FIG. 1 is a schematic diagram of node inbound links and outbound links, in which node A has two inbound links and two outbound links.

[0193] When calculating the PR value of each node, the specific calculation formula is as follows:

[0194]

[0195] Among them, PR(Ti): PR value of other nodes (pointing to node a), L(Ti): number of outbound links of other nodes (pointing to node a), i: number of cycles. The main process of the algorithm is to give each node a PR value (hereinafter PR value refers to PageRank value). The (voting) algorithm is continuously iterated until a stable distribution is reached, and the importance of each node is determined according to the PR value at this time.

[0196] With the above Figure 3 Take the node graph shown in the figure as an example, where the PR value corresponding to each iteration is shown in Table 9:

[0197] Table 9

[0198]

[0199] The initialized PR value is 1 / N=1 / 4. Taking node C as an example, when i=1, the corresponding PR value calculation process is as follows. The PR value calculation method corresponding to other nodes is similar and will not be repeated here.

[0200]

[0201] In another calculation method, in order to improve the calculation efficiency, the above can be represented by the transition probability matrix, namely the Markov matrix. Figure 3 The node graph shown is as follows:

[0202]

[0203] Among them, taking the first column as an example, it means that the probability that A jumps to B or C is 1 / 2, and the probability that D jumps to A is 1. The columns of the matrix represent out-links. Through matrix representation, the PR value can be calculated quickly. The specific calculation method is: PR(a) = M * V, where M is the matrix corresponding to the node graph, and V is the PR value of each current node. The specific calculation method is as Figure 4 shown Figure 4 is a schematic flowchart of calculating the PR value based on the transition probability matrix according to an embodiment of the present invention.

[0204] First, the PR value of each initialized node is 1 / 4. After the first iteration, by multiplying the transition probability matrix by the PR value corresponding to the current node, the PR values corresponding to the four nodes A, B, C, and D are 3 / 8, 1 / 8, 3 / 8, and 1 / 8 respectively.

[0205] Both of these calculation methods determine the specific PR value through the PageRank algorithm, and there are only differences in the calculation methods. The specific calculation method can be selected according to the actual situation.

[0206] In this embodiment, a threat path tracing device is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can implement a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0207] This embodiment provides a threat path tracing device, as Figure 5 shown, including:

[0208] A log information acquisition module 401, configured to acquire security log information of the network security system in a preset historical period. The security log information includes: source IP addresses and destination IP addresses corresponding to multiple threat traffic data;

[0209] A dangerous address determination module 402, configured to generate a first IP address set based on the source IP addresses and destination IP addresses corresponding to each threat traffic data, and determine multiple dangerous IP addresses in the first IP address set;

[0210] An address list generation module 403, configured to determine associated IP addresses corresponding to multiple dangerous IP addresses respectively according to the security log information, and generate an associated IP address list;

[0211] An association relationship determination module 404, configured to determine frequently-occurring IP addresses in the associated IP address list and the association relationships between each frequently-occurring IP address;

[0212] The association graph generation module 405 is used to generate a node association graph based on the association relationship between multiple frequent IP addresses, and the node association graph is used to characterize the path tracing of threat traffic.

[0213] In some optional implementations, the address list generating module 403, when determining the associated IP addresses corresponding to each of the plurality of dangerous IP addresses according to the security log information and generating the associated IP address list, includes:

[0214] According to the security log information, determine the corresponding associated destination IP address when each dangerous IP address is the source IP address, and the corresponding associated source IP address when each dangerous IP address is the destination IP address;

[0215] Generate a corresponding first sub-table according to each dangerous IP address and the corresponding associated destination IP address, and generate a corresponding second sub-table according to each dangerous IP address and the corresponding associated source IP address;

[0216] A list of associated IP addresses is constructed according to the first sub-table and the second sub-table corresponding to each dangerous IP address.

[0217] In some optional implementations, the dangerous address determination module 402, when determining a plurality of dangerous IP addresses in the first IP address set, includes:

[0218] Determine the public information corresponding to each IP address in the first IP address set;

[0219] Based on the public information corresponding to each IP address, a plurality of dangerous IP addresses in the first IP address set are determined.

[0220] In some optional implementations, the association relationship determination module 404, when determining the frequent IP addresses in the associated IP address list and the association relationship between each frequent IP address, includes:

[0221] By using a preset association rule mining algorithm, the frequent IP addresses in the associated IP address list and the corresponding association confidences between the frequent IP addresses in different association directions are determined;

[0222] The association relationship between each frequent IP address is determined according to the association direction and the corresponding association confidence.

[0223] In some optional implementations, before generating a node association graph based on the association relationships between multiple frequent IP addresses, the association determination module 404 is also used to define an association relationship with an association confidence less than a preset value as a low association relationship, and to release the low association relationship between each frequent IP address.

[0224] In some alternative embodiments, when generating a node association graph based on the association relationships between multiple frequent IP addresses, the association graph generation module 405 includes:

[0225] Generating nodes corresponding to each frequent IP address, and generating directed edges between the corresponding nodes based on the association directions and corresponding association confidence levels between the frequent IP addresses;

[0226] Generating a node association graph based on the nodes corresponding to each frequent IP address and the directed edges between the nodes.

[0227] In some alternative embodiments, the association graph generation module 405 is further configured to determine the influence weights of the nodes corresponding to each frequent IP address in the node association graph through a preset link analysis algorithm; the influence weights are used to characterize the participation degree of the corresponding frequent IP address in the threat traffic behavior.

[0228] The further functional descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be elaborated here.

[0229] The threat path tracing device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0230] The embodiment of the present invention further provides a computer device having the above Figure 5 shown threat path tracing device.

[0231] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As Figure 6 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other through different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as a server array, a set of blade servers, or a multi-processor system). Figure 6Take a processor 10 as an example.

[0232] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.

[0233] Among them, the memory 20 stores instructions that can be executed by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiments.

[0234] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and a combination thereof.

[0235] The memory 20 can include a volatile memory, for example, a random access memory; the memory can also include a non-volatile memory, for example, a flash memory, a hard disk, or a solid-state drive; the memory 20 can also include a combination of the above types of memories.

[0236] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 can be connected through a bus or other means. Figure 6 Take the connection through the bus as an example.

[0237] The input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 can include a display device, an auxiliary lighting device (for example, an LED), and a tactile feedback device (for example, a vibration motor), etc. The above display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device can be a touch screen.

[0238] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0239] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A threat path tracing method, characterized in that, The method comprises: Obtain security log information of a network security system in a preset historical period, wherein the security log information includes: a source IP address and a destination IP address corresponding to a plurality of threat traffic data; Generate a first IP address set based on the source IP address and the destination IP address corresponding to each threat traffic data, and determine a plurality of dangerous IP addresses in the first IP address set; According to the security log information, determine the associated IP addresses corresponding to each of the multiple dangerous IP addresses, and generate an associated IP address list, wherein the associated IP address list includes multiple sub-tables generated by each dangerous IP address and the corresponding associated source IP address and multiple sub-tables generated by the corresponding associated destination IP address; Determine the frequent IP addresses in the associated IP address list and the association relationship between each frequent IP address; Generate a node association graph based on the association relationship between the frequent IP addresses, where the node association graph is used to characterize the path tracing of the threat traffic; The determining of the frequent IP addresses in the associated IP address list and the association relationship between the frequent IP addresses includes: By using a preset association rule mining algorithm, according to the frequency of a single IP address appearing in all sub-tables and the frequency of two IP addresses appearing in the same sub-table in all sub-tables, the frequent IP addresses in the associated IP address list and the corresponding association confidences between the frequent IP addresses in different association directions are determined; The association relationship between the frequent IP addresses is determined according to the association direction and the corresponding association confidence.

2. The method according to claim 1, wherein The step of determining, based on the security log information, the associated IP addresses corresponding to each of the plurality of dangerous IP addresses, and generating an associated IP address list includes: Determine, based on the security log information, when each dangerous IP address is a source IP address, the corresponding associated destination IP address, and when each dangerous IP address is a destination IP address, the corresponding associated source IP address; Generate a corresponding first sub-table according to each dangerous IP address and the corresponding associated destination IP address, and generate a corresponding second sub-table according to each dangerous IP address and the corresponding associated source IP address; A list of associated IP addresses is constructed according to the first sub-table and the second sub-table corresponding to each dangerous IP address.

3. The method according to claim 1, wherein The determining of a plurality of dangerous IP addresses in the first IP address set includes: Determine the public information corresponding to each IP address in the first IP address set; Based on the public information corresponding to each IP address, a plurality of dangerous IP addresses in the first IP address set are determined.

4. The method according to claim 3, characterized in that Before generating a node association graph based on the association relationship between the frequent IP addresses, the method further includes: An association relationship whose association confidence is less than a preset value is defined as a low association relationship, and the low association relationship between the frequent IP addresses is released.

5. The method according to claim 3, characterized in that, The generating a node association graph based on the association relationship between the frequent IP addresses includes: Generate a node corresponding to each frequent IP address, and generate a directed edge between the corresponding nodes based on the association direction and the corresponding association confidence between the frequent IP addresses; Generate a node association graph based on the nodes corresponding to each frequent IP address and the directed edges between the nodes.

6. The method according to claim 1, characterized in that, The method further includes: Determine the influence weights of the nodes corresponding to each frequent IP address in the node association graph through a preset link analysis algorithm; the influence weights are used to characterize the participation degree of the corresponding frequent IP address in the threat traffic behavior.

7. A threat path tracing device, characterized in that, The device includes: A log information acquisition module, configured to acquire security log information of the network security system in a preset historical period, where the security log information includes source IP addresses and destination IP addresses corresponding to a plurality of threat traffic data; A dangerous address determination module, configured to generate a first IP address set based on the source IP addresses and destination IP addresses corresponding to each threat traffic data, and determine a plurality of dangerous IP addresses in the first IP address set; An address list generation module, configured to determine the associated IP addresses corresponding to each of the plurality of dangerous IP addresses according to the security log information, and generate an associated IP address list, where the associated IP address list includes a plurality of sub-tables generated by each dangerous IP address and the corresponding associated source IP addresses respectively, and a plurality of sub-tables generated by the corresponding associated destination IP addresses; An association relationship determination module, configured to determine the frequent IP addresses in the associated IP address list and the association relationships between the frequent IP addresses; The determination of the frequent IP addresses in the associated IP address list and the association relationships between the frequent IP addresses includes: Determine the frequent IP addresses in the associated IP address list and the association confidence levels corresponding to the frequent IP addresses in different association directions according to the frequency of a single IP address appearing in all sub-tables and the frequency of two IP addresses appearing in the same sub-table in all sub-tables through a preset association rule mining algorithm; Determine the association relationships between the frequent IP addresses according to the association direction and the corresponding association confidence levels; An association graph generation module, configured to generate a node association graph based on the association relationships between the frequent IP addresses, where the node association graph is used to represent the path traceability of threat traffic.

8. A computer device, characterized in that, Includes: A memory and a processor, where the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the threat path traceability method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the threat path traceability method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Flow control device and method based on flow prediction and trusted network address learning

    CN101729389A

  • Analytical method for security log based on Apriori algorithm

    CN108255996A