Method and system for collecting network security threat information
By extracting and standardizing cyber attack behavior data, combining chronological order and target consistency analysis, identifying and combining attack chains and calculating event priorities, the limitations of attack path identification and deep merger of threat events in the existing technology are solved, and an efficient threat intelligence analysis and emergency response mechanism are achieved.
Patent Information
- Application Number
- CN202510260959.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The existing technology has limitations in the identification of attack paths and the deep merger of threat events. It fails to fully consider the consistency of targets and the evolution of attack methods, resulting in the break of the attack chain, affecting the overall traceability of attack behavior, and lacks in-depth analysis of time dependence, making it impossible to effectively extract high-correlation attack paths and reduce the ability to identify complex attack chains.
By extracting cyber attackers and attack methods, identifying and classifying attacker identification and target types, the data is standardized to form a standardized threat behavior data set. Then, the attack events are arranged based on chronological order, combined with time threshold and target consistency analysis, the same attack chain is identified and merged, the attack chain impact factor is calculated, the traceability path is divided, the attack target type frequency is counted, the event attribute weight is evaluated, the event attribute weight is judged, whether the event is merged or split, and finally the node weight value is calculated, the event priority is sorted, and the emergency response path is filtered.
Effectively identify and merge the same attack chain, enhance the consistency and deep correlation of threat intelligence, improve the accuracy of attack path identification, optimize the hierarchical classification of threat intelligence, avoid information redundancy or misclassification, strengthen the predictiveness of attack behavior, and ensure that the emergency response mechanism focuses on the most threatening attack paths, thereby improving the pertinence and efficiency of overall security response.
Smart Images

Figure CN119743335B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a method and system for collecting network security threat information. Background Art
[0002] The field of network security technology includes network data protection, attack detection, intrusion prevention, threat intelligence analysis and other aspects. Its core content is to identify, analyze and protect security risks in the network environment to ensure the security and stability of the information system. This technical field involves data encryption, identity authentication, access control, log auditing, malicious behavior analysis and other technical means, and through the establishment of network monitoring systems, threat intelligence platforms, vulnerability management mechanisms and other means to achieve comprehensive management of security threats. In the network security protection system, the collection, analysis and early warning of threat information are key links. Its main purpose is to organize and classify security threat information from different sources to improve the pertinence and real-time nature of security protection.
[0003] Among them, the network security threat information collection method refers to the use of structured processing, data correlation analysis and multi-source information integration to achieve systematic collection of threat intelligence for scattered network security threat data. This method covers technical aspects such as threat data collection, formatting, correlation analysis and storage management. Specifically, it includes extracting threat information from different network security devices, log records, vulnerability databases and other channels, standardizing the collected data, identifying the correlation features through pattern matching and semantic analysis, and classifying and hierarchically organizing the collected information based on time series analysis, credibility assessment and other means to form a threat intelligence data set that can be used for security monitoring and early warning.
[0004] Existing technologies mainly rely on multi-source data collection and correlation analysis, but they are limited in attack path identification and deep merging of threat events. As existing methods mostly use a single time dimension to classify events, they do not fully consider the consistency of targets and the evolution of attack methods, resulting in a break in the attack chain and affecting the overall traceability of attack behaviors. The hierarchical division of events mainly relies on preset rules, and the analysis of the co-occurrence relationship between different targets is insufficient, which can easily lead to inaccurate classification of some attack events, resulting in redundant or missing intelligence. In terms of attack pattern recognition, existing methods lack in-depth analysis of time dependence, cannot effectively extract highly correlated attack paths, and reduce the ability to identify complex attack chains. Existing methods lack a weight ranking mechanism for attack behaviors in emergency response, resulting in unreasonable resource allocation, inability to prioritize the most urgent security threats, and reducing the effectiveness of overall security defense. Summary of the invention
[0005] In order to solve the problem that the existing technology mainly relies on multi-source data collection and correlation analysis, but has limitations in attack path identification and deep merging of threat events, since the existing methods mostly use a single time dimension to classify events, they do not fully consider the consistency of targets and the evolution of attack methods, resulting in a break in the attack chain and affecting the overall traceability of attack behaviors. The hierarchical division of events mainly relies on preset rules, and the co-occurrence relationship analysis between different targets is insufficient, which can easily cause inaccurate classification of some attack events, resulting in redundant or missing intelligence. In terms of attack pattern recognition, the existing methods lack in-depth analysis of time dependence, cannot effectively extract highly correlated attack paths, and reduce the ability to identify complex attack chains. The existing methods lack a weight ranking mechanism for attack behaviors in emergency response, resulting in unreasonable resource allocation, and cannot prioritize the most urgent security threats, reducing the effectiveness of overall security defense. Technical problems, the embodiments of the present invention provide a method and system based on network security threat information collection. The technical solution is as follows:
[0006] On the one hand, a method for collecting network security threat information is provided, the method comprising:
[0007] S1: Based on network security threat information, extract network attackers and attack methods, identify and classify attacker identification and target type, standardize the data, and obtain a standardized threat behavior data set;
[0008] S2: Based on the standardized threat behavior dataset, extract attack timestamps, arrange them in chronological order, identify the time intervals between adjacent attack events, determine whether the events are continuous, filter related events and classify them into the same attack chain, calculate the impact factor of the attack chain, divide the tracing path, and obtain the attack event time series;
[0009] S3: calling the attack event time series, counting the frequency of the attack target type in the differentiated events, evaluating the event attribution weight, determining whether the events should be merged or split, and obtaining a hierarchical threat event analysis result;
[0010] S4: Based on the hierarchical threat event analysis results, extract key behavior nodes, analyze time series dependencies, filter out related behaviors and classify them into similar attack paths, and obtain attack behavior path matching results;
[0011] S5: Call the attack behavior path matching result, calculate the node weight value, sort the event priority according to the weight value, filter the weight path and classify it into emergency response, and obtain the network attack path threat aggregation result.
[0012] As a further solution of the present invention, the standardized threat behavior data set includes attacker identification, target type, and subject-predicate-object structure parsing results; the attack event time series includes attack timestamp, event continuity determination results, and attack chain influencing factors; the hierarchical threat event analysis results include event attribution weights, target co-occurrence probabilities, and event hierarchical classifications; the attack behavior path matching results include behavior time offsets, dependency analysis results, and correlation behavior classifications; and the network attack path threat aggregation results include event priorities, node weight values, and emergency response paths.
[0013] As a further solution of the present invention, the step of standardizing the threat behavior data set is specifically:
[0014] S101: Based on network security threat information, including security event logs and intrusion detection alarms, extract attacker identification, attack methods and target information, remove duplicate records, and obtain attack behavior matching data volume;
[0015] S102: Based on the attack behavior matching data volume, the semantic relationship between the attacker identification, attack method and target is analyzed, the subject-predicate-object structure is extracted, the attack behavior field is mapped to the corresponding attack method, the attacker identification under the same attack method is merged, the number of attackers with multiple attack methods is counted, and the attack methods are classified according to the attack method number threshold to obtain the attack method distribution range;
[0016] S103: Call the attack mode distribution interval, analyze the target type distribution by attacker identification classification, group the target IP, industry category and asset type, cross-match the attack modes and summarize the key attack modes, and generate a standardized threat behavior data set.
[0017] As a further solution of the present invention, the steps of the attack event time sequence are specifically:
[0018] S201: Arrange the attack timestamps of the standardized threat behavior data set in chronological order, identify the time intervals between adjacent events, filter out events that do not exceed a threshold, and generate an attack event time interval sequence;
[0019] S202: Based on the attack event time interval sequence, for events that exceed the time threshold, analyze the consistency of attack targets, extract the target IP, port number and protocol type, filter qualified events and classify them into the same attack chain, count the number of events, time span and number of targets in the chain, and obtain attack chain statistics;
[0020] S203: Call the attack chain statistical data, calculate the attack chain impact factor, divide the tracing path according to the number of events, time span and target number, and generate the attack event time series.
[0021] As a further solution of the present invention, the attack chain impact factor adopts the formula:
[0022] ;
[0023] in, Represents the attack chain impact factor, Represents the total number of attack events, Representative The severity coefficient of an attack event, Representative The time span of the attack events, Representative The number of targets affected by an attack event, Represents the average number of targets affected by the attack event, Indicates the total number of attack events.
[0024] As a further solution of the present invention, the step of hierarchical threat event analysis results is specifically:
[0025] S301: calling the attack event time series, counting the occurrence frequency of attack target types in differentiated events, classifying them by target IP, port number and protocol type, and obtaining attack target frequency distribution data;
[0026] S302: Based on the attack target frequency distribution data, the target co-occurrence probability is calculated, the target type combination in the differentiated event is extracted, the number of occurrences of the combination is counted, and the event attribution weight is calculated according to the co-occurrence number and the total number of events according to the probability threshold, and an event attribution weight matrix is generated;
[0027] S303: calling the event attribution weight matrix, dividing the events according to the attribution weights, merging the events above the threshold, splitting the events below the threshold, and obtaining hierarchical threat event analysis results.
[0028] As a further solution of the present invention, the step of attack behavior path matching result is specifically:
[0029] S401: extracting key behavior nodes based on the hierarchical threat event analysis results, arranging the behavior nodes in chronological order, screening nodes with offset characteristics on the time axis, calculating the time difference between adjacent nodes, and obtaining the key behavior time offset;
[0030] S402: Analyze the dependency relationship in the time series according to the key behavior time offset, identify the time sequence relationship between adjacent behavior nodes, calculate the dependency strength value between nodes according to the time interval and the triggering order between nodes, and select the behavior node combination with a dependency greater than a preset threshold to obtain a time dependency relationship data set;
[0031] S403: Based on the time dependency relationship data set, the behavior nodes with high correlation are screened, the nodes belonging to the same attack target are clustered, the nodes of the same type are merged and the belonging paths are divided, and the attack behavior path matching result is generated.
[0032] As a further solution of the present invention, the step of collecting the threat results of the network attack path is specifically as follows:
[0033] S501: calling the attack behavior path matching result, extracting the behavior node triggering frequency, counting the number of occurrences of the node in the differentiated path, and obtaining the behavior node triggering ratio;
[0034] S502: Calculate the node weight value according to the behavior node trigger ratio, combine the trigger frequency and path distribution to count the node contribution, arrange the event priorities in descending order according to the weight value, select the paths with high weights and classify them into emergency response, and obtain the key threat path;
[0035] S503: Based on the key threat path, threat events under the emergency response path are aggregated, the impact scope and the number of attack targets are counted, and a network attack path threat aggregation result is generated.
[0036] As a further solution of the present invention, the behavior node trigger frequency adopts the formula:
[0037] ;
[0038] in, Represents the behavior node trigger frequency, Representing behavior nodes In the The time interval between triggers, Representing behavior nodes The average time interval between triggers, Representing behavior nodes The total number of triggers in the path, Representing behavior nodes In the The path weight at the time of triggering, Representing behavior nodes In the The influence coefficient of the trigger.
[0039] On the other hand, based on the network security threat information collection system, according to the above-mentioned network security threat information collection method, the system includes:
[0040] The attack behavior identification module extracts the attacker identification, attack method, and attack target based on network security threat information, calculates the corresponding offset of the attack time, compares the cross-degree of the attack target, screens the target attack behavior, summarizes the characteristic combination of the attack mode, and obtains the attack behavior mode;
[0041] Based on the attack behavior pattern, the attack chain construction module compares the cross frequency of the attack target, screens the attack behaviors that meet the time window, analyzes the influence relationship of the differentiated nodes, summarizes the attribution of multiple attack behaviors, adjusts the transmission path of the attack chain, screens the chain with large influence factors, and obtains the attack chain structure information;
[0042] The threat event attribution module calculates the crossover frequency of attack targets in the chain based on the attack chain structure information, analyzes the co-occurrence relationship of the targets, identifies the attack event attribution boundary, screens attack events with attribution conflicts, adjusts the chain attribution of unevenly distributed targets, and obtains the attack event attribution mapping;
[0043] The attack path reconstruction module extracts the core behavior nodes of the path based on the attack event attribution mapping, adjusts the attack path sequence, compares the behavior correlation, summarizes the logically similar paths, screens the key path derivative relationship, and obtains the attack behavior correlation path;
[0044] The threat impact assessment module extracts the attack node triggering frequency based on the attack behavior association path, calculates the impact diffusion value, screens the attack nodes with a large impact range, summarizes the high-risk attack paths, and obtains the network attack path threat aggregation result.
[0045] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0046] By extracting attack behavior information from security event logs and intrusion detection alerts and parsing the subject-predicate-object structure, we ensure the standardization and structuring of data, improve the ability to accurately identify the attacker's identity and attack method, arrange attack events in chronological order, and combine time thresholds and target consistency analysis to effectively identify and merge the same attack chain, enhance the coherence and deep correlation of threat intelligence, and improve the accuracy of attack path identification. We use target co-occurrence probability and consistency weight calculation to ensure the rationality of event merging and splitting, optimize the threat intelligence hierarchical classification, and avoid analysis bias caused by information redundancy or misclassification. By analyzing the time dependency of behavior nodes, we can summarize similar attack paths, improve the ability to identify attack patterns, strengthen the predictability of attack behaviors, prioritize attack paths according to node weight values, accurately screen high-risk attack chains, and ensure that the emergency response mechanism focuses on the most threatening attack paths, thereby improving the pertinence and efficiency of the overall security response. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a schematic diagram of the workflow of the present invention;
[0048] Figure 2 This is a detailed flow chart of S1 of the present invention;
[0049] Figure 3 This is a detailed flow chart of S2 of the present invention;
[0050] Figure 4 This is a detailed flow chart of S3 of the present invention;
[0051] Figure 5 This is a detailed flow chart of S4 of the present invention;
[0052] Figure 6 This is a detailed flow chart of S5 of the present invention;
[0053] Figure 7 It is a system structure diagram of the present invention. DETAILED DESCRIPTION
[0054] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0055] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0056] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0057] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.
[0058] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0059] See also Figure 1 The embodiment of the present invention provides a method for collecting network security threat information. The processing flow of the method may include the following steps:
[0060] S1: Based on network security threat information, including security event logs and intrusion detection alerts, network attackers and attack methods are extracted, subject-verb-object structures are parsed, text formats are converted, and standardized threat behavior data sets are obtained based on attacker identification and target type classification;
[0061] S2: Based on the standardized threat behavior dataset, extract attack timestamps, arrange them in chronological order, identify the time intervals between adjacent attack events, and determine whether the events are continuous based on the time threshold. If the time interval exceeds the threshold, analyze the consistency of the attack target, filter related events and classify them into the same attack chain, calculate the impact factor of the attack chain, divide the tracing path, and obtain the time series of attack events;
[0062] S3: Call the attack event time series, count the frequency of attack target types in differentiated events, calculate the target co-occurrence probability, evaluate the event attribution weight according to the target consistency, determine whether the event should be merged or split, and obtain the hierarchical threat event analysis results;
[0063] S4: Based on the hierarchical threat event analysis results, extract key behavior nodes, calculate behavior time offsets, analyze dependencies in the time series, filter out related behaviors and classify them into similar attack paths, and obtain attack behavior path matching results;
[0064] S5: Call the attack behavior path matching results, extract the behavior node triggering frequency, calculate the node weight value, sort the event priority according to the weight value, filter the high-weight path and classify it into emergency response, and obtain the network attack path threat aggregation result.
[0065] The standardized threat behavior data set includes attacker identification, target type, and subject-predicate-object structure parsing results. The attack event time series includes attack timestamp, event continuity determination results, and attack chain influencing factors. The hierarchical threat event analysis results include event attribution weights, target co-occurrence probabilities, and event hierarchical classifications. The attack behavior path matching results include behavior time offsets, dependency analysis results, and correlation behavior classifications. The network attack path threat aggregation results include event priority, node weight values, and emergency response paths.
[0066] Specifically, if Figure 2 As shown in the figure, the steps to standardize the threat behavior dataset are as follows:
[0067] S101: Based on network security threat information, including security event logs and intrusion detection alarms, extract attacker identification, attack methods and target information, remove duplicate records, and obtain attack behavior matching data volume;
[0068] First, determine the data source, call the network log analysis tool, filter out information containing attack behaviors from security log data from different sources, perform structured analysis on the filtered data, extract key parameters such as IP address, port number, attack time, etc. At the same time, match the alarm information of the intrusion detection system (IDS) and the security information and event management (SIEM) system to identify events related to network attacks. Subsequently, perform duplication detection on the extracted attacker identification (such as IP address, device fingerprint), attack method (such as scanning, denial of service attack, SQL injection) and target information (such as affected server IP, open port), use hash algorithm to verify the uniqueness of each data, filter duplicate records, and ensure that each attack behavior data is independent and valid. Assume that in a certain period of time, the original attack records extracted from the log system are 2000, of which 1500 are left after hash deduplication, then the amount of attack behavior matching data after removing duplicate records is 1500.
[0069] S102: Based on the attack behavior matching data volume, the semantic relationship between the attacker identification, attack method and target is analyzed, the subject-predicate-object structure is extracted, the attack behavior field is mapped to the corresponding attack method, the attacker identification under the same attack method is merged, the number of attackers with multiple attack methods is counted, and the attack methods are classified according to the attack method number threshold to obtain the attack method distribution range;
[0070] First, the extracted attack behavior data is subjected to dependency syntactic analysis to identify the subject-predicate-object structure between the attacker, attack method, and target. For example, in "a certain IP address attempts to SQL injection attack the target server", the attacker is the IP address, the attack method is SQL injection, and the target is the server. The attack behavior field and the attack method classification are further mapped in combination with the network security knowledge base, and the attacker identification under the same attack method is merged. Multiple IP addresses for the same attack method are merged. For example, when counting SQL injection attack behaviors, if multiple attacker IPs from different subnet segments are found, association analysis is performed to determine whether they belong to the same attack organization, and thus they are classified into the same attacker identification. The number of attackers with multiple attack methods is counted, and the attacker behavior is classified by setting a threshold T (such as T=3). If the number of attack methods of an attacker exceeds the threshold T, it is classified as a "high-risk attacker", otherwise it is classified as a "low-risk attacker", and finally an attack method distribution interval is formed. For example, if an attacker IP involves 5 attack methods in total, it is classified as a high-risk attacker, while an attacker involving 2 attack methods is classified as a low-risk attacker.
[0071] S103: Call the attack method distribution interval, analyze the target type distribution by attacker identification classification, group the target IP, industry category and asset type, cross-match the attack methods and summarize the key attack methods, and generate a standardized threat behavior data set;
[0072] First, the target IP is classified and the industry category and asset type to which it belongs are extracted. For example, the threat intelligence database is queried to see whether the IP address belongs to the government, finance, energy, education and other industries, and its asset type (such as Web server, database server, industrial control system, etc.) is obtained. The IP address is grouped according to the interactive relationship between the attacker identification and the target category, and a mapping table of attack methods and target types is constructed. For example, if an IP address mainly carries out SQL injection attacks on database servers in the financial industry, then the IP address belongs to the "financial database attacker" category. The attack methods are cross-matched to screen out key attack methods. For example, in the attack method statistics process, if the servers in a certain industry are mainly attacked by DDoS, then DDoS is classified as the key attack method of the industry, and Web servers are generally attacked by cross-site scripting (XSS), then XSS becomes the key attack method of Web servers. Based on the statistical results, a standardized threat behavior data set is generated. The data set contains attacker identification, attack method, target category and attack frequency information. For example, in a certain period of time, an attacker launches a total of 50 XSS attacks on Web servers in the education industry, and his behavior is recorded in the data set to form a complete threat analysis report.
[0073] Specifically, if Figure 3 As shown in Figure 1, the steps of the attack event time series are as follows:
[0074] S201: based on the attack timestamps of the standardized threat behavior data set, the attack timestamps are arranged in chronological order, the time intervals between adjacent events are identified, the events that do not exceed the threshold are screened, and a time interval sequence of attack events is generated;
[0075] First, the attack event data set is parsed, the timestamp of each record is extracted, and converted into a unified format (such as Unix timestamp) to ensure consistent time accuracy. The timestamps of all attack events are arranged in ascending order to form a time series, and the time interval between adjacent events is identified, that is, the time difference T between two adjacent attack events is calculated. The time interval calculation formula is T=t(i+1)-t(i), where t(i) is the timestamp of the i-th attack event, and t(i+1) is the timestamp of the next adjacent attack event. For example, if the timestamp of an attack event is 17 00000000 seconds, the timestamp of the next attack event is 1700000300 seconds, then the time interval T = 300 seconds, filter out events that do not exceed the threshold ΔT and retain them in the time series. Assuming that the threshold ΔT is set to 600 seconds, all events with a time interval T ≤ 600 seconds are retained, and events that exceed the threshold are marked as independent events with too long intervals. For example, if the interval between adjacent events is 800 seconds, the event will not be included in the current time series. Generate an attack event time interval sequence to ensure that all selected events meet the time constraint.
[0076] S202: Based on the attack event time interval sequence, for events that exceed the time threshold, analyze the consistency of the attack target, extract the target IP, port number and protocol type, filter the qualified events and classify them into the same attack chain, count the number of events, time span and number of targets in the chain, and obtain attack chain statistics;
[0077] First, extract the attack events that exceed the threshold ΔT, obtain their corresponding target IP, port number and protocol type, and determine whether the target IP is consistent, that is, compare the target IP of the current event with the target IP of the previous event. If they are the same, continue the analysis, otherwise it is considered as a different attack chain and compare whether the port numbers are the same, that is, check whether the ports opened by the attacker to the target server are consistent.
[0078] S203: Call the attack chain statistical data, calculate the attack chain impact factor, divide the tracing path according to the number of events, time span and number of targets, and generate the attack event time series;
[0079] The attack chain impact factor uses the formula:
[0080] ;
[0081] in, Represents the attack chain impact factor, Represents the total number of attack events, Representative The severity coefficient of an attack event, Representative The time span of the attack events, Representative The number of targets affected by an attack event, Represents the average number of targets affected by the attack event, represents the total number of attack events;
[0082] Attack chain impact factor Used to quantify the combined impact of a series of attack events on the target system. This formula takes into account the severity, duration, and degree of deviation of the number of targets affected by each attack event.
[0083] Total number of attacks : Determined by the number of independent attack events detected by the security monitoring system within a specific period of time;
[0084] No. The severity coefficient of the attack event :Based on the type, means and impact of each attack event, a quantitative assessment is conducted with reference to the internationally accepted vulnerability scoring standards (such as CVSS) to derive the severity coefficient;
[0085] No. Time span of attack events : Calculate the duration of each attack event from the start to the end through log records, in hours;
[0086] No. The number of targets affected by the attack event : Count the number of independent targets (such as servers, workstations, etc.) directly affected by each attack event;
[0087] The average number of targets affected by the attack : Calculate the average number of targets affected by all attack events. The formula is: ;
[0088] Assume that three independent attacks are detected within a week ( ), whose parameters are as follows:
[0089] Attack 1: Severity Factor , time span Hours, number of affected targets ;
[0090] Attack 2: Severity Factor , time span Hours, number of affected targets ;
[0091] Attack 3: Severity Factor , time span Hours, number of affected targets ;
[0092] Calculate the average number of affected targets : ;
[0093] Calculate the numerator part:
[0094] ;
[0095] Calculate the denominator:
[0096] ;
[0097] Calculate the attack chain impact factor : ;
[0098] Calculation results It shows that within the monitored week, considering the severity, duration and impact scope of the attack incident, the comprehensive impact factor of the attack chain on the target system is 2.84. This value can be used to evaluate the current security situation and provide a quantitative basis for formulating defense strategies.
[0099] Specifically, if Figure 4 As shown in the figure, the steps for hierarchical threat event analysis results are as follows:
[0100] S301: Call the attack event time series, count the occurrence frequency of the attack target type in the differentiated events, classify them by target IP, port number and protocol type, and obtain attack target frequency distribution data;
[0101] First, the time series is sorted according to the timestamp to ensure that the time sequence is clear. Then, each attack record in the sequence is traversed in turn, and the key fields such as the target IP, port number and protocol type of the record are extracted. The field information is stored in a temporary data table. Each record is arranged according to the time when the attack event occurred. During the data storage process, the combinations of the same target IP, port number and protocol type are counted and accumulated to obtain the preliminary frequency distribution data of the target type in the entire attack event time series. In order to improve the accuracy of statistics, the target IP needs to be normalized during the storage process, such as distinguishing private IP addresses from public IP addresses, and standardizing the port numbers. For example, common ports and non-standard ports are counted separately. The statistics of protocol types also need to be classified in combination with common attack methods. For example, TCPSYNFlood attacks mainly target TCP protocols. The DNS amplification attack is mainly based on the UDP protocol. After the data accumulation reaches a certain scale, the stored data is cleaned and deduplicated. For example, records with the same target IP, port number and protocol type at adjacent time points are counted only once to prevent repeated attack events from causing data redundancy. At the same time, some extreme cases are handled. For example, if the number of attack events for a target IP exceeds 50% of the total number of events, the target IP is an abnormal case and can be eliminated or analyzed separately. To facilitate subsequent analysis, a hash map structure can be used to store frequency data during the statistical process. For example, (target IP, port number, protocol type) is used as the key and the number of occurrences is used as the value. After the data statistics are completed, a frequency distribution data table of the attack target is formed and classified according to different dimensions, such as the frequency distribution table according to the target IP, the frequency distribution table according to the port number, and the frequency distribution table according to the protocol type, and finally the frequency distribution data of the attack target is obtained.
[0102] S302: Calculate the target co-occurrence probability based on the attack target frequency distribution data, extract the target type combination in the differentiated event, count the number of times the combination occurs, calculate the event attribution weight according to the co-occurrence number and the total number of events according to the probability threshold, and generate the event attribution weight matrix;
[0103] First, extract the target type combination, arrange the target type according to different combinations of IP, port number, and protocol type, and count the number of occurrences of each combination in the entire event sequence. For each combination type, calculate its target co-occurrence probability. The specific calculation method is to divide the joint frequency by the frequency of occurrence of a single event. For example, assuming that the number of times the target IPA and port number 80 appear together is N(A, 80), and the number of times the target IPA appears alone is N(A), then its co-occurrence probability P(A, 80) = N(A, 80) / N(A). For different combinations, such as (target IP, port number), (target IP, protocol type), (port number, protocol type), etc., the co-occurrence probability is calculated in this way. In order to improve the accuracy of the data, it is necessary to eliminate low-frequency co-occurrence combinations. For example, set a co-occurrence threshold. If a combination If the number of occurrences is less than 5% of the total number of events, the combination is ignored. After obtaining the co-occurrence probability of all combinations, the event attribution weight is further calculated. The event attribution weight is calculated by dividing the number of co-occurrences by the total number of events. For example, the number of co-occurrences of a certain combination (target IPA, port number 80) in all events is M(A, 80), and the total number of events is N_total. Then the event attribution weight W(A, 80) = M(A, 80) / N_total. According to the set probability threshold, the combinations with attribution weights lower than the threshold are eliminated, and finally an event attribution weight matrix is formed. The matrix uses the target type as the row and column index and fills in the attribution weight value. For example, an element W(i, j) in the matrix represents the attribution weight between target type i and target type j. After the attribution weight matrix is formed, it can be used for subsequent event attribution analysis.
[0104] S303: calling the event attribution weight matrix, dividing the events according to the attribution weights, merging the events above the threshold, splitting the events below the threshold, and obtaining hierarchical threat event analysis results;
[0105] First, classify the attribution weight values in the matrix. Set an attribution weight threshold T, for example, T = 0.05. Merge the events in the matrix with attribution weights higher than T. The specific method is to traverse the rows and columns of the matrix, and regard all target type combinations that satisfy W(i, j)>T as the same category and group them into a whole. For example, if the target IPA and port number 80 have a high attribution weight, they are classified as the same category of events. For events with attribution weight values lower than T, they are split according to the weights. The splitting method is to split the event into independent target type units, and calculate the event attribution relationship for each unit independently. For example, if the target IPA is associated with port 22 and port 443, and W(A, 22)<T, W(A, 443)<T, then the individual attribution weights of A with 22 and A with 443 should be calculated separately, and independent events should be reconstructed. For some events, if the attribution weight of a certain specific combination still exceeds T after splitting, they are merged again. Finally, a hierarchical threat event analysis result is formed. This analysis result classifies different target types according to the co-occurrence probability and attribution weight, and generates threat events at different levels. The highest-level event represents the combination of attack target types with the highest co-occurrence degree, and the lowest-level event represents individual target types with lower attribution weights.
[0106] Specifically, as Figure 5 shown, the steps of the attack behavior path matching result are specifically as follows:
[0107] S401: Based on the hierarchical threat event analysis result, extract key behavior nodes, arrange the behavior nodes in chronological order, screen the nodes with offset features on the time axis, and calculate the time difference between adjacent nodes to obtain the key behavior time offset.
[0108] First, the behavior nodes in the event are extracted and arranged in chronological order, which involves parsing the event log and extracting the timestamps of node activities. For example, assuming that in a network intrusion event, the attacker first performs a port scan (node A), then performs privilege escalation (node B), and then steals data (node C). Each behavior has a clear time record. The program automatically extracts the timestamp and arranges the behavior in chronological order. Next, the behavior nodes are screened for time offset features. The specific method is to calculate the time difference between adjacent nodes, such as the time difference from node A to node B, and the time difference from node B to node C. The calculation can be obtained through simple time subtraction. If the time difference is significantly different, it means that there is an abnormal operation or an intentional delay. Nodes with significant offset characteristics are marked for further analysis. This marking can be based on threshold settings. For example, if the time difference exceeds 5 minutes, it is considered a significant offset. The key turning points in the attack behavior can be quickly identified, and the key behavior time offset is finally obtained. This quantity is the statistical result of the time difference of the key nodes. For example, in multiple simulated attacks, the calculated average time difference is 4 minutes, but in the actual attack event, the time difference of a key behavior node reaches 10 minutes, which can be identified as a key behavior time offset.
[0109] S402: Analyze the dependency relationship in the time series according to the time offset of the key behavior, identify the time sequence relationship between adjacent behavior nodes, calculate the dependency strength value between nodes according to the time interval and the triggering order between nodes, and select the behavior node combination with a dependency greater than a preset threshold to obtain a time dependency relationship data set;
[0110] Analyzing the dependency relationship in the time series requires the use of statistical methods to determine the time sequence and the dependency strength between nodes. Assuming that the time offset between nodes A, B, and C has been identified, it is now necessary to evaluate whether the dependency strength from A to B is significantly greater than that from B to C. The specific method is to use the time interval and the triggering order between nodes for calculation. This can be achieved by building a simple dependency model, such as defining dependency strength = time offset / average time offset. If the obtained ratio is greater than a certain set threshold (such as 1.5 times), the dependency strength is considered to be large. This numerical quantitative analysis can help determine the key path of the attack behavior, and the combination of behavior nodes with a dependency greater than the preset threshold is screened out. For example, in an actual case, the dependency strength between the attacker's privilege escalation behavior and the subsequent data theft behavior is calculated to be 1.8, which is significantly higher than the behavior combination. Such analysis helps security analysts focus on the most critical attack behavior combination and further obtain a time dependency relationship data set. This data set contains all behavior node combinations with dependency strength exceeding the threshold, providing data support for subsequent attack path analysis.
[0111] S403: Based on the time dependency relationship data set, select the behavior nodes with high correlation, cluster the nodes belonging to the same attack target, merge the nodes of the same type and divide the belonging paths, and generate the attack behavior path matching result;
[0112] Conduct correlation analysis and clustering of behavior nodes to identify and merge nodes belonging to the same attack target, including using clustering algorithms to group nodes in the data set. Suppose there is a group of behavior nodes that all point to illegal access to the same database server. The nodes include different methods of obtaining access rights, data theft behaviors, etc. Through clustering algorithms, behaviors can be classified into one category and then merged into one node. For example, using the K-means clustering method, the number of clusters is set to the number of expected attack targets. The clustering results show that all behaviors targeting the database server are classified into one category, and then the paths of the same type of nodes are divided. The generated attack behavior path matching results include the complete path from the initial permission acquisition to data theft. The path clearly indicates how the attacker gradually penetrates into the core part of the system and steals key data.
[0113] Specifically, if Figure 6 As shown in the figure, the steps for aggregating the threat results of network attack paths are as follows:
[0114] S501: calling the attack behavior path matching result, extracting the behavior node triggering frequency, counting the number of occurrences of the node in the differentiated path, and obtaining the behavior node triggering ratio;
[0115] The behavior node trigger frequency is calculated using the formula:
[0116] ;
[0117] in, Represents the behavior node trigger frequency, Representing behavior nodes In the The time interval between triggers, Representing behavior nodes The average time interval between triggers, Representing behavior nodes The total number of triggers in the path, Representing behavior nodes In the The path weight at the time of triggering, Representing behavior nodes In the The influence coefficient of the second trigger;
[0118] In this formula, Indicates the trigger frequency of the behavior node. The following steps are required to calculate this feature value:
[0119] Calculate the average time interval : Get behavior node The time interval between each trigger ,The time interval can be obtained by monitoring system logs or event records.,Assume that the node The time intervals in the three triggers are 2.5 seconds, 3.0 seconds, and 2.8 seconds respectively. The average time interval is calculated as follows:
[0120] ;
[0121] Calculate the average absolute deviation of the time interval: calculate the absolute difference between each trigger time interval and the average time interval, and calculate the average value;
[0122] ;
[0123] ;
[0124] Calculate the sum of the product of the path weight and the influence coefficient: For each trigger, obtain the corresponding path weight and influence coefficient , the parameters can be determined by analyzing the path importance of the system and the influence of the nodes. Assume that the path weights and influence coefficients of three triggers are as follows:
[0125] First trigger: , ;
[0126] Second trigger: , ;
[0127] The third trigger: , ;
[0128] The calculation is as follows:
[0129] ;
[0130] Calculate trigger feature value :Substitute the above results into the original formula:
[0131] ;
[0132] The result shows that the triggering frequency of the behavior node is 0.31. The value reflects the triggering frequency and influence of the node in different paths, which helps to evaluate its importance in the system.
[0133] S502: Calculate the node weight value according to the behavior node trigger ratio, combine the trigger frequency and path distribution to count the node contribution, sort the event priorities in descending order according to the weight value, select the paths with high weights and classify them into emergency response, and obtain the key threat path;
[0134] Combined with the trigger frequency of the node and the distribution of the path, in order to calculate the weight, we must first evaluate the scope of influence of each node in different attack paths, that is, the frequency of the node and the contribution of the path to the overall network security threat. The frequency of the node can be combined with the path distribution by weighted average to assign a weight value to each node. Assuming that the trigger ratio of a node is 0.4, and the criticality in the path is high (for example, it is the core step for the attacker to achieve the final goal), then the weight of the node is rated as high, such as 0.7; if the trigger ratio of another node is relatively low, and the role in the path is relatively minor (such as only used to assist other attack steps), then its weight is low at 0.3. After the weight is calculated, the system will sort the priority of the nodes in descending order according to the weight value, and filter the paths corresponding to the high-weight nodes as emergency response paths. The path is the core part of the network attack, has a high threat level, and requires immediate response. The system selects the path with a higher weight value for emergency response based on the node weight, and finally obtains a set of critical threat paths.
[0135] S503: Based on the key threat path, the threat events under the emergency response path are aggregated, the impact scope and the number of attack targets are counted, and a network attack path threat aggregation result is generated;
[0136] After the critical threat path is derived, the system will collect threat events in the emergency response path. First, all relevant attack events in the path need to be comprehensively collected and classified. The process requires in-depth analysis of all threat events in each critical path, and statistics on the scope of impact and the number of attack targets. The scope of impact refers to the degree of spread of each threat event in the network, and the number of attack targets refers to the number of devices or systems affected by the path or event. Assuming that an emergency response path involves multiple threat events such as "denial of service attack" and "remote code execution", the system will evaluate the impact of the event on each node in the network and count the number of attack targets. For example, if the path involves an attack on 50 servers, and each server is affected to varying degrees, the scope of impact is 50, and the number of attack targets is also 50. The system generates a threat collection result for the network attack path, providing the security team with a more accurate threat map so that targeted measures can be taken to deal with the impending network attack.
[0137] like Figure 7 As shown, based on the network security threat information collection system, the system includes:
[0138] The attack behavior identification module extracts the attacker identification, attack method, and attack target based on network security threat information, calculates the corresponding offset of the attack time, compares the cross-degree of the attack target, screens the target attack behavior, summarizes the characteristic combination of the attack mode, and obtains the attack behavior mode;
[0139] The attack chain construction module is based on the attack behavior pattern, compares the cross frequency of the attack target, screens the attack behavior that meets the time window, analyzes the influence relationship of the differentiated nodes, summarizes the attribution of multiple attack behaviors, adjusts the transmission path of the attack chain, screens the chain with large influence factors, and obtains the attack chain structure information;
[0140] Based on the attack chain structure information, the threat event attribution module calculates the crossover frequency of attack targets in the chain, analyzes the co-occurrence relationship of targets, identifies the attack event attribution boundary, filters out attack events with conflicting attributions, adjusts the chain attribution of unevenly distributed targets, and obtains the attack event attribution mapping;
[0141] The attack path reconstruction module extracts the core behavior nodes of the path based on the attribution mapping of the attack events, adjusts the attack path sequence, compares the behavior correlation, summarizes the logically similar paths, screens the key path derivative relations, and obtains the attack behavior correlation path;
[0142] The threat impact assessment module extracts the attack node triggering frequency based on the attack behavior association path, calculates the impact diffusion value, screens the attack nodes with a large impact range, summarizes the high-risk attack paths, and obtains the network attack path threat aggregation results.
[0143] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. Based on the network security threat information collection method, it is characterized by: The following steps are involved: S1: Based on network security threat information, extract network attackers and attack methods, identify and classify attacker identification and target type, standardize the data, and obtain a standardized threat behavior data set; S2: Based on the standardized threat behavior dataset, extract attack timestamps, arrange them in chronological order, identify the time intervals between adjacent attack events, determine whether the events are continuous, filter related events and classify them into the same attack chain, calculate the impact factor of the attack chain, divide the tracing path, and obtain the attack event time series; S3: calling the attack event time series, counting the frequency of the attack target type in the differentiated events, evaluating the event attribution weight, determining whether the events should be merged or split, and obtaining a hierarchical threat event analysis result; S4: Based on the hierarchical threat event analysis results, extract key behavior nodes, analyze time series dependencies, filter out related behaviors and classify them into similar attack paths, and obtain attack behavior path matching results; The steps of attack behavior path matching results are specifically as follows: S401: extracting key behavior nodes based on the hierarchical threat event analysis results, arranging the behavior nodes in chronological order, screening nodes with offset characteristics on the time axis, calculating the time difference between adjacent nodes, and obtaining the key behavior time offset; S402: Analyze the dependency relationship in the time series according to the key behavior time offset, identify the time sequence relationship between adjacent behavior nodes, calculate the dependency strength value between nodes according to the time interval and the triggering order between nodes, and select the behavior node combination with a dependency greater than a preset threshold to obtain a time dependency relationship data set; S403: Based on the time dependency relationship data set, select behavior nodes with high correlation, cluster nodes belonging to the same attack target, merge nodes of the same type and divide the belonging paths, and generate attack behavior path matching results; S5: Call the attack behavior path matching result, calculate the node weight value, sort the event priority according to the weight value, filter the weight path and classify it into emergency response, and obtain the network attack path threat aggregation result.
2. The method for collecting network security threat information according to claim 1 is characterized in that: The standardized threat behavior data set includes attacker identification, target type, and subject-predicate-object structure parsing results; the attack event time series includes attack timestamp, event continuity determination results, and attack chain influencing factors; the hierarchical threat event analysis results include event attribution weights, target co-occurrence probabilities, and event hierarchical classifications; the attack behavior path matching results include behavior time offsets, dependency analysis results, and correlation behavior classifications; the network attack path threat aggregation results include event priorities, node weight values, and emergency response paths.
3. The method for collecting network security threat information according to claim 1 is characterized in that: The steps of standardizing the threat behavior dataset are as follows: S101: Based on network security threat information, including security event logs and intrusion detection alarms, extract attacker identification, attack methods and target information, remove duplicate records, and obtain attack behavior matching data volume; S102: Based on the attack behavior matching data volume, the semantic relationship between the attacker identification, attack method and target is analyzed, the subject-predicate-object structure is extracted, the attack behavior field is mapped to the corresponding attack method, the attacker identification under the same attack method is merged, the number of attackers with multiple attack methods is counted, and the attack methods are classified according to the attack method number threshold to obtain the attack method distribution range; S103: Call the attack mode distribution interval, analyze the target type distribution by attacker identification classification, group the target IP, industry category and asset type, cross-match the attack modes and summarize the key attack modes, and generate a standardized threat behavior data set.
4. The method for collecting network security threat information according to claim 1, characterized in that: The steps of the attack event time series are specifically as follows: S201: Arrange the attack timestamps of the standardized threat behavior data set in chronological order, identify the time intervals between adjacent events, filter out events that do not exceed a threshold, and generate an attack event time interval sequence; S202: Based on the attack event time interval sequence, for events that exceed the time threshold, analyze the consistency of attack targets, extract the target IP, port number and protocol type, filter qualified events and classify them into the same attack chain, count the number of events, time span and number of targets in the chain, and obtain attack chain statistics; S203: Call the attack chain statistical data, calculate the attack chain impact factor, divide the tracing path according to the number of events, time span and target number, and generate the attack event time series.
5. The method for collecting network security threat information according to claim 4 is characterized in that: The attack chain impact factor adopts the formula: ; in, Represents the attack chain impact factor, Represents the total number of attack events, Representative The severity coefficient of the attack incident Representative The time span of the attack events, Representative The number of targets affected by an attack event, represents the average number of targets affected by an attack event, Indicates the total number of attack events.
6. The method for collecting network security threat information according to claim 1 is characterized in that: The steps of hierarchical threat event analysis results are specifically as follows: S301: calling the attack event time series, counting the occurrence frequency of attack target types in differentiated events, classifying them by target IP, port number and protocol type, and obtaining attack target frequency distribution data; S302: Based on the attack target frequency distribution data, the target co-occurrence probability is calculated, the target type combination in the differentiated event is extracted, the number of occurrences of the combination is counted, and the event attribution weight is calculated according to the co-occurrence number and the total number of events according to the probability threshold, and an event attribution weight matrix is generated; S303: calling the event attribution weight matrix, dividing the events according to the attribution weights, merging the events above the threshold, splitting the events below the threshold, and obtaining hierarchical threat event analysis results.
7. The method for collecting network security threat information according to claim 1 is characterized in that: The steps of collecting the threat results of the network attack path are specifically as follows: S501: calling the attack behavior path matching result, extracting the behavior node triggering frequency, counting the number of occurrences of the node in the differentiated path, and obtaining the behavior node triggering ratio; S502: Calculate the node weight value according to the behavior node trigger ratio, combine the trigger frequency and path distribution to count the node contribution, arrange the event priorities in descending order according to the weight value, select the paths with high weights and classify them into emergency response, and obtain the key threat path; S503: Based on the key threat path, threat events under the emergency response path are aggregated, the impact scope and the number of attack targets are counted, and a network attack path threat aggregation result is generated.
8. The method for collecting network security threat information according to claim 7 is characterized in that: The behavior node trigger frequency is calculated using the formula: ; in, Represents the behavior node trigger frequency, Representing behavior nodes In the The time interval between triggers, Representing behavior nodes The average time interval between triggers, Representing behavior nodes The total number of triggers in the path, Representing behavior nodes In the The path weight at the time of triggering, Representing behavior nodes In the The influence coefficient of the trigger.
9. Based on the network security threat information collection system, it is characterized by: According to the method for collecting network security threat information according to any one of claims 1 to 8, the system comprises: The attack behavior identification module extracts the attacker identification, attack method, and attack target based on network security threat information, calculates the corresponding offset of the attack time, compares the cross-degree of the attack target, screens the target attack behavior, summarizes the characteristic combination of the attack mode, and obtains the attack behavior mode; Based on the attack behavior pattern, the attack chain construction module compares the cross frequency of the attack target, screens the attack behaviors that meet the time window, analyzes the influence relationship of the differentiated nodes, summarizes the attribution of multiple attack behaviors, adjusts the transmission path of the attack chain, screens the chain with large influence factors, and obtains the attack chain structure information; The threat event attribution module calculates the crossover frequency of attack targets in the chain based on the attack chain structure information, analyzes the co-occurrence relationship of the targets, identifies the attack event attribution boundary, screens attack events with attribution conflicts, adjusts the chain attribution of unevenly distributed targets, and obtains the attack event attribution mapping; The attack path reconstruction module extracts the core behavior nodes of the path based on the attack event attribution mapping, adjusts the attack path sequence, compares the behavior correlation, summarizes the logically similar paths, screens the key path derivative relationship, and obtains the attack behavior correlation path; The threat impact assessment module extracts the attack node triggering frequency based on the attack behavior association path, calculates the impact diffusion value, screens the attack nodes with a large impact range, summarizes the high-risk attack paths, and obtains the network attack path threat aggregation result.
Citation Information
Patent Citations
Attack tracing method and device based on log association analysis
CN114615063A
Multi-stage attack process reconstruction method based on causal diagram
CN115499169A