Method for processing multi-source data, alarm analysis method, device and equipment
By preprocessing and merging multiple data sources, and combining the analysis of the vulnerability detection database and the threat data process chain rule base, the problems of large analysis workload and misjudgment in multi-source data processing are solved, thereby improving the efficiency and accuracy of security threat detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QI-ANXIN LEGENDSEC INFORMATION TECH (BEIJING) INC
- Filing Date
- 2022-09-05
- Publication Date
- 2026-04-14
AI Technical Summary
Existing technologies, when processing multi-source data, involve a large amount of analysis work and are prone to misjudgment and omission of security vulnerabilities in the network, resulting in low efficiency and low accuracy in security threat detection.
By preprocessing, normalizing, and parsing data from multiple data sources and matching it with a compromise detection database, threat alert information is obtained. Based on asset information, the data is merged and analyzed using a pre-defined detection strategy and a threat data process chain rule base to determine attack events and threat levels.
It enables simple, efficient, and accurate processing of multi-source data, reduces the number of alarms, improves the efficiency and accuracy of security threat detection, and provides convenient means for alarm analysis and threat discovery.
Smart Images

Figure CN115714662B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of security threat technology, and in particular to a method, apparatus and device for processing multi-source data, as well as an alarm analysis method, apparatus and device. Background Technology
[0002] Various entities (such as companies) generate a great deal of data, and analyzing this data is a crucial task in the field of cybersecurity. In cybersecurity, an entity that generates data can be called a data source (or intelligence source), and each piece of data generated by a data source can be called a piece of intelligence. Intelligence can be recorded through logs. Typically, entities incorporate data from multiple data sources into their security protection network through self-production or procurement. By analyzing and studying the data within this security protection network, they support cybersecurity efforts.
[0003] Because the data quality varies greatly and the amount of data from different data sources is large, the current methods for discovering security vulnerabilities in the network based on massive and complex multi-source data involve a large workload for staff to analyze and study, and are prone to misjudging and overlooking attack behaviors.
[0004] Therefore, there is an urgent need to provide a simple, efficient, and accurate method for processing multi-source data to obtain data that is informative, intuitive, and easy to understand, thereby facilitating alarm analysis and threat discovery. Summary of the Invention
[0005] This application provides a method for processing multi-source data, an alarm analysis method, an apparatus, and a device that can quickly and accurately merge and organize multi-source data, effectively reducing the number of alarms and the workload of alarm analysis. In addition, it performs alarm analysis based on reasonable strategies, thereby improving the efficiency and accuracy of security threat detection.
[0006] In a first aspect, this application provides a method for processing multi-source data. This method may include, for example,: preprocessing data from multiple data sources to obtain data to be processed; obtaining threat alarm information based on the data to be processed and a vulnerability detection database, wherein the vulnerability detection database includes vulnerability detection information, and the threat alarm information is a fusion result of the data to be processed and the vulnerability detection information matched in the vulnerability detection database; and merging the threat alarm information based on the asset information in the threat alarm information to obtain actual alarm information.
[0007] The asset information may include, but is not limited to, at least one of the following: Internet Protocol (IP) address or device identification information. For example, asset information may be an asset IP address or a device identifier; or, for example, asset information may include an asset IP address and an IOC (Indicator Code).
[0008] Optionally, for the first and second threat alarm information, the step of merging the threat alarm information based on the asset information to obtain the actual alarm information includes:
[0009] If the first asset information of the first alarm information matches the second asset information of the second alarm information, then the first alarm information and the second alarm information are merged to obtain the third alarm information, and the actual alarm information includes the third alarm information;
[0010] If the first asset information of the first alarm information and the second asset information of the second alarm information do not match, then the first alarm information and the second alarm information are created as the actual alarm information, which includes the first alarm information and the second alarm information.
[0011] Optionally, the step of merging the threat alarm information based on the asset information to obtain the actual alarm information includes:
[0012] The time window value is determined based on the amount of data to be processed, the merging processing rate, and the data collision coefficient. The data collision coefficient is determined based on the number of threat alarm messages and the data to be processed.
[0013] Based on the time window value, the step of "merging the threat alarm information based on the asset information of the threat alarm information to obtain the actual alarm information" is executed.
[0014] Optionally, the preprocessing of data from multiple data sources to obtain data to be processed includes:
[0015] The multi-source data is normalized to obtain the data to be processed, wherein each data in the data to be processed has the same format.
[0016] Optionally, the preprocessing of data from multiple data sources to obtain data to be processed includes:
[0017] The multi-source data is normalized to obtain intermediate processing data, wherein each data point in the intermediate processing data has the same format.
[0018] The intermediate processing data is parsed to obtain the data to be processed. The data to be processed is obtained by extracting key information from the intermediate processing data. The key information includes information to be inspected and data details. The information to be inspected includes at least one of the following: IP address, domain name or domain, or uniform resource locator (URL).
[0019] Optionally, the method further includes:
[0020] The actual alarm information is periodically detected according to a preset detection strategy to identify attack events.
[0021] Optionally,
[0022] The step of periodically detecting the actual alarm information according to a preset detection strategy to determine attack events includes:
[0023] If the alarm frequency of the actual alarm information matches the preset first heartbeat range within a cycle, then the attack event is determined as the first attack event according to the first heartbeat range and the saved first correspondence relationship. The first correspondence relationship includes the correspondence relationship between the first heartbeat range and the first attack event.
[0024] or,
[0025] The step of periodically detecting the actual alarm information according to a preset detection strategy to determine attack events includes:
[0026] Within a cycle, predict the attack type based on actual alarm information;
[0027] The monitoring time range is determined based on the attack type and the saved second correspondence, wherein the second correspondence includes the correspondence between the attack type and the monitoring time range;
[0028] Based on the analysis of the actual alarm information occurring on the external connection ports within the monitoring time range, the attack event is determined to be the first attack event, and the actual alarm information matches the pattern of the external connection ports where the first attack event occurred.
[0029] Optionally, the method further includes:
[0030] The first process chain that obtains the actual alarm information;
[0031] Based on the first process chain and the threat data process chain rule base, the first attack event corresponding to the actual alarm information is determined, and the threat data process chain rule base includes the correspondence between the first process chain and the first attack event.
[0032] Optionally, the method further includes:
[0033] Based on the first process chain and the threat data process chain rule base, the first threat level corresponding to the actual alarm information is determined. The threat data process chain rule base includes the correspondence between the first process chain, the first attack event, and the first threat level.
[0034] Secondly, this application also provides a method for alarm analysis, including:
[0035] Within a cycle, alarm information is obtained;
[0036] Based on the alarm information and the preset detection strategy, the alarm information is periodically detected to determine the attack event of the alarm information. The detection strategy is used to indicate the attack type or the characteristics that the attack event meets.
[0037] Optionally, the step of periodically detecting the alarm information based on the alarm information and a preset detection strategy to determine the attack event of the alarm information includes:
[0038] If the alarm frequency of the alarm information matches a preset first heartbeat range, then the attack event is determined to be a first attack event based on the first heartbeat range and a saved first correspondence. The first correspondence includes the correspondence between the first heartbeat range and the first attack event.
[0039] Optionally, the method further includes:
[0040] If the alarm frequency of the alarm information does not match the heartbeat range in all the stored correspondences, then it is determined that no attack event has occurred.
[0041] Optionally, the step of periodically detecting the alarm information based on the alarm information and a preset detection strategy to determine the attack event of the alarm information includes:
[0042] Predict the attack type based on the alarm information;
[0043] The monitoring time range is determined based on the attack type and the saved second correspondence, wherein the second correspondence includes the correspondence between the attack type and the monitoring time range;
[0044] Based on the analysis of the actual alarm information occurring on the external connection ports within the monitoring time range, the attack event is determined to be the first attack event, and the actual alarm information matches the pattern of the external connection ports where the first attack event occurred.
[0045] Optionally, the method further includes:
[0046] The first process chain that obtains the alarm information;
[0047] Based on the first process chain and the threat data process chain rule base, the second attack event corresponding to the alarm information is determined, and the threat data process chain rule base includes the correspondence between the first process chain and the second attack event.
[0048] Optionally, the method further includes:
[0049] Based on the first process chain and the threat data process chain rule base, the first threat level corresponding to the alarm information is determined. The threat data process chain rule base includes the correspondence between the first process chain, the second attack event, and the first threat level.
[0050] Optionally, the method further includes:
[0051] The second attack event is verified based on the first attack event to determine the alarm result.
[0052] Optionally, the alarm information is the actual alarm information in the method provided in the first aspect above.
[0053] Thirdly, this application also provides a method for alarm analysis, including:
[0054] The first process chain for obtaining alarm information;
[0055] Based on the first process chain and the threat data process chain rule base, the first attack event corresponding to the alarm information is determined, and the threat data process chain rule base includes the correspondence between the first process chain and the first attack event.
[0056] Optionally, the method further includes:
[0057] Based on the first process chain and the threat data process chain rule base, the first threat level corresponding to the alarm information is determined. The threat data process chain rule base includes the correspondence between the first process chain, the first attack event, and the first threat level.
[0058] Optionally, the first process chain for obtaining alarm information includes:
[0059] The process information for obtaining the alarm information includes the process identifier of the alarm information, the parent process identifier, the process command line, the parent process command line, or the process hash value.
[0060] The process information is subjected to correlation analysis (e.g., analysis engine calculation) to obtain the first process chain.
[0061] Optionally, the alarm information can be a process event log from any of the following sources: Endpoint Detection and Response (EDR) system, system monitor (Sysmon), or host security. This can be understood as obtaining the alarm information by accessing process event logs enhanced by the EDR system, Sysmon, or host security.
[0062] Optionally, the alarm information is the actual alarm information or data to be processed in the method provided in the first aspect above.
[0063] Fourthly, this application also provides a multi-source data processing apparatus, comprising:
[0064] The preprocessing unit is used to preprocess data from multiple data sources to obtain data to be processed;
[0065] The collision unit is used to obtain threat alarm information based on the data to be processed and the vulnerability detection database. The vulnerability detection database includes vulnerability detection information, and the threat alarm information is the fusion result of the data to be processed and the vulnerability detection information matched in the vulnerability detection database.
[0066] The merging unit is used to merge the threat alarm information based on the asset information of the threat alarm information to obtain the actual alarm information.
[0067] The asset information may include at least one of the following: IP address or device identification information, etc.
[0068] Optionally, for the first alarm information and the second alarm information in the threat alarm information, the merging unit is specifically used for:
[0069] If the first asset information of the first alarm information matches the second asset information of the second alarm information, then the first alarm information and the second alarm information are merged to obtain the third alarm information, and the actual alarm information includes the third alarm information;
[0070] If the first asset information of the first alarm information and the second asset information of the second alarm information do not match, then the first alarm information and the second alarm information are created as the actual alarm information, which includes the first alarm information and the second alarm information.
[0071] Optionally, the merging unit includes:
[0072] A sub-unit is defined to determine a time window value based on the amount of data to be processed, the merging processing rate, and the data collision coefficient. The data collision coefficient is determined based on the number of threat alarm messages and the data to be processed.
[0073] The merging subunit is used to perform the step of "merging the threat alarm information based on the asset information of the threat alarm information to obtain the actual alarm information" according to the time window value.
[0074] Optionally, the preprocessing unit is specifically used for:
[0075] The multi-source data is normalized to obtain the data to be processed, wherein each data in the data to be processed has the same format.
[0076] Optionally, the preprocessing unit is specifically used for:
[0077] The multi-source data is normalized to obtain intermediate processing data, wherein each data point in the intermediate processing data has the same format.
[0078] The intermediate processing data is parsed to obtain the data to be processed. The data to be processed is obtained by extracting key information from the intermediate processing data. The key information includes information to be inspected and data details. The information to be inspected includes at least one of the following: IP address, Domain, or URL.
[0079] Optionally, the device further includes: a detection unit,
[0080] The detection unit is used to periodically detect the actual alarm information according to a preset detection strategy to determine the attack event.
[0081] Optionally, the detection unit is specifically used for:
[0082] If the alarm frequency of the actual alarm information matches the preset first heartbeat range within a cycle, then the attack event is determined as the first attack event according to the first heartbeat range and the saved first correspondence relationship. The first correspondence relationship includes the correspondence relationship between the first heartbeat range and the first attack event.
[0083] Alternatively, the detection unit is specifically used for:
[0084] Within a cycle, predict the attack type based on actual alarm information;
[0085] The monitoring time range is determined based on the attack type and the saved second correspondence, wherein the second correspondence includes the correspondence between the attack type and the monitoring time range;
[0086] Based on the analysis of the actual alarm information occurring on the external connection ports within the monitoring time range, the attack event is determined to be the first attack event, and the actual alarm information matches the pattern of the external connection ports where the first attack event occurred.
[0087] Optionally, the apparatus further includes: an acquisition unit and a detection unit.
[0088] The obtaining unit is used to obtain the first process chain of the actual alarm information;
[0089] The detection unit is used to determine the first attack event corresponding to the actual alarm information based on the first process chain and the threat data process chain rule base, wherein the threat data process chain rule base includes the correspondence between the first process chain and the first attack event.
[0090] Optionally, the detection unit is further configured to determine the first threat level corresponding to the actual alarm information based on the first process chain and the threat data process chain rule base, wherein the threat data process chain rule base includes the correspondence between the first process chain, the first attack event and the first threat level.
[0091] It should be noted that the specific implementation method and the effects achieved by the device provided in the fourth aspect can be found in the description of the relevant embodiments of the method described in the first aspect.
[0092] Fifthly, this application also provides an alarm analysis apparatus, comprising:
[0093] The first acquisition unit is used to acquire alarm information within one cycle;
[0094] The first detection unit is used to periodically detect the alarm information based on the alarm information and a preset detection strategy to determine the attack event of the alarm information. The detection strategy is used to indicate the attack type or the characteristics that the attack event meets.
[0095] Optionally, the first detection unit is specifically used for:
[0096] If the alarm frequency of the alarm information matches a preset first heartbeat range, then the attack event is determined to be a first attack event based on the first heartbeat range and a saved first correspondence. The first correspondence includes the correspondence between the first heartbeat range and the first attack event.
[0097] Optionally, the first detection unit is further configured to:
[0098] If the alarm frequency of the alarm information does not match the heartbeat range in all the stored correspondences, then it is determined that no attack event has occurred.
[0099] Optionally, the first detection unit further includes:
[0100] The prediction subunit is used to predict the attack type based on the alarm information;
[0101] A subunit is defined to determine a monitoring time range based on the attack type and a stored second correspondence, wherein the second correspondence includes the correspondence between the attack type and the monitoring time range.
[0102] The analysis subunit is used to determine the attack event as the first attack event based on the analysis of the actual alarm information occurring on the external ports within the monitoring time range, and the actual alarm information is matched with the pattern of the external ports where the first attack event occurred.
[0103] Optionally, the device further includes:
[0104] The second obtaining unit is used to obtain the first process chain of the alarm information;
[0105] The second detection unit is used to determine the second attack event corresponding to the alarm information based on the first process chain and the threat data process chain rule base, wherein the threat data process chain rule base includes the correspondence between the first process chain and the second attack event.
[0106] Optionally, the second detection unit is further configured to:
[0107] Based on the first process chain and the threat data process chain rule base, the first threat level corresponding to the alarm information is determined. The threat data process chain rule base includes the correspondence between the first process chain, the second attack event, and the first threat level.
[0108] Optionally, the device further includes: a verification unit.
[0109] The verification unit is used to verify the second attack event based on the first attack event and determine the alarm result.
[0110] Optionally, the alarm information is the actual alarm information in the device provided in the fourth aspect above.
[0111] It should be noted that the specific implementation method and the effects achieved by the device provided in the fifth aspect can be found in the description of the relevant embodiments of the method described in the second aspect.
[0112] Sixthly, this application also provides an alarm analysis apparatus, comprising:
[0113] The acquisition unit is the first process chain used to acquire alarm information.
[0114] The matching unit is used to determine the first attack event corresponding to the alarm information based on the first process chain and the threat data process chain rule base, wherein the threat data process chain rule base includes the correspondence between the first process chain and the first attack event.
[0115] Optionally, the matching unit is further configured to:
[0116] Based on the first process chain and the threat data process chain rule base, the first threat level corresponding to the alarm information is determined. The threat data process chain rule base includes the correspondence between the first process chain, the first attack event, and the first threat level.
[0117] Optionally, the obtaining unit is specifically used for:
[0118] The process information for obtaining the alarm information includes the process identifier of the alarm information, the parent process identifier, the process command line, the parent process command line, or the process hash value.
[0119] The process information is subjected to correlation analysis (e.g., analysis engine calculation) to obtain the first process chain.
[0120] Optionally, the alarm information is a process event log from any of the following sources: EDR system, Sysmon, or host security.
[0121] Optionally, the alarm information is the actual alarm information or data to be processed in the device provided in the fourth aspect above.
[0122] It should be noted that the specific implementation method and the effects achieved by the device provided in the sixth aspect can be found in the description of the relevant embodiments of the method described in the third aspect.
[0123] Seventhly, embodiments of this application also provide an electronic device, the electronic device including a processor and a memory:
[0124] The memory is used to store computer programs;
[0125] The processor is configured to execute the methods provided in the first, second, or third aspects described above, according to the computer program.
[0126] Eighthly, embodiments of this application also provide a computer-readable storage medium for storing a computer program for performing the methods provided in the first, second, or third aspects described above.
[0127] Therefore, this application has the following beneficial effects:
[0128] This application provides a method for processing multi-source data. This method may include, for example, preprocessing data from multiple data sources to obtain data to be processed; obtaining threat alarm information based on the data to be processed and a vulnerability detection database, wherein the vulnerability detection database includes vulnerability detection information, and the threat alarm information is the fusion result of the data to be processed and the vulnerability detection information matched in the vulnerability detection database; and merging the threat alarm information based on asset information to obtain actual alarm information. As can be seen, the method provided in this application preprocesses the accessed multi-source data, ensuring that the obtained data to be processed meets the requirements for unified processing. Through matching the data to be processed with a known vulnerability detection database and merging asset dimensions, fewer actual alarms are obtained than the data to be processed. This not only achieves simple, efficient, and accurate processing of multi-source data but also effectively reduces the number of alarms, facilitating subsequent alarm analysis and threat discovery based on reasonable alarm analysis strategies.
[0129] Furthermore, in the alarm analysis method provided in this application, alarm information obtained within a period is periodically detected based on the alarm information and a preset detection strategy to determine the attack events associated with the alarm information. The detection strategy is used to indicate the attack type or characteristics that the attack event conforms to. In this way, by analyzing and summarizing the periodic correlation characteristics of already occurred attack events through an alarm analysis device to obtain a detection strategy, and periodically detecting the alarm information to be detected within each period based on the detection strategy, the purpose of alarm analysis based on a reasonable detection strategy is achieved, thereby improving the efficiency and accuracy of security threat detection. The alarm information to be detected can be actual alarm information obtained based on the multi-source data processing method provided in this application, enabling reasonable analysis of massive and complex alarms from multiple sources.
[0130] Furthermore, this application also provides another method for alarm analysis. In this method, firstly, a first process chain of alarm information is obtained. Then, based on the first process chain and a threat data process chain rule base, a first attack event corresponding to the alarm information is determined. The threat data process chain rule base includes at least the correspondence between the first process chain and the first attack event. It is evident that this application achieves precise alarm analysis based on process chains by matching the process chain generated from the alarm information to be analyzed with the threat data process chain rule base in real time, and deriving the corresponding attack event based on the matched process chain in the threat data process chain rule base. This improves the efficiency and accuracy of security threat detection. The alarm information to be detected can be data to be processed or actual alarm information obtained based on the multi-source data processing method provided in this application, enabling reasonable analysis of massive and complex alarms from multiple sources. Moreover, the method provided in this application can mutually verify detection methods based on the periodic correlation features of alarm information, providing a solid foundation for real-time alarms and accurate location of intrusion detection. Attached Figure Description
[0131] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0132] Figure 1 A flowchart illustrating a multi-data source processing method provided in an embodiment of this application;
[0133] Figure 2 A flowchart illustrating the creation of log parsing rules is provided for an embodiment of this application;
[0134] Figure 3 A flowchart illustrating the selection of an orchestration plugin is provided for an embodiment of this application;
[0135] Figure 4 This application provides a schematic diagram of a preprocessing flow for multi-source data, as illustrated in an embodiment of the present application.
[0136] Figure 5 A flowchart illustrating an alarm information collision provided in an embodiment of this application;
[0137] Figure 6 A schematic diagram illustrating the process of merging alarm information provided in an embodiment of this application;
[0138] Figure 7 A flowchart illustrating an alarm analysis method provided in an embodiment of this application;
[0139] Figure 8 A flowchart illustrating another alarm analysis method provided in this application embodiment;
[0140] Figure 9 A schematic diagram of a multi-data source processing device provided in an embodiment of this application;
[0141] Figure 10 A schematic diagram of the structure of an alarm analysis device provided in an embodiment of this application;
[0142] Figure 11 A schematic diagram of another alarm analysis device provided in an embodiment of this application;
[0143] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0144] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the embodiments of this application will be further described in detail below with reference to the accompanying drawings and specific implementation methods. It should be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to this application are shown in the accompanying drawings, not the entire structure.
[0145] With the rapid development of the internet, cybersecurity issues are becoming increasingly serious. Faced with increasingly complex network structures and threat sources, cyber threats are becoming more diverse and destructive. Traditional security tools (such as Snort and firewalls) analyze individual log entries based on rules, which cannot capture complete attack behaviors, often resulting in false positives and false negatives. Furthermore, detecting compromised devices typically involves installing antivirus software, but current antivirus software is easily bypassed. For example, some antivirus software can be made undetectable by exploiting vulnerabilities, or powerful malware can directly disable antivirus software, rendering it ineffective. Additionally, malware can use rootkits to hide its processes and ports, preventing antivirus detection.
[0146] As the requirements for accuracy and quality in security threat detection capabilities continue to increase, the production, operation, and sharing of databases (i.e., intelligence repositories) have become a fundamental capability for companies. Companies increasingly incorporate open-source threat data into their security protection networks through self-production and procurement, relying on multi-source data to provide a richer data foundation for security threat detection. However, multi-source data comes from different companies, and the formats and quality of data from different companies vary, posing challenges to the analysis work of staff. To ensure that the data used for security threat detection is more information-rich, it is necessary to comprehensively consider multi-source data during the detection process. It should be noted that the data mentioned in the embodiments of this application refers to data that can be detected by security threats; the data source is the entity providing the data (such as a company), and the database is the carrier storing the data.
[0147] Currently, the methods for discovering security vulnerabilities in networks based on massive and complex multi-source data not only involve a large workload for staff in analysis and research, but also easily lead to misjudgment and omission of attack behaviors.
[0148] Based on this, embodiments of this application provide a method for processing multi-source data, which can perform simple, efficient, and accurate processing on multi-source data to obtain data that is information-rich, intuitively presented, and simple, thus facilitating alarm analysis and threat discovery. Specifically, the multi-source data processing method provided in this application may include, for example: a multi-source data processing device preprocesses data from multiple data sources to obtain data to be processed; the multi-source data processing device obtains threat alarm information based on the data to be processed and a vulnerability detection database, wherein the vulnerability detection database includes vulnerability detection information, and the threat alarm information is the fusion result of the data to be processed and the vulnerability detection information matched in the vulnerability detection database; the multi-source data processing device merges the threat alarm information based on the asset information in the threat alarm information to obtain actual alarm information.
[0149] As can be seen, the multi-source data processing method provided in this application preprocesses the accessed multi-source data so that the obtained data to be processed meets the requirements for unified processing. Thus, through matching the data to be processed with the known vulnerability detection database and merging the asset dimensions, the actual alarm information is less than the data to be processed. This not only achieves simple, efficient and accurate processing of multi-source data, but also effectively reduces the number of alarms, providing convenience for subsequent alarm analysis and threat discovery based on reasonable alarm analysis strategies.
[0150] It should be noted that the main body implementing this multi-source data processing method can be the multi-source data processing device provided in the embodiments of this application, which can be housed in an electronic device or a functional module of an electronic device. The electronic device in the embodiments of this application can be any device capable of implementing the multi-source data processing method in the embodiments of this application.
[0151] Furthermore, this application also provides a method for alarm analysis. In this method, alarm information obtained within a period is periodically detected based on the alarm information and a preset detection strategy to determine the attack events associated with the alarm information. The detection strategy indicates the attack type or characteristics that the attack event conforms to. Thus, by analyzing and summarizing the periodic correlation characteristics of already occurred attack events using an alarm analysis device to obtain a detection strategy, and then periodically detecting the alarm information to be detected within each period based on the detection strategy, the purpose of alarm analysis based on a reasonable detection strategy is achieved, thereby improving the efficiency and accuracy of security threat detection. The alarm information to be detected can be actual alarm information obtained based on the multi-source data processing method provided in this application, enabling reasonable analysis of massive and complex alarms from multiple sources.
[0152] It should be noted that the main body implementing this alarm analysis method can be the alarm analysis apparatus provided in the embodiments of this application, which can be carried in an electronic device or a functional module of an electronic device. The electronic device in the embodiments of this application can be any device capable of implementing the alarm analysis method in the embodiments of this application.
[0153] Furthermore, this application embodiment also provides another alarm analysis method. In this method, firstly, a first process chain of alarm information is obtained. Then, based on the first process chain and a threat data process chain rule base, a first attack event corresponding to the alarm information is determined. The threat data process chain rule base includes at least the correspondence between the first process chain and the first attack event. It is evident that this application embodiment achieves accurate alarm analysis based on process chains by real-time matching of the process chain generated from the alarm information to be analyzed with the threat data process chain rule base, and deriving the corresponding attack event based on the matched process chain in the threat data process chain rule base. This improves the efficiency and accuracy of security threat detection. The alarm information to be detected can be data to be processed or actual alarm information obtained based on the multi-source data processing method provided in this application, enabling reasonable analysis of massive and complex alarms from multiple sources. Moreover, the method provided in this application can mutually verify detection methods based on the periodic correlation features of alarm information, providing a solid foundation for real-time alarms and accurate location of intrusion detection.
[0154] It should be noted that the main body implementing this alarm analysis method can be the alarm analysis apparatus provided in the embodiments of this application, which can be carried in an electronic device or a functional module of an electronic device. The electronic device in the embodiments of this application can be any device capable of implementing the alarm analysis method in the embodiments of this application.
[0155] It should be noted that the meanings of intelligence, log, and alarm information are not the same in the embodiments of this application. After the log data is preprocessed, it is analyzed by intelligence. After the intelligence is matched, the log data and intelligence data are combined and enriched to generate alarm information.
[0156] To facilitate understanding of the specific implementation of the evaluation method for processing multi-source data provided in the embodiments of this application, the following will be combined with the appendix. Figure 1 To be continued Figure 6 Please provide an explanation.
[0157] Figure 1 This is a schematic flowchart illustrating a multi-source data processing method provided in an embodiment of this application. The method is applied to a multi-source data processing device. Figure 1 As shown, the method may include the following steps S101 to S103:
[0158] S101 preprocesses data from multiple data sources to obtain data to be processed.
[0159] It is understandable that the data from multiple data sources have different formats and cannot be processed based on the same standard. Therefore, when implementing the method provided in this application, it is necessary to first preprocess the data from multiple data sources so that the preprocessed data to be processed can be processed according to the same standard.
[0160] As an example, S101 may include: S101a, normalizing the multi-source data to obtain the data to be processed, wherein each data in the data to be processed has the same format.
[0161] As another example, S101 may also include: S101b1, normalizing the multi-source data to obtain intermediate processing data, wherein each data point in the intermediate processing data has the same format; S101b2, parsing the intermediate processing data to obtain the data to be processed, wherein the data to be processed is obtained by extracting key information from the intermediate processing data, the key information including information to be inspected and data details, the information to be inspected including at least one of the following: IP address, domain name (or Domain), or Uniform Resource Locator (URL). Data details may include, for example, but not limited to, the log file (also known as the log payload).
[0162] Whether it's S101a or S101b1~S101b2 implementing S101, they all belong to the data normalization of multiple data sources. The data normalization process is completed through flexible arrangement, without the need to write independent parsers for data with different formats from multiple data sources. This effectively reduces the cost of connecting data from multiple data sources and greatly improves the work efficiency of developers.
[0163] In some implementations, S101a or S101b1 may be implemented as a normalization module in a multi-data source processing device. This normalization module can be used for log access and object extraction based on orchestration tasks. This normalization module supports various log access methods, such as file access, system logs (Syslog), or Kafka message queue access. Specific file types can include, but are not limited to, log traffic from JavaScript Object Notation (JSON), Syslog, Netflow, Domain Name System (DNS), and Hypertext Transfer Protocol (HTTP), as well as SEIM system log records. After access, orchestration tasks can be set up for logs from multiple data sources, parsing them into normalized logs and sending them to a Redis cache queue for temporary storage. Normalized logs refer to logs with the same format.
[0164] In this implementation, the first step is to create log parsing rules through a simple interface configuration; the second step is to select orchestration plugins as needed during process orchestration.
[0165] For the first step, such as Figure 2 As shown, the implementation process of this step may include: S11, basic information entry, i.e., defining the rule name; S12, providing log samples such as JSON, comma-separated values (CSV, sometimes also called character-separated values, because the separator character may not be a comma), etc.; S13, setting pre-filter conditions, logs that do not meet the pre-filter conditions are directly discarded, which can effectively reduce the processing pressure of subsequent processes; S14, selecting the filtering method according to the log sample, such as JSON, CSV, delimiter, etc.; S15, mapping the extracted fields to the normalized fields; Optionally, the implementation process of this step may also include after S15: S16, verifying the effect of the rule, checking whether it meets the requirements, if it does, saving it and using it in subsequent process orchestration, if it does not, it can fall back to S14 to reconfigure.
[0166] In this process, the fields extracted from S15 are mapped to the normalized fields. The main consideration during configuration is whether the semantics of the fields are consistent. For example, for Nginx JSON format logs, the mapping can be seen in Table 1 below:
[0167]
[0168] Through the above mapping, for example, a collision detection function in S102 can be performed on connect_ip to check for malicious or suspicious IPs accessing business sites and update blocking policies in a timely manner. In other words, the above normalization process prepares for the implementation of S102. Furthermore, if the pre-set normalization fields cannot meet the needs of business scenarios, they can be dynamically expanded according to actual requirements. Moreover, various plugins in the process orchestration support custom combinations; by combining input plugins, various parsing rules, and output plugins, a process-oriented processing of log parsing can be achieved. Utilizing... Figure 2 The method shown can meet most log parsing requirements and complete the normalization processing after receiving logs from multiple data sources through interface configuration alone.
[0169] In this way, data normalization is completed through flexible arrangement. Data from multiple data sources with different formats only requires one parser, which effectively reduces the cost of connecting multiple data sources and greatly improves the work efficiency of developers.
[0170] For the implementation process of the second step, as follows: Figure 3 As shown, it may include: S21, selecting the access source, such as Syslog or Kafka as the input method; S22, defining the parsing strategy, where the rules created above can be directly referenced; S23, defining the output, outputting the normalized logs to custom storage (such as a Redis cache queue). It should be noted that in this embodiment, a data source may include one or more parsing strategies, and a parsing strategy may correspond to one or more rules.
[0171] The advantages of using orchestration tasks include: 1) Dynamic input configuration, such as receiving Syslog logs without pre-allocating ports, automatically listening to the port to receive data after starting the orchestration task; 2) Support for multiple parsing formats, such as JSON, CSV, delimiters, regular expressions, Grok decoding, etc.; 3) Visual creation of parsing rules, supporting accurate and complete mapping of parsing results to normalized fields, and parsing rules can be referenced in multiple processes; 4) Filtering rules can be set for different log sources, automatically filtering useless logs and effectively reducing the analysis pressure of subsequent detection processes; 5) Normalized logs include, but are not limited to, fields such as: <asset IP, port>, <network connection IP, port>, access domain name, access URL, log details (also known as log text, usually the content in the log payload), etc. Among them, the definition of network connection takes into account the fact that network traffic may be bidirectional in actual business scenarios. When the asset IP is determined, only the asset IP can be considered, without considering the network direction.
[0172] In some implementations, S101b may be implemented by a parsing module in a multi-data source processing device. This parsing module processes the normalized data. For example, the parsing module retrieves data in batches from the Redis cache queue that stores the normalized data in real time, extracts the IP, Domain, URL, and other data to be checked, as well as the corresponding log details, and sends them to the Redis queue to be checked for temporary storage.
[0173] like Figure 4 As shown, taking multiple data sources including data source A, data source B, and data source C as an example, assuming data source A is in Syslog format, data source B is in Kafka format, and data source C is in Kafka format, then S101 may include, for example: S31, obtaining logs of different formats, and parsing the logs through an orchestration tool to complete the normalization processing of log information; S32, outputting the normalized logs to a Redis cache queue; S33, parsing the normalized logs to extract data such as IP, Domain, and URL, as well as corresponding log details; S34, temporarily storing the parsed content in a Redis waiting queue.
[0174] Thus, by configuring multiple data source log inputs in S101, the logs from multiple data sources are normalized through an orchestration task process, and then the normalized data is processed by a parser and pushed into the inspection queue, preparing for the subsequent execution of S102 and S103.
[0175] S102, based on the data to be processed and the vulnerability detection database, a threat alarm information is obtained. The vulnerability detection database includes vulnerability detection information, and the threat alarm information is the fusion result of the data to be processed and the vulnerability detection information matched in the vulnerability detection database.
[0176] Between S101 and S103, the multi-data source processing device can obtain data to be processed from the queue of data to be inspected and compare the data with threat intelligence big data (such as a vulnerability detection database). The multi-data source processing device pre-stores a vulnerability detection database, which contains vulnerability detection information for known vulnerabilities.
[0177] As an example, the implementation process of S102 can be carried out by a collision module in a multi-data source processing device, such as... Figure 5 As shown, the implementation process of the collision module may include: S41, the collision module retrieves data in batches from the Redis queue to be detected in real time; S42, the retrieved data is assigned to a multi-process task and collided with the vulnerability detection database; S43, it is determined whether vulnerability detection information is matched in the vulnerability detection database. If yes, S44 is executed; otherwise, S45 is executed; S44, the vulnerability detection information and its corresponding log details are merged (or fused) to generate threat alarm information; S45, the data is discarded.
[0178] The vulnerability detection database stores vulnerability detection information that confirms vulnerability. In S43, "determine whether a vulnerability detection information is matched in the vulnerability detection database" can refer to determining whether the characteristics of the data to be processed are consistent with the characteristics of a vulnerability detection information entry in the vulnerability detection database. If they are consistent, it means that the data to be processed matches that vulnerability detection information entry in the vulnerability detection database. Therefore, according to S44, the matched vulnerability detection information is merged with the data to be processed to generate a threat alarm. For example, if the vulnerability detection database contains a vulnerability detection information A: 1.1.1.1 accessed 2.2.2.2 in the afternoon, and the data to be detected, a, also belongs to the access from 1.1.1.1 to 2.2.2.2, it can be determined that the data to be detected, a, matches the vulnerability detection information A. Therefore, threat alarm information X is obtained based on a and A.
[0179] In S44, fusion can refer to simple data merging, or it can refer to integration processing (including integration, padding, and conflict resolution). Taking integration processing as an example, the fields included in the vulnerability detection information A and the data to be processed a are different, including inconsistent field descriptions and different field richness. Processing example one: For "threat level", the field name of A is Risk, and the field name of a is severity. Based on the name of this field in A, the field corresponding to a can be uniformly named Risk, and the value of the Risk field in A can be assigned to the corresponding field in a. Processing example two: For "threat level", A does not have this related field, and the field corresponding to a is Risk. Then, the corresponding field can be mapped to Risk, and the value of the Risk field in a can be used to supplement the value of the Risk field in the final result. Processing example three: For tags, A includes [a1, b1], a includes [b1, c1], or a has no tags. Then, the corresponding field can be mapped to tags, and then the union of all tags appearing in the data source can be taken, i.e., tags : [a1, b1, c1]. In this way, the integrated processing results of the defect detection information and the data to be processed can be stored in the database, and a corresponding query interface can be provided to complete the conflict resolution of IOC data and the output of high-quality data.
[0180] S103, based on the asset information of the threat alarm information, merge the threat alarm information to obtain the actual alarm information.
[0181] S103 is designed to merge threat alerts with results based on the asset dimension, thereby obtaining the actual alert information that needs to be displayed to staff.
[0182] The asset information may include at least one of the following: Internet Protocol (IP) address or device identification information, etc. Threat indicators (IOCs). For example, asset information may be the asset's IP address; alternatively, asset information may include both the asset's IP address and IOCs.
[0183] It should be noted that the data to be processed, threat alarm information, real-time alarm information, etc. mentioned in the embodiments of this application all refer to a type of data. This type of information includes multiple data with the same characteristics (such as the same format, matching the loss detection information in the loss detection database, or the same asset information, etc.), and does not specifically refer to a certain information.
[0184] As an example, regarding the first and second alarm messages in the threat alarm information, S103 may include: if the first asset information of the first alarm message and the second asset information of the second alarm message match, then the first alarm message and the second alarm message are merged to obtain a third alarm message, and the actual alarm message includes the third alarm message; if the first asset information of the first alarm message and the second asset information of the second alarm message do not match, then the first alarm message and the second alarm message are created as the actual alarm message, and the actual alarm message includes the first alarm message and the second alarm message.
[0185] The S103 can be implemented by the merging module in a multi-data source processing device. The merging module can be used to merge massive threat alerts based on the asset dimension. By matching with threat intelligence, the merging module provides an efficient log storage and alert detection mechanism, which can accurately detect and alert using existing intelligence information in the case of large log volume heterogeneous data sources.
[0186] To merge threat alerts with collision results based on the asset dimension, for example, when configuring multiple data source log inputs, based on the asset IP and / or port of the logs, alarm IOC merging can be implemented, or threat alerts based on the asset IP plus the same alarm IOC can be merged. Figure 6As shown, S103 may include: S51, the merging task reads the threat alarm information; S52, check if the asset IP of the threat alarm information is appearing for the first time. If it is, execute S53; otherwise, execute S54; S53, create an aggregated threat alarm, that is, use the first-time occurrence of the threat alarm information as the actual alarm information; S54, determine whether the IOC of the threat alarm information to be merged is consistent with that of the actual alarm information that has appeared. If they are consistent, execute S55; otherwise, execute S56; S55, update the aggregated threat information, that is, merge the threat alarm information to be merged that has consistent asset IP and IOC with other content of the actual alarm information that has appeared except for asset IP and IOC, to obtain the updated actual alarm information; S56, add an aggregated threat alarm, that is, merge the threat alarm information to be merged that has consistent asset IP but inconsistent IOC with other content of the actual alarm information that has appeared except for asset IP, to obtain the updated actual alarm information. S54~S56 can be understood as: generating actual alarm information based on the threat alarm information; otherwise, determining whether to add or update actual alarm information based on whether the IOC appears repeatedly. Regardless of whether to update or add actual alarm information, the threat alarm information to be merged is merged with the actual alarm information that has appeared. That is, the fields that appear in both are taken as the union, and the values of the fields are selected.
[0187] In some implementations, to improve the processing efficiency of S103 and avoid the backlog of threat alarm information, the merging module in this embodiment supports adaptive dynamic adjustment of the merging time window size. The time window value T can be configured automatically based on the log volume H, the system merging rate S, and the intelligence collision coefficient K. As an example, S103 may include: determining the time window value based on the amount of data to be processed, the merging processing rate, and the data collision coefficient, wherein the data collision coefficient is determined based on the number of threat alarm information and the data to be processed; and executing step S103 based on the time window value.
[0188] The merging time window T is dynamically adjusted according to the log volume H, the system merging rate S, and the intelligence collision coefficient K. For example, the following formula (1) can be used to achieve T's self-configuration.
[0189] (H) K T) <= S T …… Formula (1)
[0190] Among them, the log volume H is the number of logs per second or gigabytes per second when the traffic peaks. The number of gigabytes per second can be calculated by converting each log into 1kb in size. The system merging rate S is obtained through performance testing. For example, it can be the maximum number of alarm logs processed per second under log traffic saturation. The intelligence collision coefficient K is automatically calculated by the system. For example, the system samples the current amount of received log data N every 15 minutes, and the number of collision hit logs is M. The average value of M / N data over 6 hours is taken. The time window size T can be automatically adjusted according to the changes in the intelligence coefficient K, or it can be adaptively adjusted according to the aggregation time specified by the user. It can be initially set to merge once every 900 seconds.
[0191] For example, assuming the strategy adopted is that the system performs a merge every 900 seconds, and the system's maximum rate S is 5000 records / second, then the maximum amount of data processed in 900 seconds is (900... 5000) = 4,500,000; The intelligence collision coefficient K is taken as an example of 10 out of 10,000 log alarm data in the last 6 hours. The value of K is 10 / 10000 = 0.001; The system generates a maximum peak of 10,000 logs per second. It can be seen that H = 10,000 logs / second, K = 0.001, T = 900 seconds. According to the above formula (1), we have: 900 (seconds) 10,000 messages / second 0.001 = 9000, since 9000 < 5000 From 900, we can see that S103's processing can consume all the log records.
[0192] When the log volume increases, a dynamic adjustment mechanism can be used to adjust the merging time window size. Specifically, the system records the amount of log data received over a continuous 6-hour period (sampled every 15 minutes and stored). If a sudden increase in the log volume is detected at a certain sampling time, the merging time window size will be automatically adjusted in a multiple relationship. For example, if the received data volume suddenly increases from 10,000 to 20,000, the merging time window changes from 900 seconds to 450 seconds. If the data volume continues to increase, the minimum merging time window size is set to once every 60 seconds to prevent the merging time from continuously being divided by a factor as the traffic increases indefinitely. Furthermore, through continuous sampling and monitoring of the received data volume, the merging time window will revert to the default 900-second time window when the system detects a decrease in traffic.
[0193] It is evident that by adaptively adjusting the size of the merging time window, the utilization rate of system resources can be optimized, the processing pressure on database components can be effectively reduced, and data processing capabilities can be improved.
[0194] The advantages of this method through the merging process in S103 include: 1) effectively reducing the actual number of alarms while also reducing the database storage space usage; 2) providing a direct overview of the attack on an asset and quickly accessing detailed information about related IOCs at the asset level; and 3) effectively reducing database pressure by dynamically adjusting the merging time window, making reasonable use of the central processing unit (CPU) and memory, maximizing system resource utilization, and avoiding unnecessary resource waste.
[0195] As can be seen, the method provided in this application preprocesses the multi-source data to make the obtained data to be processed meet the requirements of unified processing. Then, through matching the data to be processed with the known vulnerability detection database and merging the asset dimensions, the actual alarm information is less than the data to be processed. This not only achieves simple, efficient and accurate processing of multi-source data, but also effectively reduces the number of alarms, which facilitates subsequent alarm analysis and threat discovery based on reasonable alarm analysis strategies.
[0196] In some implementations, this method can also perform policy-based detection and analysis of actual alarm information or data to be processed. For example, based on the generated alarm information, the frequency of log alarms can be analyzed. By utilizing the differences between Trojan horses, viruses, and normal network communication behavior, the frequency of alarms can be analyzed to determine whether it conforms to the regularity of "heartbeat intervals," efficiently detecting abnormal attacks and providing tracing to promptly sever the connection between the controlled end and the attacker. Another example is the process chain-based threat analysis method, which analyzes and correlates the process events of assets by accessing process event logs, tracing the parent and ancestor processes of the current process, and matching the generated process chain with the threat intelligence process chain rule base in real time. Alarms are then issued for the process behavior of assets that match the rules in the threat intelligence process chain rule base.
[0197] As an example, the method may further include: periodically detecting the actual alarm information according to a preset detection strategy to determine attack events. In one case, if the alarm frequency of the actual alarm information matches a preset first heartbeat range within a period, then the attack event is determined as a first attack event based on the first heartbeat range and a saved first correspondence, where the first correspondence includes the correspondence between the first heartbeat range and the first attack event. In another case, firstly, within a period, the attack type is predicted based on the actual alarm information; then, a monitoring time range is determined based on the attack type and a saved second correspondence, where the second correspondence includes the correspondence between the attack type and the monitoring time range; then, based on the analysis of the actual alarm information occurring on external ports within the monitoring time range, the attack event is determined as a first attack event, where the actual alarm information matches the pattern of the external ports where the first attack event occurs.
[0198] Here, the external port refers to port information such as the sequence number or time of the connection port between the control end and the controlled end. Considering the characteristics of attack behavior, the controlled end will connect to the control end at a fixed period and frequency when updating its status, obtaining new instructions, etc. Although it may use random delays as a disguise, there are still patterns to be found between attack events and external ports. For example, the external ports used may be consecutive or form an arithmetic sequence. Based on this, in this embodiment, the known correspondence between various attack events and external port patterns can be preset. Thus, within the monitoring time range, the external ports where the actual alarm information occurs are matched with the pre-saved correspondence. If a match is found, the attack type of the actual alarm information is determined to be the attack event corresponding to the matched external port pattern.
[0199] For periodic detection, based on the generated alarm information, the actual alarm information of assets + IOC can be displayed in reverse chronological order, and the frequency of log alarms can be analyzed. The heartbeat frequency in Trojans, viruses, and C&C attacks is usually very obvious. Because the controlled end generally connects to the control end periodically to update its status and obtain new instructions, the connection period and frequency are relatively fixed. Although random delays may be used as a disguise, they are still significantly different from normal network communication behavior. Therefore, analyzing whether the alarm frequency conforms to the regularity of "heartbeat intervals" can identify whether there is heartbeat behavior, which is one of the important connection characteristics for detecting Advanced Persistent Threat (APT) attacks. In addition to conforming to heartbeat behavior, periodicity is also reflected in the fact that the ports used are usually consecutive ports or form an arithmetic sequence to avoid the traffic characteristics being too concentrated.
[0200] Based on the above characteristics, the following two detection strategies can be set, and security operations personnel can dynamically adjust them according to the actual situation:
[0201] 1) Set a heartbeat range and monitor whether there are alarm assets that meet the preset heartbeat range. If so, the alarm may be considered an attack event.
[0202] 2) Set a time range. After predicting a possible attack type, data analysis can be performed on the external ports within the preset time range corresponding to the attack type to determine whether they meet the requirements of consecutive ports or an arithmetic sequence. If they do, the alarm can be considered to be an attack event.
[0203] In this way, by setting heartbeat ranges and / or time ranges, periodic detection can efficiently detect abnormal attacks and provide traceability, effectively helping security operations staff to promptly cut off the connection between the controlled terminal and the attacker.
[0204] As another example, the method may further include: obtaining a first process chain of the actual alarm information; determining a first attack event corresponding to the actual alarm information based on the first process chain and a threat data process chain rule base, wherein the threat data process chain rule base includes the correspondence between the first process chain and the first attack event. Alternatively, if the threat data process chain rule base includes the correspondence between the first process chain, the first attack event, and the first threat level, then the method may further include: determining a first threat level corresponding to the actual alarm information based on the first process chain and the threat data process chain rule base. The step of obtaining the first process chain of the actual alarm information may, for example, include: for the actual alarm information, accessing process event logs, extracting fields, normalizing data, analyzing and associating process events corresponding to the asset (e.g., calculation by an analysis engine), tracing the parent and ancestor processes of the current process, thereby obtaining the first process chain of the actual alarm information.
[0205] For process chain-based threat analysis methods, process event logs provided by EDR, Sysmon, and host security tools can be accessed to analyze and correlate process events corresponding to assets, tracing the parent and ancestor processes of the current process. Taking EDR process event logs as an example, information such as Processor Identity (PID), parent PID, process command line, parent process command line, and process hash can be provided. After logging, field extraction, and data normalization, the analysis engine can calculate a process chain similar to cmd.exe -> WmiPrvse.exe -> notepad.exe. The generated process chain is then matched in real time with the threat intelligence process chain rule base, providing alerts for the process behavior of assets that match the rules (and also obtaining the threat level).
[0206] In addition, by increasing the detection priority and storing all logs for periodic backscanning of compromise detection intelligence, it can complement the above-mentioned analysis method based on network feature matching of compromise detection information, providing a solid foundation for real-time alerts and accurate location of intrusion detection.
[0207] As can be seen, this method preprocesses the multi-source data to ensure that the data to be processed meets the requirements for unified processing. By matching the data to be processed with the known vulnerability detection database and merging asset dimensions, fewer actual alarm messages are obtained than the data to be processed. This not only achieves simple, efficient, and accurate processing of multi-source data but also effectively reduces the number of alarms, facilitating subsequent alarm analysis and threat discovery based on reasonable alarm analysis strategies.
[0208] It should be noted that the embodiments of this application can be applied to threat intelligence in customers' APT threat tracking and detection projects to discover APT attacks. In addition, the threat detection scheme based on the embodiments of this application greatly improves the ability to extract and discover threats from massive logs, reduces the analysis burden of on-site security analysis personnel, improves analysis efficiency, and effectively tracks threats.
[0209] In this embodiment, the alarm overview and asset threat alarm details based on the asset dimension in the log detection method of threat intelligence can be visualized and displayed in a page format. The log access methods of multiple data sources and the results of log detection are displayed. The log access types meet common log data types (such as DNS, Transmission Control Protocol (TCP), User Datagram Protocol (UDP), HTTP, etc.), which greatly improves the detection efficiency of network security detection of massive data, reduces the manpower cost of R&D and analysis personnel, and improves the work efficiency of threat analysis and threat assessment of security operations personnel.
[0210] Furthermore, embodiments of this application also provide a method for alarm analysis, such as... Figure 7 As shown, the method may include:
[0211] S701 receives alarm information within a cycle;
[0212] S702, based on the alarm information and a preset detection strategy, the alarm information is periodically detected to determine the attack event of the alarm information. The detection strategy is used to indicate the attack type or the characteristics that the attack event meets.
[0213] Optionally, S702 may include, for example, the following: if the alarm frequency of the alarm information matches a preset first heartbeat range, then the attack event is determined to be a first attack event based on the first heartbeat range and a saved first correspondence, wherein the first correspondence includes the correspondence between the first heartbeat range and the first attack event.
[0214] Optionally, the method may further include: if the alarm frequency of the alarm information does not match the heartbeat range in all stored correspondences, then it is determined that no attack event has occurred.
[0215] Optionally, S702 may include, for example,: predicting the attack type based on the alarm information; determining a monitoring time range based on the attack type and a stored second correspondence, the second correspondence including the correspondence between the attack type and the monitoring time range; and determining the attack event as a first attack event based on the analysis of the alarm information occurring on the external connection port within the monitoring time range, wherein the alarm information matches the pattern of the external connection port where the first attack event occurred.
[0216] Optionally, the method may further include: obtaining a first process chain of the alarm information; determining a second attack event corresponding to the alarm information based on the first process chain and a threat data process chain rule base, wherein the threat data process chain rule base includes the correspondence between the first process chain and the second attack event.
[0217] Optionally, the method may further include: determining a first threat level corresponding to the alarm information based on the first process chain and the threat data process chain rule base, wherein the threat data process chain rule base includes the correspondence between the first process chain, the second attack event and the first threat level.
[0218] Optionally, the method may further include: verifying the second attack event based on the first attack event to determine an alarm result.
[0219] Optionally, the alarm information is as described above. Figure 1 The actual alarm information provided in the method.
[0220] In this way, by analyzing and summarizing the periodic correlation characteristics of past attack events through an alarm analysis device, a detection strategy is obtained. Based on the detection strategy, the alarm information to be detected within each period is periodically detected, achieving the goal of alarm analysis based on a reasonable detection strategy, thereby improving the efficiency and accuracy of security threat detection. The alarm information to be detected can be actual alarm information obtained based on the multi-source data processing method provided in this application, enabling reasonable analysis of massive and complex alarms from multiple sources.
[0221] Furthermore, embodiments of this application also provide a method for alarm analysis, such as... Figure 8 As shown, for example, it may include:
[0222] S801, the first process chain to obtain alarm information;
[0223] S802, based on the first process chain and the threat data process chain rule base, determine the first attack event corresponding to the alarm information, wherein the threat data process chain rule base includes the correspondence between the first process chain and the first attack event.
[0224] Optionally, the method may further include: determining a first threat level corresponding to the alarm information based on the first process chain and the threat data process chain rule base, wherein the threat data process chain rule base includes the first process chain, the correspondence between the first attack event and the first threat level.
[0225] Optionally, S801 may include, for example, obtaining process information of the alarm information, the process information including the process identifier of the alarm information, the parent process identifier, the process command line, the parent process command line, or the process hash value; and performing analysis engine calculations on the process information to obtain the first process chain.
[0226] Optionally, the alarm information is a process event log from any of the following sources: EDR system, Sysmon, or host security.
[0227] Optionally, the alarm information is as described above. Figure 1 The actual alarm information or pending data in the provided method.
[0228] As can be seen, in this embodiment, the process chain generated from the alarm information to be analyzed is matched in real time with the threat data process chain rule base. Based on the matched process chain in the threat data process chain rule base, the corresponding attack event is derived, achieving precise alarm analysis based on the process chain, thereby improving the efficiency and accuracy of security threat detection. The alarm information to be detected can be data to be processed or actual alarm information obtained based on the multi-source data processing method provided in this embodiment, enabling reasonable analysis of massive and complex alarms from multiple sources. Furthermore, the method provided in this embodiment can be mutually verified with detection methods based on the periodic correlation features of alarm information, providing a solid foundation for real-time alarms and accurate location of intrusions.
[0229] Accordingly, embodiments of this application also provide a multi-source data processing apparatus 900, such as... Figure 9 As shown, the device 900 may include:
[0230] The preprocessing unit 901 is used to preprocess data from multiple data sources to obtain data to be processed;
[0231] The collision unit 902 is used to obtain threat alarm information based on the data to be processed and the trap detection database. The trap detection database includes trap detection information, and the threat alarm information is the fusion result of the data to be processed and the trap detection information matched in the trap detection database.
[0232] The merging unit 903 is used to merge the threat alarm information based on the asset information of the threat alarm information to obtain the actual alarm information.
[0233] The asset information may include at least one of the following: IP address or device identification information, etc.
[0234] Optionally, for the first alarm information and the second alarm information in the threat alarm information, the merging unit 903 is specifically used for:
[0235] If the first asset information of the first alarm information matches the second asset information of the second alarm information, then the first alarm information and the second alarm information are merged to obtain the third alarm information, and the actual alarm information includes the third alarm information;
[0236] If the first asset information of the first alarm information and the second asset information of the second alarm information do not match, then the first alarm information and the second alarm information are created as the actual alarm information, which includes the first alarm information and the second alarm information.
[0237] Optionally, the merging unit 903 includes:
[0238] A sub-unit is defined to determine a time window value based on the amount of data to be processed, the merging processing rate, and the data collision coefficient. The data collision coefficient is determined based on the number of threat alarm messages and the data to be processed.
[0239] The merging subunit is used to perform the step of merging the threat alarm information based on the asset information of the threat alarm information to obtain the actual alarm information according to the time window value.
[0240] Optionally, the preprocessing unit 901 is specifically used for:
[0241] The multi-source data is normalized to obtain the data to be processed, wherein each data in the data to be processed has the same format.
[0242] Optionally, the preprocessing unit 901 is specifically used for:
[0243] The multi-source data is normalized to obtain intermediate processing data, wherein each data point in the intermediate processing data has the same format.
[0244] The intermediate processing data is parsed to obtain the data to be processed. The data to be processed is obtained by extracting key information from the intermediate processing data. The key information includes information to be inspected and data details. The information to be inspected includes at least one of the following: IP address, Domain, or URL.
[0245] Optionally, the device 900 further includes: a detection unit,
[0246] The detection unit is used to periodically detect the actual alarm information according to a preset detection strategy to determine the attack event.
[0247] Optionally, the detection unit is specifically used for:
[0248] If the alarm frequency of the actual alarm information matches the preset first heartbeat range within a cycle, then the attack event is determined as the first attack event according to the first heartbeat range and the saved first correspondence relationship. The first correspondence relationship includes the correspondence relationship between the first heartbeat range and the first attack event.
[0249] Alternatively, the detection unit is specifically used for:
[0250] Within a cycle, predict the attack type based on actual alarm information;
[0251] The monitoring time range is determined based on the attack type and the saved second correspondence, wherein the second correspondence includes the correspondence between the attack type and the monitoring time range;
[0252] Based on the analysis of the alarm information occurring on the external connection ports within the monitoring time range, the attack event is determined to be the first attack event, and the alarm information matches the pattern of the external connection ports where the first attack event occurred.
[0253] Optionally, the device 900 further includes: an acquisition unit and a detection unit.
[0254] The obtaining unit is used to obtain the first process chain of the actual alarm information;
[0255] The detection unit is used to determine the first attack event corresponding to the actual alarm information based on the first process chain and the threat data process chain rule base, wherein the threat data process chain rule base includes the correspondence between the first process chain and the first attack event.
[0256] Optionally, the detection unit is further configured to determine the first threat level corresponding to the actual alarm information based on the first process chain and the threat data process chain rule base, wherein the threat data process chain rule base includes the correspondence between the first process chain, the first attack event and the first threat level.
[0257] It should be noted that the specific implementation method and effects of the device 900 can be found in [reference needed]. Figure 1 Description of related embodiments of the method.
[0258] Accordingly, this application also provides an alarm analysis device 1000, such as... Figure 10 As shown, it includes:
[0259] The first acquisition unit 1001 is used to acquire alarm information within one cycle;
[0260] The first detection unit 1002 is used to periodically detect the alarm information based on the alarm information and a preset detection strategy to determine the attack event of the alarm information. The detection strategy is used to indicate the attack type or the characteristics that the attack event meets.
[0261] Optionally, the first detection unit 1002 is specifically used for:
[0262] If the alarm frequency of the alarm information matches a preset first heartbeat range, then the attack event is determined to be a first attack event based on the first heartbeat range and a saved first correspondence. The first correspondence includes the correspondence between the first heartbeat range and the first attack event.
[0263] Optionally, the first detection unit 1002 is further configured to:
[0264] If the alarm frequency of the alarm information does not match the heartbeat range in all the stored correspondences, then it is determined that no attack event has occurred.
[0265] Optionally, the first detection unit 1002 further includes:
[0266] The prediction subunit is used to predict the attack type based on the alarm information;
[0267] A subunit is defined to determine a monitoring time range based on the attack type and a stored second correspondence, wherein the second correspondence includes the correspondence between the attack type and the monitoring time range.
[0268] The analysis subunit is used to determine the attack event as a first attack event based on the analysis of the alarm information occurring on the external ports within the monitoring time range, and the alarm information matches the pattern of the external ports where the first attack event occurred.
[0269] Optionally, the device 1000 further includes:
[0270] The second obtaining unit is used to obtain the first process chain of the alarm information;
[0271] The second detection unit is used to determine the second attack event corresponding to the alarm information based on the first process chain and the threat data process chain rule base, wherein the threat data process chain rule base includes the correspondence between the first process chain and the second attack event.
[0272] Optionally, the second detection unit is further configured to:
[0273] Based on the first process chain and the threat data process chain rule base, the first threat level corresponding to the alarm information is determined. The threat data process chain rule base includes the correspondence between the first process chain, the second attack event, and the first threat level.
[0274] Optionally, the device 1000 further includes: a verification unit,
[0275] The verification unit is used to verify the second attack event based on the first attack event and determine the alarm result.
[0276] Optionally, the alarm information is as described above. Figure 9 The actual alarm information provided by device 900.
[0277] It should be noted that the specific implementation method and effects of the device 1000 can be found in [reference needed]. Figure 7 Description of related embodiments of the method.
[0278] Accordingly, this application also provides an alarm analysis device 1100, such as... Figure 11 As shown, it includes:
[0279] The acquisition unit 1101 is used to acquire the first process chain of alarm information;
[0280] The matching unit 1102 is used to determine the first attack event corresponding to the alarm information based on the first process chain and the threat data process chain rule base, wherein the threat data process chain rule base includes the correspondence between the first process chain and the first attack event.
[0281] Optionally, the matching unit 1102 is further configured to:
[0282] Based on the first process chain and the threat data process chain rule base, a first threat level corresponding to the alarm information is determined. The threat data process chain rule base includes the first process chain, the correspondence between the first attack event and the first threat level, and so on.
[0283] Optionally, the obtaining unit 1101 is specifically used for:
[0284] The process information for obtaining the alarm information includes the process identifier of the alarm information, the parent process identifier, the process command line, the parent process command line, or the process hash value.
[0285] The process information is analyzed by an analysis engine to obtain the first process chain.
[0286] Optionally, the alarm information is a process event log from any of the following sources: EDR system, Sysmon, or host security.
[0287] Optionally, the alarm information is as described above. Figure 9 The actual alarm information or pending data provided in the device 900.
[0288] It should be noted that the specific implementation method and effects of the device 1100 can be found in [reference needed]. Figure 8 Description of related embodiments of the method.
[0289] Furthermore, embodiments of this application also provide an electronic device 1200, such as... Figure 12 As shown, the electronic device 1200 includes a processor 1201 and a memory 1202:
[0290] The memory 1202 is used to store computer programs;
[0291] The processor 1201 is used to execute the method provided in the embodiments of this application according to the computer program.
[0292] Furthermore, embodiments of this application also provide a computer-readable storage medium for storing a computer program for executing the method provided in embodiments of this application.
[0293] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the methods of the above embodiments can be implemented by means of software plus a general-purpose hardware platform. Based on this understanding, the technical solution of this application can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as a read-only memory (ROM) / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, a server, or a network communication device such as a router) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0294] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system and device embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The device and system embodiments described above are merely illustrative. Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0295] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for processing multi-source data, characterized in that, include: Preprocess data from multiple data sources to obtain data to be processed; Based on the data to be processed and the vulnerability detection database, threat alarm information is obtained. The vulnerability detection database includes vulnerability detection information, and the threat alarm information is the fusion result of the data to be processed and the vulnerability detection information matched in the vulnerability detection database. Based on the asset information of the threat alert information, the threat alert information is merged to obtain the actual alert information; The asset information includes at least one of the following: Internet Protocol (IP) address or device identification information; The method further includes: The actual alarm information is periodically detected according to a preset detection strategy to determine attack events; The step of periodically detecting the actual alarm information according to a preset detection strategy to determine attack events includes: If the alarm frequency of the actual alarm information matches the preset first heartbeat range within a cycle, then the attack event is determined as the first attack event according to the first heartbeat range and the saved first correspondence relationship. The first correspondence relationship includes the correspondence relationship between the first heartbeat range and the first attack event. Alternatively, the step of periodically detecting the actual alarm information according to a preset detection strategy to determine attack events includes: Within a cycle, predict the attack type based on actual alarm information; The monitoring time range is determined based on the attack type and the saved second correspondence, wherein the second correspondence includes the correspondence between the attack type and the monitoring time range; Based on the analysis of the actual alarm information occurring on the external connection ports within the monitoring time range, the attack event is determined to be the first attack event, and the actual alarm information matches the pattern of the external connection ports where the first attack event occurred.
2. The method according to claim 1, characterized in that, For the first and second threat alerts, the step of merging the threat alerts based on the asset information to obtain the actual alert information includes: If the first asset information of the first alarm information matches the second asset information of the second alarm information, then the first alarm information and the second alarm information are merged to obtain the third alarm information, and the actual alarm information includes the third alarm information; If the first asset information of the first alarm information and the second asset information of the second alarm information do not match, then the first alarm information and the second alarm information are created as the actual alarm information, which includes the first alarm information and the second alarm information.
3. The method according to claim 1, characterized in that, The asset information based on the threat alert information is used to merge the threat alert information to obtain the actual alert information, including: The time window value is determined based on the amount of data to be processed, the merging processing rate, and the data collision coefficient. The data collision coefficient is determined based on the number of threat alarm messages and the data to be processed. Based on the time window value, the step of merging the threat alarm information with the asset information based on the threat alarm information to obtain the actual alarm information is performed.
4. The method according to claim 1, characterized in that, The preprocessing of data from multiple data sources to obtain data to be processed includes: The multi-source data is normalized to obtain the data to be processed, wherein each data in the data to be processed has the same format.
5. The method according to claim 1, characterized in that, The preprocessing of data from multiple data sources to obtain data to be processed includes: The multi-source data is normalized to obtain intermediate processing data, wherein each data point in the intermediate processing data has the same format. The intermediate processing data is parsed to obtain the data to be processed. The data to be processed is obtained by extracting key information from the intermediate processing data. The key information includes information to be inspected and data details. The information to be inspected includes at least one of the following: IP address, domain, or Uniform Resource Locator (URL).
6. The method according to any one of claims 1-5, characterized in that, The method further includes: The first process chain that obtains the actual alarm information; Based on the first process chain and the threat data process chain rule base, the first attack event corresponding to the actual alarm information is determined, and the threat data process chain rule base includes the correspondence between the first process chain and the first attack event.
7. The method according to claim 6, characterized in that, The method further includes: Based on the first process chain and the threat data process chain rule base, the first threat level corresponding to the actual alarm information is determined. The threat data process chain rule base includes the correspondence between the first process chain, the first attack event, and the first threat level.
8. A method for alarm analysis, characterized in that, include: Within a cycle, alarm information is obtained; Based on the alarm information and the preset detection strategy, the alarm information is periodically detected to determine the attack event of the alarm information. The detection strategy is used to indicate the attack type or the characteristics that the attack event meets.
9. The method according to claim 8, characterized in that, The method of periodically detecting the alarm information based on the alarm information and a preset detection strategy to determine the attack event of the alarm information includes: If the alarm frequency of the alarm information matches a preset first heartbeat range, then the attack event is determined to be a first attack event based on the first heartbeat range and a saved first correspondence. The first correspondence includes the correspondence between the first heartbeat range and the first attack event.
10. The method according to claim 9, characterized in that, The method further includes: If the alarm frequency of the alarm information does not match the heartbeat range in all the stored correspondences, then it is determined that no attack event has occurred.
11. The method according to claim 8, characterized in that, The method of periodically detecting the alarm information based on the alarm information and a preset detection strategy to determine the attack event of the alarm information includes: Predict the attack type based on the alarm information; The monitoring time range is determined based on the attack type and the saved second correspondence, wherein the second correspondence includes the correspondence between the attack type and the monitoring time range; Based on the analysis of the alarm information occurring on the external connection ports within the monitoring time range, the attack event is determined to be the first attack event, and the actual alarm information matches the pattern of the external connection ports where the first attack event occurred.
12. The method according to any one of claims 8-11, characterized in that, The method further includes: The first process chain that obtains the alarm information; Based on the first process chain and the threat data process chain rule base, the second attack event corresponding to the alarm information is determined, and the threat data process chain rule base includes the correspondence between the first process chain and the second attack event.
13. The method according to claim 12, characterized in that, The method further includes: Based on the first process chain and the threat data process chain rule base, the first threat level corresponding to the alarm information is determined. The threat data process chain rule base includes the correspondence between the first process chain, the second attack event, and the first threat level.
14. The method according to claim 12, characterized in that, The method further includes: The second attack event is verified based on the first attack event to determine the alarm result.
15. The method according to any one of claims 8-11, characterized in that, The alarm information is the actual alarm information in the method provided by any one of claims 1-7.
16. A method for alarm analysis, characterized in that, include: The first process chain for obtaining alarm information; Based on the first process chain and the threat data process chain rule base, the first attack event corresponding to the alarm information is determined, and the threat data process chain rule base includes the correspondence between the first process chain and the first attack event.
17. The method according to claim 16, characterized in that, The method further includes: Based on the first process chain and the threat data process chain rule base, the first threat level corresponding to the alarm information is determined. The threat data process chain rule base includes the correspondence between the first process chain, the first attack event, and the first threat level.
18. The method according to claim 16, characterized in that, The first process chain for obtaining alarm information includes: The process information for obtaining the alarm information includes the process identifier of the alarm information, the parent process identifier, the process command line, the parent process command line, or the process hash value. The process information is analyzed to obtain the first process chain.
19. The method according to claim 16, characterized in that, The alarm information is a process event log from any of the following sources: Terminal Detection and Response (EDR) system, System Monitor Sysmon, or Host Security.
20. The method according to any one of claims 16-19, characterized in that, The alarm information is the actual alarm information or data to be processed in the method provided by any one of claims 1-7.
21. A multi-source data processing apparatus, characterized in that, include: The preprocessing unit is used to preprocess data from multiple data sources to obtain data to be processed; The collision unit is used to obtain threat alarm information based on the data to be processed and the vulnerability detection database. The vulnerability detection database includes vulnerability detection information, and the threat alarm information is the fusion result of the data to be processed and the vulnerability detection information matched in the vulnerability detection database. The merging unit is used to merge the threat alarm information based on the asset information of the threat alarm information to obtain the actual alarm information; The asset information includes at least one of the following: Internet Protocol (IP) address or device identification information; The device further includes: a detection unit, The detection unit is used to periodically detect the actual alarm information according to a preset detection strategy to determine attack events; The detection unit is specifically used for: If the alarm frequency of the actual alarm information matches the preset first heartbeat range within a cycle, then the attack event is determined as the first attack event according to the first heartbeat range and the saved first correspondence relationship. The first correspondence relationship includes the correspondence relationship between the first heartbeat range and the first attack event. Alternatively, the detection unit is specifically used for: Within a cycle, predict the attack type based on actual alarm information; The monitoring time range is determined based on the attack type and the saved second correspondence, wherein the second correspondence includes the correspondence between the attack type and the monitoring time range; Based on the analysis of the alarm information occurring on the external connection ports within the monitoring time range, the attack event is determined to be the first attack event, and the alarm information matches the pattern of the external connection ports where the first attack event occurred.
22. An alarm analysis device, characterized in that, include: The acquisition unit is used to acquire alarm information within a cycle; The detection unit is used to periodically detect the alarm information based on the alarm information and a preset detection strategy to determine the attack event of the alarm information. The detection strategy is used to indicate the attack type or the characteristics that the attack event meets.
23. An alarm analysis device, characterized in that, include: The acquisition unit is the first process chain used to acquire alarm information. The matching unit is used to determine the first attack event corresponding to the alarm information based on the first process chain and the threat data process chain rule base, wherein the threat data process chain rule base includes the correspondence between the first process chain and the first attack event.
24. An electronic device, characterized in that, The electronic device includes a processor and a memory: The memory is used to store computer programs; The processor is configured to perform the method according to any one of claims 1-20 according to the computer program.
25. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program for performing the method according to any one of claims 1-20.
Citation Information
Patent Citations
Threat intelligence-based network threat identification method and identification system
CN110719291A
Data processing method and device, computer system and storage medium
CN111988341A
Threat tracing method and related equipment
CN112131571A
Penetration attack identification method, device and system, storage medium and electronic device
CN112398786A
Attack traffic detection method and device, storage medium and electronic equipment
CN112583774A