Network security threat analysis method and apparatus, and electronic device and storage medium

By screening and analyzing network communication logs, combining historical threat intelligence data, using reverse estimation data leakage model to calculate data leakage volume, and generating threat risk warning messages, the problem of untimely network security threat analysis in the existing technology is solved, and rapid and efficient detection and early warning of network security is achieved.

WO2025130600A1PCT designated stage expired Publication Date: 2025-06-26CHINA TELECOM NETWORK SECURITY TECH CO LTD

Patent Information

Application Number
PCT/CN2024/136463
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-03
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The existing technology is difficult to conduct timely and accurately analyze network security threats, resulting in the inability to promptly early warning and respond to cyber attacks.

Method used

By obtaining the communication logs of the target network, and filtering out a set of threat logs that meet preset rules based on historical threat intelligence data, the reverse estimation data leakage model is used to calculate the reverse estimation data leakage amount to generate threat risk warning messages.

Benefits of technology

It realizes rapid and efficient detection and early warning of network security threats, ensuring timely response to network security and risk management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136463_26062025_PF_FP_ABST
    Figure CN2024136463_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of network security, and in particular to a network security threat analysis method and apparatus, and an electronic device and a storage medium. The method comprises: acquiring a plurality of communication logs of a target network within a set time range, and performing screening on the basis of historical threat intelligence data and source IP addresses and destination IP addresses of the plurality of communication logs, so as to obtain a threat log set that meets a preset rule; comparing the data volume of each threat log with a preset data volume threshold value, so as to obtain a first log set having a high threat risk analysis value and a second log set having a low threat risk analysis value; on the basis of the data volume and number of log packets in the first log set and the data volume and number of log packets in the second log set, using a reverse estimation model for data leakage to determine a reverse-estimated data leakage volume; and when the reverse-estimated data leakage volume is not less than a set threat threshold value, generating a threat risk early-warning message. By means of the solution, threat analysis for network security can be performed in a timely manner, and the accuracy of threat analysis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Network security threat analysis method, device, electronic device and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on December 22, 2023, with application number 202311785338.8 and application name "A network security threat analysis method, device, electronic device and storage medium", the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of network security technology, and in particular to a network security threat analysis method, device, electronic device, and storage medium. Background Art

[0004] Against the backdrop of informatization, digitization, and intelligentization, networks have become a crucial development direction. However, today's networks face numerous attack threats, such as APTs, information theft, and distributed denial of service (DDoS) attacks. These attacks often threaten highly confidential intelligence information in sectors such as politics, scientific research, military industry, aerospace, foreign trade, and economics, causing losses to various organizations.

[0005] Therefore, how to conduct threat analysis on network security in a timely manner has become an urgent problem that needs to be solved. Summary of the Invention

[0006] Embodiments of the present application provide a network security threat analysis method, apparatus, electronic device, and storage medium for detecting threat activity status based on threat intelligence data in order to make accurate risk assessments and threat warnings.

[0007] In a first aspect, an embodiment of the present application provides a network security threat analysis method, the method comprising:

[0008] Acquire multiple communication logs of a target network within a set time range, and filter the multiple communication logs based on historical threat intelligence data and the source IP addresses and destination IP addresses of the multiple communication logs to obtain a threat log set that meets preset rules, where the threat log set includes multiple threat logs, and the historical threat intelligence data includes multiple historical threat logs that are valuable for threat risk analysis;

[0009] Comparing the data volume of each threat log with a preset data volume threshold to obtain a first log set with high threat risk analysis value and a second log set with low threat risk analysis value, wherein the data volume of any threat log in the first log set is not less than the data volume threshold, and the data volume of any threat log in the second log set is less than the data volume threshold, where the data volume threshold is a preset effective communication data packet size critical value or a threat intelligence communication data basic packet size;

[0010] Based on the log packet data volume and the number of log packets in the first log set and the log packet data volume and the number of log packets in the second log set, a reverse estimation data leakage model is used to determine a reverse estimation data leakage amount, where the reverse estimation data leakage model is trained based on historical threat logs in historical threat intelligence data whose data leakage amount is not less than the threat threshold;

[0011] When the reverse estimated data leakage amount is not less than the set threat threshold, a threat risk warning message is generated.

[0012] In the above method, based on the first log set with high threat risk analysis value and the second log set with low threat risk analysis value, a reverse estimation data leakage model is used to calculate the reverse estimated data leakage amount. This can ensure that the reverse estimated data leakage amount is calculated accurately. At the same time, based on the comparison between the reverse estimated data leakage amount and the set threat threshold, it is determined whether to generate a threat risk warning message. This can ensure that threat risk warning messages are generated quickly and efficiently at any time point to ensure network security. At the same time, the present application realizes automatic detection of network security activities, obtains multiple communication logs of the target network to monitor the data leakage amount in real time, analyzes the data leakage risk status of the network security activities during the attack process, and thus determines whether there is a threat in the network security activities.

[0013] Optionally, the above-mentioned filtering of multiple communication logs based on historical threat intelligence data and the source IP addresses and destination IP addresses of multiple communication logs to obtain a threat log set that meets preset rules specifically includes:

[0014] Filter multiple communication logs based on historical threat intelligence data to obtain multiple communication logs with threat risk analysis value;

[0015] The confidence level of each communication log is calculated based on the source and destination IP addresses of multiple communication logs with threat risk analysis value, as well as a pre-trained two-way oscillation model based on the source and destination IP addresses of historical threat logs with data leakage volumes no less than the threat threshold.

[0016] A set of threat logs whose confidence levels are not less than a set confidence threshold is used as an initial threat log set;

[0017] The destination IP address of each threat log in the initial threat log set is determined, and a set consisting of multiple threat logs whose destination IP address is the preset target IP address is taken as the threat log set.

[0018] In the above method, the method of screening multiple communication logs based on historical threat intelligence data to obtain multiple threat logs with threat risk analysis value can realize data analysis of historical threat intelligence data and detect whether the acquired communication logs are threat logs in real time based on historical threat intelligence data. At the same time, by comparing the data volume of each threat log with a preset data volume threshold to obtain a first set of logs with high threat risk analysis value and a second set of logs with low threat risk analysis value, it can facilitate the subsequent extraction of key indicators for calculating the amount of data leakage. This facilitates more accurate calculation and estimation of data leakage, improving the accuracy and reliability of subsequent threat analysis of network security.

[0019] Optionally, the above reverse estimation data leakage model satisfies the following formula: Dbv=(Dbhv-Dbhlc*Dbv) / Pl

[0020] Among them, Dbv represents the reverse estimation of data leakage, Dbhv represents the data leakage volume of the first log set, Dbhlc represents the number of log packets in the first log set, and Pl represents the collection ratio of communication logs.

[0021] Optionally, the above Dbhv satisfies the following formula: Dbhv=Dbhlc*Dbtv

[0022] Among them, Dbtv represents basic data inclusion.

[0023] In the above method, the key indicator data extracted by analysis, namely the number of log packets in the first log set and the basic data inclusion calculation are used to obtain the data leakage amount of the first log set, which can determine the transmission data value with high threat risk analysis value in the data leakage traffic, so as to facilitate the subsequent more accurate reverse estimation of the data leakage amount.

[0024] Optionally, the above Dbtv satisfies the following formula: Dbtv=Dblv / Dbllc

[0025] Wherein, Dblv represents the data leakage volume of the second log set, and Dbllc represents the number of log packets in the second log set.

[0026] In this method, the basic data inclusion is calculated using the key metrics extracted from the analysis, namely the amount of data leakage and the number of log packets in the second log set. This allows the average basic traffic data inclusion of the traffic packets in the data leakage to be obtained, which is then used to accurately calculate the data leakage value of the first log set.

[0027] In a second aspect, the present application provides a network security threat analysis device, comprising:

[0028] an acquisition module configured to acquire multiple communication logs of a target network within a set time range, and filter the multiple communication logs based on historical threat intelligence data and source IP addresses and destination IP addresses of the multiple communication logs to obtain a threat log set that meets preset rules, wherein the threat log set includes multiple threat logs, and the historical threat intelligence data includes multiple historical threat logs with threat risk analysis value;

[0029] a processing module, configured to compare the data volume of each threat log with a preset data volume threshold to obtain a first log set with a high threat risk analysis value and a second log set with a low threat risk analysis value, wherein the data volume of any threat log in the first log set is not less than the data volume threshold, and the data volume of any threat log in the second log set is less than the data volume threshold, wherein the data volume threshold is a preset effective communication data packet size critical value or a threat intelligence communication data basic packet size;

[0030] The processing module is further configured to determine a reverse estimated data leakage amount using a reverse estimation data leakage model based on the log packet data volume and the number of log packets in the first log set and the log packet data volume and the number of log packets in the second log set, wherein the reverse estimation data leakage model is trained based on historical threat logs in the historical threat intelligence data whose data leakage amount is not less than the threat threshold;

[0031] The alarm module is used to generate a threat risk warning message when the reverse estimated data leakage amount is not less than a set threat threshold.

[0032] In a third aspect, the present application provides an electronic device comprising a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor executes the steps of the network security threat analysis method of the first aspect above.

[0033] In a fourth aspect, the present application provides a computer-readable storage medium comprising a program code. When the program code is run on an electronic device, the program code is used to enable the electronic device to execute the steps of the network security threat analysis method of the first aspect.

[0034] In a fifth aspect, the present application provides a computer program product, which, when called by a computer, enables the computer to execute the steps of the network security threat analysis method of the first aspect.

[0035] The technical effects brought about by any implementation method in the second to fifth aspects can refer to the technical effects brought about by the corresponding implementation method in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. In the drawings:

[0037] FIG1 is a schematic diagram of an optional system architecture applicable to an embodiment of the present application;

[0038] FIG2 is a schematic diagram of an implementation flow of a network security threat analysis method provided in an embodiment of the present application;

[0039] FIG3 is a schematic diagram of an implementation flow of another network security threat analysis method provided in an embodiment of the present application;

[0040] FIG4 is a schematic diagram of a visual mapping method provided in an embodiment of the present application;

[0041] FIG5 is a schematic diagram of the structure of a network security threat analysis device provided in an embodiment of the present application;

[0042] FIG6 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0043] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of the technical solutions of this application, but not all of them. Based on the embodiments described in this application document, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the technical solutions of this application.

[0044] It should be noted that in the description of this application, "multiple" is understood to mean "at least two." "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. A and B are connected, which can mean: A and B are directly connected, and A and B are connected through C. In addition, in the description of this application, words such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying an order.

[0045] In addition, the collection, dissemination, and use of data in the technical solution of this application comply with relevant national laws and regulations.

[0046] The following is a brief introduction to the design concept of the embodiment of this application:

[0047] Against the backdrop of informatization, digitization, and intelligentization, networks have become a crucial development direction. However, today's networks face numerous attack threats, such as Advanced Persistent Threats (APTs), information theft, and distributed denial of service (DDoS) attacks. These attacks often threaten highly confidential intelligence information in sectors such as politics, scientific research, military industry, aerospace, foreign trade, and economics, causing losses to various organizations. Therefore, timely threat analysis and early warning for network security threats have become a pressing issue.

[0048] To address the above-mentioned issues, embodiments of the present application provide a method, device, and electronic device for determining indicators based on IoT data. For example, multiple communication logs of a target network within a set time range are obtained, and based on historical threat intelligence data and the source IP addresses and destination IP addresses of the multiple communication logs, the multiple communication logs are screened to obtain a threat log set that meets preset rules. The threat log set includes multiple threat logs, and the historical threat intelligence data includes multiple historical threat logs with threat risk analysis value. The data volume of each threat log is compared with a preset data volume threshold to obtain a first log set with high threat risk analysis value and a second log set with low threat risk analysis value. The data volume of any threat log in the first log set is not less than the data volume threshold. The data volume of any threat log in the second log set is less than the data volume threshold. The data volume threshold is a preset critical value of the size of a valid communication data packet or the size of the basic packet of threat intelligence communication data. Based on the data volume and number of log packets of the log packets of the first log set and the data volume and number of log packets of the log packets of the second log set, a reverse estimation data leakage model is used to determine the reverse estimated data leakage amount. The reverse estimation data leakage model is trained based on historical threat logs in the historical threat intelligence data whose data leakage amount is not less than the threat threshold. Generate a threat risk warning message when the reverse estimated data leakage amount is not less than the set threat threshold

[0049] In this way, based on the first log set with high threat risk analysis value and the second log set with low threat risk analysis value, the reverse estimation data leakage model is used to calculate the reverse estimated data leakage amount. This can ensure that the reverse estimated data leakage amount is calculated accurately. At the same time, based on the comparison between the reverse estimated data leakage amount and the set threat threshold, it is determined whether to generate a threat risk warning message. This can ensure that threat risk warning messages are generated quickly and efficiently at any time to ensure network security. At the same time, the present application realizes automatic detection of network security activities, obtains multiple communication logs of the target network to monitor the data leakage amount in real time, analyzes the data leakage risk status of the network security activities during the attack process, and thus determines whether there is a threat in the network security activities.

[0050] In particular, the preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. In addition, the embodiments of the present application and the features in the embodiments may be combined with each other if there is no conflict.

[0051] Referring to FIG. 1 , which is a schematic diagram of a system architecture applicable to an embodiment of the present application, the system architecture includes a terminal 101 and a server 100. Information exchange between the terminal 101 and the server 100 can be performed via a communication network, wherein the communication network may employ wireless communication and wired communication.

[0052] Exemplarily, the terminal 101 can access the network and communicate with the server 100 via cellular mobile communication technology, wherein the cellular mobile communication technology includes, for example, the fifth generation mobile communication (5th Generation Mobile Networks, 5G) technology.

[0053] Optionally, the terminal 101 may access the network and communicate with the server 100 via short-range wireless communication, wherein the short-range wireless communication includes, for example, Wireless Fidelity (Wi-Fi) technology.

[0054] The embodiment of the present application does not impose any restriction on the number of communication devices involved in the above system architecture. For example, there may be more terminals 101, or no terminal 101, or other network devices may be included. As shown in Figure 1, only the terminal 101 and the server 100 are described as an example. The following is a brief introduction to the above devices and their respective functions.

[0055] The terminal 101 is a device that can provide voice and / or data connectivity to a user, and can be a device that supports wired and / or wireless connection.

[0056] Exemplarily, terminal 101 includes but is not limited to: mobile phones, tablet computers, laptop computers, PDAs, mobile Internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminal devices in industrial control, wireless terminal devices in unmanned driving, wireless terminal devices in smart grids, wireless terminal devices in transportation safety, wireless terminal devices in smart cities, or wireless terminal devices in smart homes, etc.

[0057] In addition, a related client can be installed on the terminal 101, and the client can be software, such as an application (APP), a browser, a short video software, etc., or a web page, a mini-program, etc.; it should be noted that in an embodiment of the present application, the terminal 101 can enable the above-mentioned client related to network security threat analysis to send multiple communication logs of the target network within a set time range to the server 100, so as to perform subsequent network security threat analysis and other method steps.

[0058] Server 100 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0059] In an embodiment of the present application, the server 100 can be used to obtain multiple communication logs of the target network within a set time range, and based on historical threat intelligence data, as well as the source IP addresses and destination IP addresses of the multiple communication logs, screen the multiple communication logs to obtain a threat log set that meets preset rules. The threat log set includes multiple threat logs, and the historical threat intelligence data includes multiple historical threat logs with threat risk analysis value. Then, the data volume of each threat log is compared with the preset data volume threshold to obtain a first log set with high threat risk analysis value and a second log set with low threat risk analysis value. The data volume of any threat log in the first log set is not less than the data volume threshold, and the data volume of any threat log in the second log set is less than the data volume threshold. The data volume threshold is a preset effective communication data packet size critical value or the threat intelligence communication data basic packet size. Based on the log packet data volume and log packet number of the first log set, the log packet data volume and log packet number of the second log set, a reverse estimation data leakage model is used to determine the reverse estimated data leakage amount. The reverse estimation data leakage model is trained based on historical threat intelligence data. Finally, when the reverse estimated data leakage amount is not less than the set threat threshold, a threat risk warning message is generated.

[0060] Optionally, in an embodiment of the present application, a model / device for network security threat analysis may be deployed on the server, that is, a network security threat analysis model / device pre-trained on the server may be used to implement threat risk index analysis of threat intelligence data.

[0061] The following describes the threat risk index analysis method provided by the exemplary embodiment of the present application in combination with the above-mentioned system architecture and with reference to the accompanying drawings. It should be noted that the above-mentioned system architecture is only shown to facilitate understanding of the spirit and principles of the present application, and the implementation of the present application is not limited in this respect.

[0062] Refer to FIG2 , which is a schematic diagram of an implementation process of a network security threat analysis method provided in an embodiment of the present application. The execution subject is a server as an example. The specific implementation process of the method is as follows:

[0063] S201: Acquire multiple communication logs of a target network within a set time range, and filter the multiple communication logs based on historical threat intelligence data and source IP addresses and destination IP addresses of the multiple communication logs to obtain a threat log set that meets preset rules.

[0064] The threat log set includes multiple threat logs, and the historical threat intelligence data includes multiple historical threat logs that are valuable for threat risk analysis.

[0065] Optionally, the communication log may also include log information, including but not limited to: the user's source IP address, destination IP address, access Uniform Resource Locator (URL) request time and response time, and other information.

[0066] It is understood that the above-mentioned time range and target network can be pre-set by those skilled in the art. The above-mentioned time range and target network can also be changed according to the specific application scenario. This application does not specifically limit this. For example, the above-mentioned time range can be: a time interval such as one day, several days, one week, or one month. The target network can be a backbone network. Among them, the backbone network is also called the core network, which is shared by all users, is responsible for transmitting backbone data, and is usually based on optical fiber, enabling large-scale data transmission.

[0067] It should be noted that the multiple historical threat logs with threat risk analysis value in the above-mentioned historical threat intelligence data can be used to extract a list of associated threatened assets, and to call a resource management system to locate the threatened assets.

[0068] In an optional embodiment, when executing step S201, after obtaining multiple communication logs, the server can filter the multiple communication logs based on historical threat intelligence data and the source IP addresses and destination IP addresses of the multiple communication logs to obtain a set of threat logs that meet preset rules. As shown in Figure 3, the following steps may be included:

[0069] S301. Filter multiple communication logs based on historical threat intelligence data to obtain multiple threat logs with threat risk analysis value.

[0070] Among them, historical threat intelligence data is stored in the historical threat database.

[0071] Specifically, the server can obtain historical threat intelligence data from a historical threat database. Based on multiple historical threat logs with threat risk analysis value in the historical threat intelligence data, the server can monitor and filter multiple communication logs to obtain multiple threat logs with threat risk analysis value.

[0072] In the above method, multiple communication logs are screened based on historical threat intelligence data to obtain multiple threat logs with threat risk analysis value. This method can realize data analysis of historical threat intelligence data and detect in real time whether the obtained communication logs are threat logs based on historical threat intelligence data.

[0073] S302. Calculate the confidence level of each threat log based on the source IP addresses and destination IP addresses of multiple threat logs with threat risk analysis value and a pre-trained two-way oscillation model for the source IP addresses and destination IP addresses of historical threat logs with data leakage amounts not less than a threat threshold.

[0074] In the embodiment of the present application, the confidence level is used to characterize the similarity between the threat log and the historical threat log whose data leakage amount is not less than the threat threshold. The higher the similarity, the higher the confidence level, and vice versa.

[0075] For example, the server can calculate the confidence of multiple threat logs with threat risk analysis value based on a two-way oscillation model to obtain corresponding confidence levels. This makes it easier to filter out threat logs with lower confidence levels (such as brute force logs and scanning logs) based on the confidence levels and retain threat logs with higher confidence levels.

[0076] S303: Taking a set of threat logs whose confidence levels are not less than a set confidence threshold as an initial threat log set.

[0077] Specifically, after determining the confidence level of each communication log valuable for threat risk analysis, the confidence level can be compared with a set confidence threshold. From the multiple communication logs valuable for threat risk analysis, multiple threat logs with confidence levels not less than the set confidence threshold are selected. This set of threat logs with confidence levels not less than the set confidence threshold is used as the initial threat log set.

[0078] For example, assuming the confidence threshold is set to 98%, when performing confidence analysis on multiple threat logs, if the confidence of any threat log is 92%, it can be determined that the confidence of the threat log is less than the set confidence threshold (i.e., 92% < 98%). This threat log does not belong to the initial threat log set.

[0079] On the contrary, if the confidence level of any threat log is 99%, it can be known that the confidence level of the threat log is not less than the set confidence level threshold (ie, 99%>98%) and the threat log belongs to the initial threat log set.

[0080] S304: Determine the destination IP address of each threat log in the initial threat log set, and take a set consisting of multiple threat logs whose destination IP address is the preset target IP address as the threat log set.

[0081] For example, assuming that the preset target IP address includes multiple IP addresses for communication directed to the asset, the server may select a set of threat logs in the initial threat log set whose destination IP address is the preset target IP address as the threat log set.

[0082] Optionally, after determining the threat log set, the server may extract threat log features of the threat log, which may include features such as log packet data volume, number of log packets, threat log type, source IP address, destination IP address, corresponding port, and time.

[0083] Log files are files used to record events occurring during the operation of an operating system or other application software, or messages between different users of communication software. Log files can be used to track and locate errors, debug and analyze code, and monitor application performance. Technicians can use the events recorded in log files to determine the operating status of the operating system or application software. When an operating system or application encounters an error or crash, analyzing the log files can reveal the underlying issue.

[0084] The threat log sets include a first log set with a high threat risk analysis value and a second log set with a low threat risk analysis value.

[0085] S202: Compare the data volume of each threat log with a preset data volume threshold to obtain a first log set with high threat risk analysis value and a second log set with low threat risk analysis value.

[0086] The data volume of any threat log in the first log set is no less than a data volume threshold. The data volume of any threat log in the second log set is less than a data volume threshold. The data volume threshold is a preset critical value of the size of a valid communication data packet or a basic packet size of threat intelligence communication data.

[0087] It is understood that the above-mentioned data volume threshold can be pre-set by those skilled in the art. The above-mentioned data volume threshold can also be changed according to the specific application scenario. This application does not specifically limit this. Moreover, the above-mentioned data volume threshold can be a different value set by those skilled in the art based on a set time range. For example, the data volume threshold can be 20MB / day. For another example, the data volume threshold can be 5MB / hour.

[0088] In one possible scenario, the data volume thresholds of different types of threat logs may be the same.

[0089] For example, assume the time range is set to one day. Assume the target network is a backbone network. Multiple communication logs from the backbone network are obtained within one day. From these multiple communication logs, a set of threat logs that meets the pre-set criteria is determined. Assuming the data volume threshold is set to 20MB / day, if a particular threat log in the set has a data volume of 10MB, it can be determined that the data volume of this threat log is less than the data volume threshold, i.e., 10MB < 20MB. Therefore, this threat log belongs to the second log set.

[0090] If the data volume of a threat log in the threat log set is 30MB, it can be determined that the data volume of this threat log is less than the data volume threshold, that is, 30MB > 20MB. Therefore, this threat log belongs to the first log set. In the above method, by comparing the data volume of each threat log with the preset data volume threshold to obtain a first log set with high threat risk analysis value and a second log set with low threat risk analysis value, it is possible to facilitate the subsequent extraction of key indicators for calculating the amount of data leakage. This facilitates more accurate calculation and estimation of the amount of data leakage, improving the accuracy and reliability of subsequent threat analysis of network security.

[0091] In another possible scenario, the data volume thresholds for different types of threat logs may be different.

[0092] For example, suppose a threat log set includes multiple different types of threat logs. The server can determine the corresponding data volume threshold based on the log type of the threat log. For example, the data volume threshold for the first threat log type is 28 MB / day. The data volume threshold for the second threat log type is 18 MB / day. Each threat log is then compared with the corresponding data volume threshold to determine whether it belongs to the first log set.

[0093] In the above method, by presetting data volume thresholds for different types of threat logs, it is possible to more accurately determine whether different types of threat logs are the first log set with high threat risk analysis value or the second log set with low threat risk analysis value.

[0094] S203 : Based on the log packet data volume and the number of log packets of the first log set and the log packet data volume and the number of log packets of the second log set, a reverse estimation data leakage model is used to determine a reverse estimation data leakage amount.

[0095] The reverse estimation data leakage model is obtained by training based on historical threat logs in which the amount of data leakage in historical threat intelligence data is not less than the threat threshold.

[0096] In one possible scenario, the historical threat intelligence data may include multiple historical threat logs for a certain type of network security incident. The reverse estimation data leakage model is trained based on the historical threat logs in which the data leakage amount is not less than the threat threshold in the historical threat intelligence data.

[0097] In this way, training a reverse estimation data leakage model based on multiple historical threat logs for a certain type of network security incident can make the subsequent reverse estimation of data leakage amount more targeted and improve the accuracy of the reverse estimation of data leakage amount.

[0098] In another possible scenario, the historical threat intelligence data may also include multiple historical threat logs of multiple types of network security incidents. The reverse estimation data leakage model is trained based on the historical threat logs in which the data leakage amount is not less than the threat threshold in the historical threat intelligence data.

[0099] In this way, the data leakage value during the attack process can be reversely estimated based on the collected historical threat intelligence data. This allows for relatively accurate reverse estimation of data leakage during threat event review without the need for terminal information and network traffic.

[0100] Optionally, the server may also identify multiple historical threat logs in the historical threat intelligence data whose destination IP address is a preset target IP address. A reverse estimation data leakage model is trained based on the multiple historical threat logs whose destination IP address is the preset target IP address. This allows the reverse estimation data leakage model to be more accurate in calculating the reverse estimation data leakage amount.

[0101] For example, assuming that the preset target IP address includes multiple IP addresses for communication directed to assets, a reverse estimation data leakage model is obtained by training multiple historical threat logs whose target IP addresses are the IP addresses for communication directed to assets.

[0102] Optionally, after determining the first log set, the server may determine the log packet data volume and the number of log packets of the first log set and the log packet data volume and the number of log packets of the second log set for subsequent reverse calculation and estimation of the data leakage amount.

[0103] Specifically, the reverse estimation data leakage model satisfies the following formula: Dbv = (Dbhv - Dbhlc * Dbv) / Pl

[0104] Among them, Dbv represents the reverse estimation of data leakage, Dbhv represents the data leakage volume of the first log set, Dbhlc represents the number of log packets in the first log set, and Pl represents the collection ratio of communication logs.

[0105] It will be understood that the communication log collection ratio Pl represents the ratio of the number of communication logs acquired within a set time range to the actual number of communication logs in the target network when acquiring multiple communication logs from the target network within a set time range. Pl is a collection ratio preset by those skilled in the art. Furthermore, Pl can be modified based on the specific application scenario. For example, Pl can be 1 / 5000. Another example is Pl can be 1 / 1000.

[0106] The above Dbhv satisfies the following formula: Dbhv=Dbhlc*Dbtv

[0107] Among them, Dbtv represents basic data inclusion.

[0108] In this method, the key indicator data extracted by analysis, namely the number of log packets in the first log set and the basic data inclusion calculation are used to obtain the data leakage amount of the first log set, which can determine the transmission data value with high threat risk analysis value in the data leakage traffic, so as to facilitate the subsequent more certain reverse estimation of the data leakage amount.

[0109] The above Dbtv satisfies the following formula: Dbtv=Dblv / Dbllc

[0110] Wherein, Dblv represents the data leakage volume of the second log set, and Dbllc represents the number of log packets in the second log set.

[0111] This method uses the extracted key metrics, namely the amount of data leakage from the second log set and the number of log packets in the second log set, to calculate basic data inclusion. This method can then obtain the average basic traffic data inclusion of the traffic packets in the data leakage, which is then used to accurately calculate the data leakage value for the first log set.

[0112] Obviously, by using the reverse estimation data leakage analysis model, in addition to detecting the current data leakage volume and risk status, it is also possible to make high-confidence data leakage risk assessments and warnings;

[0113] S204: When the reverse estimated data leakage amount is not less than the set threat threshold, a threat risk warning message is generated.

[0114] In the above method, by comparing the amount of data leakage with the set threat threshold, a warning of the threat can be issued in a timely manner, thereby playing a role in real-time detection of the network security threat risk status.

[0115] For example, suppose the threat threshold is set at 20MB / day. If the preset reverse estimation data leakage model is used to determine that the data leakage volume of threat intelligence data is 15MB / day, then a value threat risk warning message can be generated without targeting the threat intelligence data.

[0116] For example, suppose the threat threshold is set at 20MB / day. If the preset reverse estimation data leakage model is used to determine that the data leakage volume of threat intelligence data is 30MB / day, a value threat risk warning message will be generated for the threat intelligence data.

[0117] Optionally, the above-mentioned threat threshold may be set according to the threat intelligence level, threat log type, and accuracy of the threat log in the threat log set.

[0118] For example, after extracting the threat log features of the threat log to obtain the threat log type, the server determines the corresponding threat threshold based on the threat log type. Assume that the threat threshold corresponding to the first threat log type is 30MB / day. The threat threshold corresponding to the second threat log type is 20MB / day.

[0119] In this way, setting a threat threshold based on the threat intelligence level, threat type, and accuracy of the threat log collection can more accurately assess whether the reverse estimation of data leakage exceeds the corresponding threshold, achieving more accurate threat analysis at the network security level.

[0120] For example, assuming the threat is an APT attack, this can effectively display APT attack behavior analysis. Assume the threat threshold is set at 20MB / day. In a domestic attack by an APT organization, using the data leakage estimation model to estimate the data leakage volume, it can be determined that a true data leakage state (leakage volume > 20MB / day) only began on November 17, 2022, reaching the threat threshold and generating a threat risk warning message. The data leakage peaked from December 30, 2022, to January 3, 2023, primarily targeting the same attack target. A preliminary estimate indicates that 1.61075GB of data was leaked from the attack target. After January 3, 2023, data leakage from other attack targets gradually began. Although the data leakage volume did not reach a particularly high peak, the overall leakage volume was greater than before January 3, 2023, and the daily data leakage volume remained above the data leakage risk warning threshold.

[0121] It is understood that the above threat risk index and value threat risk index can be visualized and mapped, so that the threat activity status can be observed more intuitively. As shown in Figure 4, the present application provides a schematic diagram of visualization mapping.

[0122] Clearly, based on the above approach, it is possible to filter communication logs with threat risk analysis value, obtaining a first log set with high threat risk analysis value and a second log set with low threat risk analysis value. Furthermore, an automated method is implemented to determine the reverse estimated data leakage amount using a reverse estimation data leakage model based on the log packet data volume and number of log packets in the first log set, and the log packet data volume and number of log packets in the second log set. Furthermore, a determination is made based on the reverse estimated data leakage amount whether a threat risk warning needs to be generated. This enables timely and accurate network security threat analysis and threat warnings.

[0123] At the same time, from the perspective of network security construction, the technical solution of this application can be integrated into the cyberspace threat hunting system, which can fill the market demand of threat intelligence enterprise startups (To Business, To B). It can also fill the "vacuum zone" of Advanced Persistent Threat (APT) attacks. Based on the calculation of reverse estimation of data leakage, a new perspective of network security threat analysis is achieved, ensuring that threat risk warning messages are generated quickly and efficiently at any time to protect network security.

[0124] Furthermore, based on the same technical concept, the embodiment of the present application provides a threat risk index analysis device, which is used to implement the above-mentioned method flow of the embodiment of the present application. Referring to Figure 5, the threat risk index analysis device includes: an acquisition module 501, a processing module 502 and an alarm module 503, wherein:

[0125] An acquisition module 501 is configured to acquire multiple communication logs of a target network within a set time range, and filter the multiple communication logs based on historical threat intelligence data and the source IP addresses and destination IP addresses of the multiple communication logs to obtain a threat log set that meets preset rules, wherein the threat log set includes multiple threat logs, and the historical threat intelligence data includes multiple historical threat logs that are valuable for threat risk analysis;

[0126] Processing module 502 is configured to compare the data volume of each threat log with a preset data volume threshold to obtain a first log set with high threat risk analysis value and a second log set with low threat risk analysis value, wherein the data volume of any threat log in the first log set is not less than the data volume threshold, and the data volume of any threat log in the second log set is less than the data volume threshold, where the data volume threshold is a preset effective communication data packet size critical value or a threat intelligence communication data basic packet size;

[0127] The processing module 502 is further configured to determine a reverse estimated data leakage amount using a reverse estimation data leakage model based on the log packet data volume and the number of log packets in the first log set and the log packet data volume and the number of log packets in the second log set, where the reverse estimation data leakage model is trained based on historical threat logs in which the data leakage amount is not less than a threat threshold in the historical threat intelligence data;

[0128] The alarm module 503 is used to generate a threat risk warning message when the reverse estimated data leakage amount is not less than a set threat threshold.

[0129] Optionally, the above-mentioned filtering of multiple communication logs based on historical threat intelligence data and the source IP addresses and destination IP addresses of multiple communication logs to obtain a threat log set that meets preset rules, the acquisition module 501 is specifically used to:

[0130] Filter multiple communication logs based on historical threat intelligence data to obtain multiple communication logs with threat risk analysis value;

[0131] The confidence level of each communication log is calculated based on the source and destination IP addresses of multiple communication logs with threat risk analysis value, as well as a pre-trained two-way oscillation model based on the source and destination IP addresses of historical threat logs with data leakage volumes no less than the threat threshold.

[0132] A set of threat logs whose confidence levels are not less than a set confidence threshold is used as an initial threat log set;

[0133] The destination IP address of each threat log in the initial threat log set is determined, and a set consisting of multiple threat logs whose destination IP address is the preset target IP address is taken as the threat log set.

[0134] Optionally, the above reverse estimation data leakage model satisfies the following formula: Dbv=(Dbhv-Dbhlc*Dbv) / Pl

[0135] Among them, Dbv represents the reverse estimation of data leakage, Dbhv represents the data leakage volume of the first log set, Dbhlc represents the number of log packets in the first log set, and Pl represents the collection ratio of communication logs.

[0136] Optionally, the above Dbhv satisfies the following formula: Dbhv=Dbhlc*Dbtv

[0137] Among them, Dbtv represents basic data inclusion

[0138] Optionally, the above Dbtv satisfies the following formula: Dbtv=Dblv / Dbllc

[0139] Wherein, Dblv represents the data leakage volume of the second log set, and Dbllc represents the number of log packets in the second log set.

[0140] Based on the same technical concept, the embodiment of the present application also provides an electronic device that can implement the network security threat analysis method provided in the above embodiment of the present application. In one embodiment, the electronic device can be a server, a terminal device, or other electronic device. As shown in Figure 6, the electronic device may include:

[0141] At least one processor 601, and a memory 602 connected to at least one processor 601. In the embodiments of the present application, the specific connection medium between the processor 601 and the memory 602 is not limited. FIG6 takes the connection between the processor 601 and the memory 602 via the bus 600 as an example. The bus 600 is represented by a bold line in FIG6. The connection method between other components is only for schematic illustration and is not intended to be limiting. The bus 600 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, FIG6 only uses a bold line to represent it, but this does not mean that there is only one bus or one type of bus. Alternatively, the processor 601 can also be called a controller, and there is no limitation on the name.

[0142] In this embodiment of the present application, memory 602 stores instructions executable by at least one processor 601. At least one processor 601 can execute the threat risk index analysis method discussed above by executing the instructions stored in memory 602. Processor 601 can implement the functions of each module in the apparatus shown in FIG5.

[0143] Among them, the processor 601 is the control center of the device, which can use various interfaces and lines to connect the various parts of the entire control device, and monitor the device as a whole by running or executing instructions stored in the memory 602 and calling data stored in the memory 602, the various functions of the device and processing data.

[0144] In one possible design, processor 601 may include one or more processing units. Processor 601 may integrate an application processor and a modem processor. The application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily processes wireless communications. It is understood that the modem processor may not be integrated into processor 601. In some embodiments, processor 601 and memory 602 may be implemented on the same chip. In some embodiments, they may also be implemented on separate chips.

[0145] Processor 601 can be a general-purpose processor, such as a CPU, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the threat risk index analysis method disclosed in the embodiments of this application can be directly implemented and executed by a hardware processor, or by a combination of hardware and software modules in the processor.

[0146] The memory 602 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 602 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 602 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 602 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.

[0147] By designing and programming processor 601, the code corresponding to the network security threat analysis method described in the aforementioned embodiment can be embedded in the chip, thereby enabling the chip to execute the steps of the threat risk index analysis method of the embodiment shown in FIG2 during operation. Designing and programming processor 601 is well known to those skilled in the art and will not be further described here.

[0148] Based on the same inventive concept, an embodiment of the present application further provides a storage medium storing computer instructions. When the computer instructions are executed on a computer, the computer executes a threat risk index analysis method discussed above.

[0149] In some possible implementations, the present application also provides various aspects of a threat risk index analysis method that can also be implemented in the form of a program product, which includes program code. When the program product is run on an apparatus, the program code is used to enable the control device to execute the steps of a threat risk index analysis method according to various exemplary implementations of the present application described above in this specification.

[0150] It should be noted that although several units or subunits of the device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the application, the features and functions of two or more units described above can be embodied in a single unit. Conversely, the features and functions of a single unit described above can be further divided and embodied by multiple units.

[0151] Furthermore, although the operations of the method of the present application are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0152] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0153] The present application is described with reference to the flowchart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, and the combination of the process and / or box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a server, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the function specified in one process or multiple processes in the flowchart and / or one box or multiple boxes in the block diagram.

[0154] The program code used to perform the operations of the present application may be written using any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0155] Where a remote computing device is involved, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0156] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0157] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A network security threat analysis method, characterized in that: The method comprises: Acquire multiple communication logs of the target network within a set time range, and filter the multiple communication logs based on historical threat intelligence data and source IP addresses and destination IP addresses of the multiple communication logs to obtain a threat log set that meets preset rules, wherein the threat log set includes multiple threat logs, and the historical threat intelligence data includes multiple historical threat logs with threat risk analysis value; Compare the data volume of each threat log with a preset data volume threshold to obtain a first log set with a high threat risk analysis value and a second log set with a low threat risk analysis value, wherein the data volume of any threat log in the first log set is not less than the data volume threshold, and the data volume of any threat log in the second log set is less than the data volume threshold, and the data volume threshold is a preset effective communication data packet size critical value or a threat intelligence communication data basic packet size; Based on the log packet data volume and the number of log packets of the first log set and the log packet data volume and the number of log packets of the second log set, a reverse estimation data leakage model is used to determine the reverse estimation data leakage amount, wherein the reverse estimation data leakage model is obtained by training based on historical threat logs in which the data leakage amount in the historical threat intelligence data is not less than the threat threshold; When the reverse estimated data leakage amount is not less than the set threat threshold, a threat risk warning message is generated.

2. The method according to claim 1, characterized in that The filtering of the plurality of communication logs based on the historical threat intelligence data and the source IP addresses and the destination IP addresses of the plurality of communication logs to obtain a threat log set that meets a preset rule specifically includes: Filtering the plurality of communication logs based on the historical threat intelligence data to obtain a plurality of threat logs having threat risk analysis value; Calculate the confidence of each threat log according to the source IP addresses and destination IP addresses of the multiple threat logs with threat risk analysis value and the two-way oscillation model pre-trained for the source IP addresses and destination IP addresses of the historical threat logs whose data leakage amount is not less than the threat threshold; Taking a set of multiple threat logs whose confidence is not less than a set confidence threshold as an initial threat log set; The destination IP address of each threat log in the initial threat log set is determined, and a set consisting of multiple threat logs whose destination IP addresses are preset target IP addresses is taken as the threat log set.

3. The method according to claim 1, characterized in that The reverse estimation data leakage model satisfies the following formula: Dbv = (Dbhv-Dbhlc*Dbv) / Pl Among them, Dbv represents the reverse estimated data leakage amount, Dbhv represents the data leakage amount of the first log set, Dbhlc represents the number of log packets in the first log set, and Pl represents the collection ratio of communication logs.

4. The method according to claim 3, characterized in that The Dbhv satisfies the following formula: Dbhv = Dbhlc*Dbtv Among them, Dbtv represents basic data inclusion.

5. The method according to claim 4, characterized in that The Dbtv satisfies the following formula: Dbtv=Dblv / Dbllc Wherein, Dblv represents the data leakage volume of the second log set, and Dbllc represents the number of log packages in the second log set.

6. A network security threat analysis device, characterized in that: include: An acquisition module is used to acquire multiple communication logs of a target network within a set time range, and filter the multiple communication logs based on historical threat intelligence data and source IP addresses and destination IP addresses of the multiple communication logs to obtain a threat log set that meets preset rules, wherein the threat log set includes multiple threat logs, and the historical threat intelligence data includes multiple historical threat logs with threat risk analysis value; A processing module, used for comparing the data volume of each threat log with a preset data volume threshold to obtain a first log set with a high threat risk analysis value and a second log set with a low threat risk analysis value, wherein the data volume of any threat log in the first log set is not less than the data volume threshold, and the data volume of any threat log in the second log set is less than the data volume threshold, and the data volume threshold is a preset effective communication data packet size critical value or a threat intelligence communication data basic packet size; The processing module is further used to determine the reverse estimated data leakage amount by using a reverse estimated data leakage model based on the log packet data volume and the number of log packets of the first log set and the log packet data volume and the number of log packets of the second log set, wherein the reverse estimated data leakage model is obtained by training based on the historical threat intelligence data; The alarm module is used to generate a threat risk warning message when the reverse estimated data leakage amount is not less than a set threat threshold.

7. The device according to claim 6, characterized in that The method of filtering the multiple communication logs based on the historical threat intelligence data and the source IP addresses and destination IP addresses of the multiple communication logs to obtain a threat log set that meets the preset rules, wherein the acquisition module is specifically used for: Filtering the plurality of communication logs based on the historical threat intelligence data to obtain a plurality of communication logs having threat risk analysis value; Calculate the confidence of each communication log according to the source IP addresses and destination IP addresses of the plurality of communication logs having threat risk analysis value and a two-way oscillation model pre-trained for the source IP addresses and the destination IP addresses; Taking a set of multiple threat logs whose confidence is not less than a set confidence threshold as an initial threat log set; The destination IP address of each threat log in the initial threat log set is determined, and a set consisting of multiple threat logs whose destination IP addresses are preset target IP addresses is taken as the threat log set.

8. The device according to claim 7, characterized in that The reverse estimation data leakage model satisfies the following formula: Dbv = (Dbhv-Dbhlc*Dbv) / Pl Among them, Dbv represents the reverse estimated data leakage amount, Dbhv represents the data leakage amount of the first log set, Dbhlc represents the number of log packets in the first log set, and Pl represents the collection ratio of communication logs.

9. An electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Threat intelligence extraction method, system and equipment and readable storage medium

    CN110351280A

  • Threat level acquisition method and device of network attack and storage medium

    CN114124552A

  • Network security threat analysis method and device, electronic equipment and storage medium

    CN117955689A

  • Systems and methods for generating network threat intelligence

    US20150215334A1

Cited By

  • Network security equipment vulnerability scanning method and device and electronic equipment

    CN120675783A