Network security intrusion detection method and device
By establishing a communication topology relationship network and performing concurrent analysis of reference communication nodes, the traditional network intrusion detection method is solved, and the problem of low efficiency and poor real-time performance when dealing with large-scale network traffic is achieved, efficient and real-time network traffic security detection is achieved.
Patent Information
- Application Number
- CN202510223067.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Traditional network intrusion detection methods are inefficient and poor real-time when dealing with large-scale network traffic, making it difficult to meet the security needs of modern network environments.
By obtaining session segments and communication data packets in network traffic, establishing a communication topology relationship network, identifying reference communication nodes that support concurrent analysis, and performing concurrent analysis on them to improve detection efficiency and real-time.
It realizes rapid and accurate detection of network traffic, improves the real-timeness of security detection, saves the system's computing power, and improves the efficiency of network traffic security analysis.
Smart Images

Figure CN119728304B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a network security intrusion detection method and device. Background Art
[0002] With the rapid development of Internet technology and the widespread popularity of network applications, network security issues have become increasingly prominent and have become one of the key factors restricting the healthy development of the network. As an important means to ensure network security, the efficiency and accuracy of network intrusion detection are directly related to the security defense capabilities of the network system. Traditional network intrusion detection methods are often based on independent analysis of single data packets or session fragments. This method has problems such as low efficiency and poor real-time performance when processing large-scale network traffic, and it is difficult to meet the security needs in the modern network environment. Summary of the invention
[0003] The purpose of the present invention is to provide a network security intrusion detection method and device.
[0004] In the first aspect, the present invention provides a network security intrusion detection method, which is applied to a security intrusion detection device, and the method includes: obtaining network traffic to be analyzed; the network traffic covers several session segments, and the communication data packets in a session segment describe a complete network interaction process; establishing a communication topology relationship network based on the communication data packets of each session segment in the network traffic, and the association between different communication data packets; the communication topology relationship network covers several communication nodes and traffic pointers, a communication node contains a communication data packet of a session segment, and a traffic pointer represents the association between the communication data packets contained in the connected communication nodes; when performing network traffic security detection on the network traffic according to the communication topology relationship network, determining a reference communication node that supports concurrent analysis from the communication topology relationship network; performing concurrent analysis on the communication data packets contained in the reference communication node to obtain a traffic security analysis result of the network traffic.
[0005] In a second aspect, the present invention provides a security intrusion detection device, comprising: one or more processors; a memory; one or more computer programs; wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method as described above is implemented.
[0006] The embodiment of the present invention obtains the network traffic to be analyzed, the network traffic includes multiple session segments, and the communication data packet of a session segment describes a complete network interaction process; according to the communication data packets of each session segment in the network traffic and the association between different communication data packets, a communication topology relationship network is established. In the communication topology relationship network, a communication node contains a communication data packet, and a pointer represents the association between the communication data packets contained in different communication nodes. The present invention maps the network traffic into a communication topology relationship network, and can vividly and intuitively describe the association between the communication data packets of the network traffic according to the communication topology relationship network. When the network traffic is subjected to network traffic security detection according to the communication topology relationship network, a reference communication node supporting concurrent analysis can be determined in the communication topology relationship network, and the communication data packets stored in the reference communication node supporting concurrent analysis are subjected to concurrent analysis. According to the established communication topology relationship network, the reference communication node supporting concurrent analysis can be quickly and accurately determined. Concurrent analysis is performed on the reference communication node supporting concurrent analysis, so that the analysis of the communication data packets corresponding to each session segment in the network traffic is performed simultaneously, and the efficiency of the network traffic security analysis is greatly improved, so that the real-time performance of the security detection is enhanced, and the computing power of the system is saved. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 The present invention provides a flowchart of a network security intrusion detection method.
[0008] Figure 2 It is a schematic diagram of the composition of a security intrusion detection device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0009] The execution subject of the network security intrusion detection method in the embodiment of the present invention is a security intrusion detection device, including but not limited to a server, a personal computer, a laptop computer, a tablet computer, a smart phone, etc. The server includes but is not limited to a single network server, a server group consisting of multiple network servers, or a cloud consisting of a large number of computers or network servers in cloud computing. Among them, the security intrusion detection device accesses the network and interacts with other security intrusion detection devices in the network. Among them, the network where the security intrusion detection device is located includes but is not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a VPN network, etc.
[0010] The process of network traffic security analysis provided by the present invention can be mainly summarized as the following process: 1. Obtain the network traffic to be analyzed. The network traffic covers several session segments, and the communication data packets in a session segment describe a complete network interaction process. The network traffic contains session segments with associated situations. 2. A communication topology relationship network is established based on the communication data packets of each session segment in the network traffic, and the association between different communication data packets. The communication topology relationship network covers several communication nodes and traffic pointers. A communication node contains a communication data packet of a session segment, and a traffic pointer indicates the association between the communication data packets contained in the connected communication nodes. There may be an association between the communication data packets corresponding to the two communication nodes with traffic pointers in the communication topology relationship network. According to the association between the communication data packets, which is unidirectional, the traffic pointer may be a directional pointer, and the communication topology relationship network is directed acyclic. 3. When the network traffic is subjected to network traffic security detection according to the communication topology relationship network, a reference communication node supporting concurrent analysis is determined from the communication topology relationship network. The reference communication nodes supporting concurrent analysis include multiple, and there is no association between the communication data packets contained in the multiple reference communication nodes. This is manifested in the communication topology network as the absence of traffic pointers between the two reference communication nodes, and the two nodes are in the same hierarchical structure in the communication topology network. 4. Perform concurrent analysis on the communication data packets contained in the reference communication node to obtain the traffic security analysis results of the network traffic. For each communication node that can be analyzed concurrently in the communication topology network, during the network traffic security analysis, concurrent analysis is performed on the corresponding communication data packets to obtain the traffic security analysis results of each session segment. In addition, the communication topology network may also include communication nodes that are analyzed sequentially. When performing network traffic security analysis on these communication nodes, sequential analysis can be performed according to the degree of concern to obtain the corresponding traffic security analysis results.
[0011] like Figure 1 As shown, the network security intrusion detection method provided by the present invention comprises the following steps:
[0012] Step S110: Obtain the network traffic to be analyzed; the network traffic covers several session segments, and the communication data packets in a session segment describe a complete network interaction process.
[0013] In the embodiment of the present invention, network traffic refers to the total amount of data transmitted through the network within a certain period of time. It contains a large amount of information. Network traffic is usually composed of multiple session segments. A session segment represents a complete network interaction process, which includes all communication data packets from initiating a request to receiving a response. For example, when a user visits a web page, the browser will send an HTTP request data packet to the server. After receiving the request, the server will return the corresponding HTML, CSS, JavaScript and other response data packets. This process forms a session segment. In this session segment, each communication data packet contains key information such as the source address, destination address, protocol type, data content, etc., which is crucial for understanding the network interaction process.
[0014] In order to obtain the network traffic to be analyzed, the security intrusion detection device can deploy packet capture devices at key locations in the network, such as network switches, routers, or dedicated data acquisition devices. These devices can capture data packets passing through the network in real time and store them for subsequent analysis. For example, deploying a packet capture device at the network boundary can capture all data packets entering and leaving the network, thereby obtaining complete network traffic data.
[0015] After collecting network traffic data, the security intrusion detection device preprocesses the data. The main purpose of preprocessing is to extract session fragments and sort and organize them. This usually involves operations such as parsing of data packets, session reorganization, and deduplication. For example, for data packets of the TCP protocol, the security intrusion detection device can reorganize data packets belonging to the same TCP connection into a session fragment by analyzing the sequence number, confirmation number and other information of the data packet. For data packets of the UDP protocol, since it is connectionless, the security intrusion detection device can organize related data packets into a session fragment based on information such as the source address, destination address, and port number.
[0016] After extracting the session segments, the security intrusion detection device further analyzes the communication data packets in each session segment. This includes extracting key information of the data packets, such as source address, destination address, protocol type, data length, timestamp, etc., and organizing and storing this information. This information will be used to subsequently build a communication topology network and analyze abnormal behavior in network traffic.
[0017] It is worth noting that when acquiring network traffic data, the security intrusion detection device also considers the integrity and accuracy of the data. Since the amount of network traffic data is huge and the real-time requirements are high, the data collection and preprocessing process must be efficient and reliable. In addition, since there may be noise and interference in the network environment, such as packet loss, disorder, etc., the security intrusion detection device also uses appropriate data cleaning and verification technology to ensure that the acquired network traffic data is accurate and reliable.
[0018] During the execution of step S110, the security intrusion detection device can also use machine learning models to improve the efficiency and accuracy of data processing. For example, a clustering algorithm can be used to group data packets and group data packets belonging to the same session segment together; or a classification algorithm can be used to mark data packets to identify different types of network traffic (such as normal traffic, abnormal traffic, etc.). These machine learning models can be trained and optimized through training data sets to improve their performance and accuracy. In addition, when obtaining network traffic data, data privacy protection and compliance requirements are also considered. In particular, when processing data involving personal privacy or sensitive information, relevant laws, regulations and privacy policies must be strictly observed to ensure the legal and proper use of data.
[0019] Step S120: Establish a communication topology relationship network based on the communication data packets of each session fragment in the network traffic and the association between different communication data packets; the communication topology relationship network covers a number of communication nodes and traffic pointers, a communication node contains a communication data packet of a session fragment, and a traffic pointer represents the association between the communication data packets contained in the connected communication nodes.
[0020] The security intrusion detection device obtains the network traffic data that has been sorted in step S110. These data include several session segments, each of which is composed of a series of communication data packets. These communication data packets record a complete interaction process in the network, including key information such as source address, destination address, protocol type, data content, etc.
[0021] Next, analyze the association between these communication packets. Association usually refers to the logical relationship or dependency between packets, which may be based on a variety of factors, such as the time sequence of packets, protocol type, session identifier, etc. For example, in the TCP protocol, the association between packets can be determined by fields such as sequence number and confirmation number; in the HTTP protocol, the association between request packets and response packets can be established by request ID or session identifier.
[0022] After determining the association between communication data packets, the security intrusion detection device can start to build a communication topology network. This network is a directed acyclic graph consisting of several communication nodes and traffic pointers. Each communication node represents a set of communication data packets of a session segment, and the traffic pointer is used to represent the association between data packets between different communication nodes, that is, the edge.
[0023] In the process of building a communication topology relationship network, the security intrusion detection device performs the following operations: For each session segment in the network traffic, the security intrusion detection device creates a corresponding communication node. This node will contain all communication data packets of the session segment, as well as some basic attributes of these data packets, such as source address, destination address, protocol type, etc. For example, if the network traffic contains an HTTP session segment, the security intrusion detection device will create a communication node, which will contain all HTTP request and response data packets of this session segment. After creating the communication node, the security intrusion detection device analyzes the association between these nodes and determines the flow pointer. The flow pointer is directional, and they point from the source communication node to the destination communication node, indicating the association or dependency between the data packets. For example, in a TCP session, if data packet A is an acknowledgment packet of data packet B, then the security intrusion detection device will create a flow pointer in the communication topology relationship network from the communication node containing data packet A to the communication node containing data packet B. After determining all communication nodes and flow pointers, the security intrusion detection device can connect them to build a complete communication topology relationship network. This relationship network will intuitively display the communication structure and association between each session segment in the network traffic.
[0024] It is worth noting that in the process of constructing the communication topology network, the security intrusion detection device may handle some complex situations. For example, there may be cyclic dependencies or duplicate data packets in the network traffic, which may cause loops or redundant nodes in the communication topology network. In order to solve these problems, the security intrusion detection device can use some optimization algorithms or heuristic rules, such as depth-first search, breadth-first search, pruning algorithm, etc., to ensure that the constructed communication topology network is correct and efficient.
[0025] In addition, in the process of building a communication topology network, the security intrusion detection device can also use machine learning models to improve the accuracy and efficiency of analysis. For example, a clustering algorithm can be used to group communication nodes and cluster nodes with similar communication patterns or behavior characteristics; or a classification algorithm can be used to classify traffic pointers and identify different types of associations (such as sequential relationships, dependency relationships, etc.). These machine learning models can be trained and optimized through training data sets to improve their performance and accuracy.
[0026] After building the communication topology network, use it to perform network traffic security detection. By analyzing the communication nodes and traffic pointers in the communication topology network, the security intrusion detection device can identify abnormal behavior or intrusion attempts in the network traffic. For example, if a communication node suddenly receives a large number of data packets from unknown source addresses, or a traffic pointer points to a non-existent communication node, these may be signs of network intrusion.
[0027] Step S130: when performing network traffic security detection on network traffic according to the communication topology relationship network, a reference communication node supporting concurrent analysis is determined from the communication topology relationship network.
[0028] In the embodiment of the present invention, concurrent analysis refers to analyzing communication data packets in multiple communication nodes at the same time to improve detection efficiency and real-time performance. Compared with traditional sequential analysis, concurrent analysis can make full use of computing resources and shorten detection time, thereby better responding to rapidly changing network threats.
[0029] In step S130, the security intrusion detection device determines which communication nodes support concurrent analysis based on the established communication topology network. This process involves an in-depth understanding of the communication topology network structure and a detailed analysis of the association relationship between communication nodes.
[0030] First, the security intrusion detection device traverses the entire communication topology network and identifies all communication nodes and the traffic pointers between them. The traffic pointer represents the association of data packets between communication nodes and is key information for determining the feasibility of concurrent analysis. For example, in a communication topology network containing multiple HTTP session fragments, the traffic pointer may represent the correspondence between requests and responses, or the order relationship between different data packets in the same session.
[0031] Next, the security intrusion detection device determines which communication nodes support concurrent analysis based on certain rules or algorithms. This rule or algorithm is usually based on factors such as the degree of association between communication nodes, the dependency of data packets, and the characteristics of network traffic. In one implementation, the security intrusion detection device can use the concept of focus degree to measure the degree of attention of communication nodes when performing network traffic security detection. Communication nodes with higher focus degrees usually contain more important or more representative data packets, so they should be given priority when performing concurrent analysis.
[0032] The calculation of focality can be based on a variety of factors, such as the degree of the communication node (i.e., the number of traffic pointers connected to it), the centrality of the communication node (such as betweenness centrality, closeness centrality, etc.), the position of the communication node in the hierarchy, etc. For example, in a communication topology network, if a communication node is located at the intersection of multiple critical paths, then its focality may be high. Alternatively, if a communication node contains a large number of data packets from different source addresses, then its focality may also be high, because this may mean that the node is an important transit point or destination in the network.
[0033] In the process of determining the focus, the security intrusion detection device can also use machine learning models to improve the accuracy and efficiency of the calculation. For example, a supervised learning algorithm (such as support vector machine, random forest, etc.) can be used to train a focus prediction model that can predict the focus of the communication node based on its feature vector (such as degree, centrality, number of packets, etc.). The training process uses a set of labeled data sets, in which each communication node is assigned an actual focus value. By continuously iterating and optimizing model parameters, the security intrusion detection device can obtain a focus prediction model with better performance and apply it to the new communication topology relationship network.
[0034] In addition to the focus, the security intrusion detection device can also consider other factors to determine the reference communication nodes that support concurrent analysis. For example, it can be checked whether there is a direct flow pointer connection between the communication nodes. If there is, it indicates that there is a dependency between the data packets between them, and therefore it is not suitable for concurrent analysis. In addition, some thresholds or rules can be set according to the characteristics of the network traffic (such as traffic size, data packet type, etc.) to further filter out the communication nodes suitable for concurrent analysis.
[0035] After determining the reference communication nodes that support concurrent analysis, the security intrusion detection device can perform concurrent analysis on the communication packets in these nodes. The specific implementation of concurrent analysis may vary depending on the application scenario and technical requirements, but usually involves technical means such as multi-threaded programming and distributed computing. For example, multi-threading or thread pools can be used to simultaneously process packet analysis tasks for multiple communication nodes; or the analysis tasks can be distributed to multiple computing nodes, using distributed computing frameworks (such as Apache Spark, Hadoop, etc.) to achieve efficient concurrent processing.
[0036] Step S140: concurrently analyzing the communication data packets contained in the reference communication node to obtain a traffic security analysis result of the network traffic.
[0037] In the embodiment of the present invention, concurrent analysis is a technical means of parallel processing, which allows the security intrusion detection device to process multiple tasks at the same time, thereby significantly improving the processing speed and efficiency. In the field of network security, concurrent analysis is widely used in network traffic security detection. By processing multiple communication data packets in parallel, abnormal traffic patterns or potential intrusion behaviors can be quickly identified.
[0038] In step S140, the security intrusion detection device will perform concurrent analysis on the communication data packets contained in the reference communication nodes. These reference communication nodes are determined in step S130 through a series of algorithms and rules. There is no direct association between them, so they can be safely processed concurrently. The specific implementation of concurrent analysis may vary depending on the application scenario and technical requirements, but generally involves the following key steps:
[0039] 1. Task allocation: The security intrusion detection device first allocates concurrent analysis tasks to different processing units or threads. These processing units or threads can be multiple CPU cores, GPU cores in the security intrusion detection device, or multiple computing nodes in a distributed computing environment. The purpose of task allocation is to make full use of computing resources and achieve efficient parallel processing.
[0040] 2. Data packet preprocessing: Before concurrent analysis begins, the security intrusion detection device preprocesses each communication data packet. The preprocessing content may include data packet parsing, protocol identification, feature extraction, etc. For example, for TCP / IP protocol data packets, the security intrusion detection device parses out key fields such as source address, destination address, port number, sequence number, confirmation number, etc.; for HTTP protocol data packets, it extracts information such as request method, URL, response status code, etc. The purpose of preprocessing is to provide standardized data input for subsequent concurrent analysis.
[0041] 3. Concurrent analysis execution: After task allocation and data packet preprocessing are completed, the security intrusion detection device will begin to perform concurrent analysis. The specific algorithms for concurrent analysis may vary depending on the application scenario, but generally include pattern matching, statistical analysis, machine learning and other methods. For example, a pattern matching algorithm can be used to detect whether a data packet contains known attack fields or feature strings; a statistical analysis method can be used to identify abnormal traffic patterns, such as sudden increases in traffic, abnormal data packet sizes, etc.; and a machine learning model (such as a neural network, support vector machine, etc.) can be used to automatically learn and identify potential intrusion behaviors.
[0042] 4. Result aggregation and evaluation: After the concurrent analysis is completed, the security intrusion detection device aggregates and evaluates the results of each processing unit or thread. The purpose of aggregation is to integrate the scattered analysis results into a unified report or alarm; the purpose of evaluation is to verify and screen the analysis results to eliminate false positives and false negatives. For example, a majority voting mechanism can be used to integrate the judgments of multiple analysis results; threshold settings can be used to filter out alarms that are below a certain risk level.
[0043] In step S140, the machine learning model can learn the normal and abnormal patterns of network traffic through the training data set, so as to automatically identify potential intrusion behaviors during the concurrent analysis process. For example, a supervised learning algorithm (such as logistic regression, decision tree, etc.) can be used to train an intrusion detection model, which can predict whether it belongs to abnormal traffic based on the feature vector of the communication data packet (such as source address, destination address, port number, protocol type, data packet size, etc.). During the concurrent analysis process, the security intrusion detection device can input the feature vector of each communication data packet into the trained intrusion detection model to quickly obtain the judgment result of whether it is abnormal.
[0044] In addition, in order to further improve the efficiency and accuracy of concurrent analysis, security intrusion detection devices can also adopt some optimization strategies. For example, data parallelization technology can be used to accelerate the process of data packet preprocessing and feature extraction; model parallelization technology can be used to distribute machine learning models on multiple computing nodes for processing; load balancing technology can be used to ensure that the workload of each processing unit or thread is balanced to avoid overloading some nodes while other nodes are idle.
[0045] In one implementation, the communication topology relationship network established is an updated communication topology relationship network, and the traffic pointer contained in the updated communication topology relationship network has directionality. Then, step S120, based on the communication data packets of each session segment in the network traffic and the association between different communication data packets, establishes a communication topology relationship network, including:
[0046] Step S121: establishing an initial communication topology relationship network according to the communication data packets of each session segment in the network traffic and the association between different communication data packets, wherein the communication nodes included in the initial communication topology relationship network contain communication data packets, and the direction of the pointers between the communication nodes is from the communication data packet that initiates the communication to the communication data packet that receives the communication;
[0047] Step S122: updating the initial communication topology relationship network according to the communication nodes and pointers included in the initial communication topology relationship network to obtain an updated communication topology relationship network.
[0048] When executing step S121, the security intrusion detection device first parses the network traffic to extract each session segment and the communication data packets in each session segment. These data packets contain detailed information of network interaction, such as source address, destination address, protocol type, data content, etc.
[0049] Next, the security intrusion detection device analyzes the association between these communication data packets. The association may be based on a variety of factors, such as the time sequence of the data packets, the protocol type, the session identifier, etc. For example, in the TCP protocol, the association between data packets can be determined by fields such as the sequence number and the confirmation number; in the HTTP protocol, the association between the request data packet and the response data packet can be established by the request ID or the session identifier.
[0050] After determining the association between communication data packets, the security intrusion detection device can start to build the initial communication topology network. This network consists of several communication nodes and traffic pointers. Each communication node represents a set of communication data packets of a session segment, which contains all the data packets of the session segment and some basic attributes of these data packets (such as source address, destination address, protocol type, etc.). Traffic pointers are used to indicate the association between data packets between different communication nodes. They are directional, from the communication node that initiates the communication to the communication node that receives the communication.
[0051] For example, suppose that the network traffic contains an HTTP session segment, which consists of an HTTP request packet and an HTTP response packet. When constructing the initial communication topology network, the security intrusion detection device will create two communication nodes, representing the request packet and the response packet respectively. Then, it will establish a traffic pointer between the two nodes, from the communication node of the request packet to the communication node of the response packet, to indicate the association relationship between them.
[0052] Step S122 is an optimization and updating process based on the initial communication topology relationship network, and its purpose is to obtain a more concise and efficient communication topology relationship network by removing redundant information, merging duplicate nodes, etc.
[0053] When executing step S122, the security intrusion detection device first checks the communication nodes and traffic pointers in the initial communication topology network to identify possible redundant or repeated information, which may come from repeated transmission of network traffic, fragmentation of data packets, etc.
[0054] One type of redundant information is multiple different communication nodes that contain the same communication data packet content. In the initial communication topology network, these nodes may be mistakenly created as multiple nodes due to fragmentation or repeated transmission of network traffic. In order to eliminate this redundancy, the security intrusion detection device can compare the data packet contents of these nodes, and if they are exactly the same, they can be merged into one node and the direction of the traffic pointer can be updated accordingly.
[0055] For example, suppose the initial communication topology contains two communication nodes A and B, both of which contain the same HTTP response data packet. During the update process, the security intrusion detection device will recognize this and merge A and B into a new node C. Then, it will update all traffic pointers pointing to A or B so that they point to the new node C.
[0056] In addition to merging duplicate nodes, the security intrusion detection device can also update the communication topology network based on the traffic security analysis results in the historical data set. The historical data set contains unique identifiers of communication data packets for which the traffic security analysis results are known. By retrieving these identifiers, the security intrusion detection device can identify those communication nodes in the initial communication topology network that have been security analyzed. For these nodes, the security intrusion detection device can remove them and the traffic pointers associated with them because their analysis results are already known and do not need to be analyzed again.
[0057] For example, suppose the historical data set contains a unique identifier for an HTTP response packet, which matches a communication node D in the initial communication topology network. During the update process, the security intrusion detection device will recognize this and remove node D and all traffic pointers pointing to it. This can reduce the workload of subsequent security analysis and improve analysis efficiency.
[0058] In addition, it is also possible to identify and remove those no longer needed intermediate nodes based on the pointing relationship between the communication nodes and the traffic pointers in the initial communication topology network. These intermediate nodes may have played a role in forwarding or relaying in the transmission of network traffic, but they may no longer be necessary when constructing the communication topology network. By removing these intermediate nodes, the security intrusion detection device can further simplify the structure of the communication topology network and improve its readability and analysis efficiency.
[0059] For example, assume that the initial communication topology network includes an intermediate node E, which receives the communication data packet from node F and forwards it to node G. However, in the subsequent analysis process, the security intrusion detection device finds that nodes F and G can also communicate directly without passing through node E. Therefore, during the update process, the security intrusion detection device can remove node E and all traffic pointers related to it, thereby simplifying the structure of the communication topology network.
[0060] In one implementation, step S122, based on the communication nodes and pointers included in the initial communication topology relationship network, the initial communication topology relationship network is updated to obtain an updated communication topology relationship network, including:
[0061] Step S1221: Retrieving the traffic security analysis results of the communication data packets contained in the selected communication nodes in the initial communication topology relationship network from the historical data set;
[0062] Step S1222: if the traffic security analysis result of the communication data packet contained in the selected communication node is retrieved, then according to the distribution position of the selected communication node in the initial communication topology relationship network and the pointing of the pointer included in the initial communication topology relationship network, the associated communication node having a communication association with the selected communication node is obtained;
[0063] Step S1223: remove redundancy from the selected communication nodes and associated communication nodes to obtain an updated communication topology relationship network.
[0064] Step S1221 is the first step of updating the communication topology relationship network, and its purpose is to retrieve the traffic security analysis results of the communication data packets contained in the selected communication node in the initial communication topology relationship network from the historical data set. The historical data set is a database containing known traffic security analysis results, which records the previously analyzed communication data packets and their corresponding security analysis results.
[0065] When executing step S1221, the security intrusion detection device first determines which communication nodes are selected communication nodes for retrieval. These nodes may be selected based on a certain rule or algorithm, for example, they may contain a specific type of communication data packet, or they are in a specific position in the initial communication topology relationship network.
[0066] Next, the security intrusion detection device retrieves the unique identifiers of the communication data packets contained in these selected communication nodes in the historical data set. The unique identifier is a string or number used to uniquely identify the communication data packet, which usually contains some key information of the data packet, such as the source address, destination address, protocol type, timestamp, etc. By comparing these unique identifiers, the security intrusion detection device can determine whether the historical data set contains the traffic security analysis results corresponding to the selected communication node. For example, assume that the initial communication topology relationship network contains a communication node A, which contains an HTTP response data packet. In step S1221, the security intrusion detection device extracts the unique identifier of the HTTP response data packet (for example, a hash value generated based on its source address, destination address, protocol type and timestamp), and retrieves the matching records in the historical data set. If a matching record is found, it means that the traffic security analysis results of the HTTP response data packet are already included in the historical data set; if no matching record is found, it means that the data packet is new and further security analysis is performed.
[0067] Step S1222 is performed after the traffic security analysis result of the selected communication node is retrieved, and its purpose is to obtain the associated communication nodes that have communication association with the selected communication node according to the distribution position of the selected communication node in the initial communication topology relationship network and the pointing relationship of the traffic pointer. There is a direct or indirect communication relationship between these associated communication nodes and the selected communication node, and they may also contain communication data packets that are being analyzed or have been analyzed.
[0068] When executing step S1222, the security intrusion detection device first determines the position of the selected communication node in the initial communication topology relationship network. This usually involves traversing and searching the communication nodes to find all nodes directly connected to the selected communication node. Then, the security intrusion detection device further expands the search range according to the pointing relationship of the traffic pointer to find nodes indirectly connected to the selected communication node.
[0069] For example, assume that in step S1221, the security intrusion detection device retrieves the traffic security analysis results of communication node A. In step S1222, it will first find all nodes directly connected to communication node A (for example, nodes pointed to communication node A by traffic pointers, and nodes connected by traffic pointers from communication node A to other nodes). Then, it will continue to expand the search range to find the associated nodes of these directly connected nodes (for example, nodes connected by further traffic pointers). Finally, the security intrusion detection device will obtain a node set containing the selected communication node A and all its associated communication nodes.
[0070] Step S1223 is performed after the selected communication node and its associated communication nodes are obtained, and its purpose is to perform redundancy removal operations on these nodes to eliminate redundant information in the communication topology relationship network. Redundant information may come from repeated transmission of network traffic, fragmentation of data packets, forwarding of intermediate nodes, and other situations. Through the redundancy removal operation, the security intrusion detection device can simplify the structure of the communication topology relationship network and improve its readability and analysis efficiency.
[0071] When executing step S1223, the security intrusion detection device first analyzes the relationship between the selected communication node and its associated communication nodes. This includes checking whether there are repeated communication data packets between them, whether the network structure can be simplified by merging nodes, etc. Then, the security intrusion detection device performs corresponding redundancy removal operations according to the analysis results.
[0072] One redundancy removal operation is to merge duplicate nodes. If the selected communication node and its associated communication nodes contain duplicate communication data packets (i.e., their unique identifiers are the same), the security intrusion detection device can merge these nodes into one node and update the pointing relationship of the traffic pointer accordingly. This can eliminate redundant nodes and traffic pointers in the network and simplify the network structure.
[0073] For example, suppose that in step S1222, the security intrusion detection device obtains a node set including communication node A and its associated communication nodes. In this set, there are two nodes B and C that contain the same HTTP response data packet (that is, their unique identifiers are the same). In step S1223, the security intrusion detection device will recognize this and merge nodes B and C into a new node D. Then, it will update all traffic pointers pointing to node B or C so that they point to the new node D. At the same time, if node B or C also points to other nodes (such as node E), the security intrusion detection device will also update these pointing relationships accordingly so that they point from node D to node E.
[0074] In addition to merging duplicate nodes, the security intrusion detection device can also perform other types of redundancy removal operations based on the structural characteristics of the communication topology network. For example, if an intermediate node only plays the role of forwarding or relaying and does not process or modify the communication data packet, the security intrusion detection device can remove it and directly connect its previous and next nodes. Doing so can further simplify the network structure and improve analysis efficiency.
[0075] For example, suppose that in the communication topology network, there is an intermediate node F that only forwards HTTP request packets. It receives the request packet from node G and forwards it to node H. In step S1223, the security intrusion detection device can identify this and remove node F and all traffic pointers associated with it. Then, it will directly connect node G and node H so that they can communicate directly. Doing so can eliminate redundant nodes and traffic pointers in the network, making the network structure more concise and clear.
[0076] In summary, steps S1221, S1222 and S1223 are one of the important steps for constructing and optimizing the communication topology relationship network. Through the execution of these three steps, the security intrusion detection device can update and optimize the initial network based on the traffic security analysis results in the historical data set and the structural characteristics of the initial communication topology relationship network. Through the redundant removal operation, the security intrusion detection device can eliminate redundant information in the network, simplify the network structure, and improve the analysis efficiency. The updated communication topology relationship network obtained in this way provides strong support for subsequent network traffic security detection, so that the security intrusion detection device can more quickly and accurately identify potential network security threats.
[0077] In one implementation, the historical data set includes a unique identifier of a communication data packet for which the traffic security analysis result is known. Then, step S1221, in the historical data set, retrieves the traffic security analysis result of the communication data packet included in the selected communication node in the initial communication topology relationship network, including:
[0078] Step S12211: obtaining a reference unique identifier of a communication data packet contained in a selected communication node in the initial communication topology relationship network;
[0079] Step S12212: In the historical data set, a unique identifier that is consistent with the reference unique identifier is searched. If a unique identifier that is consistent with the reference unique identifier is retrieved in the historical data set, it is determined that the traffic security analysis results of the communication data packets contained in the selected communication node are retrieved; if a unique identifier that is consistent with the reference unique identifier is not retrieved in the historical data set, it is determined that the traffic security analysis results of the communication data packets contained in the selected communication node are not retrieved.
[0080] The purpose of step S12211 is to obtain the reference unique identifier of the communication data packet contained in the selected communication node in the initial communication topology relationship network. The reference unique identifier is the key information used to retrieve the traffic security analysis results in the historical data set. It must be able to uniquely identify the communication data packet so that the security intrusion detection device can accurately find the corresponding analysis results. When executing step S12211, the security intrusion detection device first determines which communication nodes are selected. These nodes may be selected based on certain rules or algorithms. For example, they may contain a specific type of communication data packet, or they are in a specific position in the initial communication topology relationship network. Once the communication nodes are selected, the security intrusion detection device extracts the reference unique identifier of the communication data packet from these nodes.
[0081] The method for generating the reference unique identifier can be selected according to the actual situation, but it must meet the requirements of uniqueness and stability. Uniqueness means that different communication data packets should have different unique identifiers to avoid confusion and misjudgment; stability means that even if the content of the communication data packet changes slightly (for example, the time stamp is updated), its unique identifier should remain unchanged to ensure that the analysis results in the historical data set are still valid.
[0082] One way to generate a reference unique identifier is to use a hash function. A hash function is a function that can map an input of any length (such as the content of a communication data packet) to an output of a fixed length. Since hash functions have good hashing and collision resistance, the generated hash value can be used as a unique identifier for a communication data packet. For example, the SHA-256 hash function can be used to hash the content of a communication data packet to obtain a 256-bit hash value as a reference unique identifier. In addition to hash functions, other methods can also be used to generate reference unique identifiers, such as strings constructed based on key fields of a communication data packet (such as source address, destination address, protocol type, etc.), digital signatures based on the content of the data packet, etc. Regardless of which method is used, it must be ensured that the generated unique identifier can uniquely identify the communication data packet and has sufficient stability.
[0083] For example, suppose that the initial communication topology network includes a communication node A, which includes an HTTP response data packet. In step S12211, the security intrusion detection device may select communication node A as the selected communication node and extract the reference unique identifier of the HTTP response data packet. If the unique identifier is generated by a hash function, the security intrusion detection device performs a SHA-256 hash operation on the content of the HTTP response data packet to obtain a hash value similar to "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855" as the reference unique identifier.
[0084] The purpose of step S12212 is to retrieve a unique identifier consistent with the reference unique identifier in the historical data set to determine whether the traffic security analysis results of the communication data packets contained in the selected communication node have been retrieved. The historical data set is a database containing known traffic security analysis results, which records previously analyzed communication data packets and their corresponding security analysis results. By retrieving the historical data set, the security intrusion detection device can quickly determine whether the selected communication node has been analyzed, thereby avoiding repeated analysis and improving detection efficiency. When executing step S12212, the security intrusion detection device will use the reference unique identifier obtained in step S12211 as a search keyword to search in the historical data set. The search process may involve database operation techniques such as querying database tables and using indexes to improve retrieval speed and accuracy.
[0085] If there is a unique identifier in the historical data set that is consistent with the reference unique identifier, it means that the communication data packets contained in the selected communication node have been security analyzed before, and the analysis results have been stored in the historical data set. At this time, the security intrusion detection device can directly read the analysis results without repeating the analysis of the communication data packets.
[0086] For example, assume that in step S12211, the security intrusion detection device obtains the reference unique identifier "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855" of the HTTP response data packet contained in communication node A. In step S12212, the security intrusion detection device searches for this hash value in the historical data set. If a matching record is found, it means that a security analysis has been performed on this HTTP response data packet, and the analysis results have been stored in the historical data set. At this point, the security intrusion detection device can directly read the analysis results and apply them to the subsequent network security detection process.
[0087] If there is no unique identifier consistent with the reference unique identifier in the historical data set, it means that the communication data packet contained in the selected communication node has not been security analyzed before, or the data packet is a new data packet (that is, its unique identifier is unique in the historical data set). At this time, the security intrusion detection device performs further security analysis on the communication data packet to determine whether it has a security risk.
[0088] For example, if no record matching the reference unique identifier "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855" is found in the historical data set, it means that no security analysis has been performed on this HTTP response data packet before. At this time, the security intrusion detection device may send it to the security analysis module for further processing and analysis to determine whether it has security risks such as attack fields and abnormal behaviors.
[0089] In summary, steps S12211 and S12212 are one of the key steps in constructing and optimizing the communication topology relationship network. Through the execution of these two steps, the security intrusion detection device can quickly retrieve the traffic security analysis results of the communication data packets contained in the selected communication nodes in the initial communication topology relationship network. If the analysis results are found, the results can be directly used to improve the detection efficiency; if the analysis results are not found, the communication data packets are further security analyzed. The implementation of this process is of great significance to improving the efficiency and accuracy of network traffic security detection.
[0090] In another implementation, step S122, according to the communication nodes and pointers included in the initial communication topology relationship network, the initial communication topology relationship network is updated to obtain an updated communication topology relationship network, including:
[0091] Step S122A: Based on the communication data packets contained in the communication nodes included in the initial communication topology relationship network, a plurality of different communication nodes having the same communication data packet content are determined in the initial communication topology relationship network;
[0092] Step S122B: removing the multiple different communication nodes and retaining one of the multiple different communication nodes;
[0093] Step S122C: Based on the distribution of multiple different communication nodes in the initial communication topology relationship network and the pointing of the pointers included in the initial communication topology relationship network, the communication nodes that have communication associations with the removed communication nodes and the remaining communication nodes are connected to obtain an updated communication topology relationship network.
[0094] The purpose of step S122A is to identify multiple different communication nodes with the same communication data packet content in the initial communication topology relationship network. These nodes may be mistakenly created as multiple nodes due to network traffic fragmentation, repeated transmission or other reasons, but in fact they contain the same communication data packet content. Identifying these nodes is a prerequisite for subsequent redundancy removal.
[0095] When executing step S122A, the security intrusion detection device first traverses each communication node in the initial communication topology relationship network to extract the communication data packet content they contain. Then, the security intrusion detection device compares the data packet content to identify communication data packets with the same content. The comparison process may involve operations such as hashing and feature extraction of the data packet content to improve comparison efficiency and accuracy.
[0096] Once multiple different communication nodes with the same communication data packet content are identified, the security intrusion detection device will record them for processing in subsequent steps. For example, suppose that the initial communication topology relationship network contains two communication nodes A and B, both of which contain an identical HTTP response data packet. In step S122A, the security intrusion detection device will traverse the initial communication topology relationship network, extract the communication data packet content contained in nodes A and B, and compare them. If the comparison result shows that the two data packet contents are the same, then the security intrusion detection device will identify nodes A and B as multiple different communication nodes with the same communication data packet content.
[0097] The purpose of step S122B is to remove multiple different communication nodes with the same communication data packet content identified in step S122A, and only retain one of the nodes. The purpose of this step is to eliminate redundant nodes in the network, simplify the network structure, and improve analysis efficiency. When executing step S122B, the security intrusion detection device determines which communication node to retain and removes other nodes. This decision may be based on a variety of factors, such as the creation time of the node, the location of the node, the importance of the node in the network, etc. However, in most cases, which node to retain will not have a substantial impact on the subsequent security analysis results, because all nodes with the same communication data packet content contain the same information.
[0098] Once the communication nodes to be retained are determined, the security intrusion detection device will remove other nodes with the same communication data packet content from the initial communication topology relationship network. The removal operation may involve operations such as updating the network structure and adjusting the pointing of the flow pointer. For example, in step S122A, the security intrusion detection device identifies that nodes A and B have the same communication data packet content. In step S122B, the security intrusion detection device may decide to retain node A and remove node B. In order to achieve this, the security intrusion detection device will update the structure of the initial communication topology relationship network, delete node B and its related flow pointer. At the same time, the security intrusion detection device also adjusts the pointing of the flow pointer connected to node B so that they point to the retained node A.
[0099] The purpose of step S122C is to reconnect the associated communication nodes of the removed communication nodes and the remaining communication nodes after removing the redundant nodes to ensure the connectivity and integrity of the network. This step is a key link in updating the communication topology relationship network, which ensures that the structure and function of the network will not be damaged while eliminating redundant information.
[0100] When executing step S122C, the security intrusion detection device first determines the associated communication nodes of the removed communication node. These nodes are nodes directly connected to the removed node in the initial communication topology relationship network. Then, the security intrusion detection device will reconnect these associated communication nodes with the remaining communication nodes based on the pointing relationship of the flow pointer. The connection process may involve operations such as creating a new flow pointer and adjusting the pointing of an existing flow pointer.
[0101] For example, in step S122B, the security intrusion detection device removes node B and retains node A. In step S122C, the security intrusion detection device finds the associated communication nodes (assuming they are node C and node D) directly connected to node B, and reconnects them with the retained node A. To achieve this, the security intrusion detection device may create a new traffic pointer from node A to node C and node D to indicate the communication association between them. At the same time, the security intrusion detection device also deletes or adjusts the old traffic pointer associated with node B to ensure the connectivity and consistency of the network.
[0102] In summary, steps S122A, S122B and S122C are one of the important steps in building and optimizing the communication topology relationship network. Through the execution of these three steps, the security intrusion detection device can identify and remove redundant nodes in the initial communication topology relationship network while ensuring the connectivity and integrity of the network. The implementation of this process is of great significance for improving the efficiency and accuracy of network traffic security detection. By eliminating redundant information and simplifying the network structure, the security intrusion detection device can more quickly identify potential network security threats and take corresponding defensive measures.
[0103] In one implementation, in step S130, determining a reference communication node supporting concurrent analysis from the communication topology relationship network includes:
[0104] Step S131: determining the focus of each communication node based on the communication topology relationship network, wherein the focus represents the degree of attention paid to the communication data packets contained in the corresponding communication node when performing network traffic security detection, and at the same time, the communication data packets contained in the communication nodes with the same focus have the same degree of attention paid to the network traffic security detection;
[0105] Step S132: Determine the communication nodes corresponding to the same focus as reference communication nodes supporting concurrent analysis.
[0106] The purpose of step S131 is to determine the focus of each communication node based on the communication topology relationship network. Focus is a measurement indicator that indicates the degree of attention paid to the communication data packets contained in the corresponding communication node when performing network traffic security detection. The higher the focus of the communication node, the more important the communication data packets contained in it are when performing security detection, and more attention is given to it.
[0107] When executing step S131, the security intrusion detection device comprehensively considers various factors in the communication topology relationship network to determine the focus of each communication node. These factors may include the degree of the communication node (i.e., the number of communication nodes directly connected to it), the centrality of the communication node (such as betweenness centrality, closeness centrality, etc.), the position of the communication node in the network (such as whether it is on the critical path), the number and content of communication data packets contained in the communication node, etc.
[0108] For example, suppose that the communication topology network contains a communication node A, which is directly connected to multiple other communication nodes and is on the critical path in the network. In addition, communication node A also contains a large number of communication data packets, some of which involve sensitive information or high-risk operations. Based on these factors, the security intrusion detection device may give a higher focus to communication node A, indicating that it pays more attention when performing network traffic security detection.
[0109] In order to more accurately calculate the focus of the communication node, the security intrusion detection device can adopt some mathematical models or algorithms. For example, the betweenness centrality algorithm can be used to calculate the centrality of the communication node in the network. Nodes with higher betweenness centrality usually have higher focus. In addition, a weight function can be defined based on the number and content of communication data packets contained in the communication node, taking into account factors such as the number, type, and importance of the data packets, so as to calculate a comprehensive focus value. The calculation formula for focus varies depending on the specific application scenario, but in general, it should be able to comprehensively reflect the importance and attention of the communication node in network traffic security detection. For example, a possible focus calculation formula is:
[0110] ;
[0111] in, , , , is the weight coefficient, which is used to adjust the contribution of each factor in the focus calculation. These weight coefficients can be adjusted and optimized according to actual conditions.
[0112] The purpose of step S132 is to determine the communication nodes corresponding to the same focus as the reference communication nodes supporting concurrent analysis. In concurrent analysis, multiple communication nodes can be processed and analyzed simultaneously to improve detection efficiency. However, in order to avoid dependencies and conflicts between data packets, those communication nodes that are not directly associated or dependent are selected for concurrent analysis.
[0113] When executing step S132, the security intrusion detection device will first check the focus of each communication node and group the communication nodes with the same focus. Then, for each group of communication nodes with the same focus, the security intrusion detection device will further check the association relationship between them. If there is no direct traffic pointer connecting two communication nodes (i.e., there is no direct communication association between them), and their attention levels in network traffic security detection are consistent (i.e., their focus is the same), then the two communication nodes can be regarded as reference communication nodes that support concurrent analysis.
[0114] For example, suppose there are two communication nodes B and C in the communication topology network, and their focus is 0.8 (assuming that the focus ranges from 0 to 1). At the same time, there is no direct traffic pointer between communication nodes B and C, that is, there is no direct communication association between them. Based on this information, the security intrusion detection device may determine communication nodes B and C as reference communication nodes that support concurrent analysis.
[0115] When determining the reference communication node, the security intrusion detection device may also consider other factors, such as the processing time of the communication node, the computing resource requirements, etc. If two communication nodes have similar processing times and similar computing resource requirements, they are more suitable for concurrent analysis. In addition, in order to balance the efficiency and accuracy of concurrent analysis, the security intrusion detection device may also set a concurrent analysis threshold, and concurrent analysis is performed only when the number of communication nodes exceeds the threshold.
[0116] In summary, steps S131 and S132 are key steps in determining reference communication nodes that support concurrent analysis. By comprehensively considering factors such as the focus, association, processing time, and computing resource requirements of communication nodes, the security intrusion detection device can more reasonably select and determine reference communication nodes, thereby improving the efficiency and accuracy of network traffic security detection. The implementation of this process is of great significance for improving network security protection capabilities and ensuring the normal operation of the network.
[0117] In one implementation, in step S131, determining the focus of each communication node based on the communication topology relationship network includes:
[0118] Step S1311: obtaining the distribution level of the selected communication node in the communication topology relationship network;
[0119] Step S1312: Determine the focus degree of the selected communication node according to the distribution level of the selected communication node.
[0120] The purpose of step S1311 is to obtain the distribution level of the selected communication node in the communication topology network. The distribution level is an indicator that describes the position of a communication node in the network structure, which reflects the importance and influence of the node in the network. Generally speaking, nodes at the center of the network or on the critical path have a higher distribution level because they play a more important role in the network and carry more communication traffic and information transmission.
[0121] When executing step S1311, the security intrusion detection device first determines the structural characteristics of the communication topology relationship network, including the number of nodes, the number of edges, the degree distribution of nodes, etc. Then, the security intrusion detection device will use a hierarchical division algorithm, such as breadth-first search (BFS), depth-first search (DFS) or PageRank algorithm, to traverse the communication topology relationship network and divide it into different levels according to the connection relationship and importance of the nodes.
[0122] For example, assuming that the communication topology network is a directed acyclic graph (DAG), the security intrusion detection device can use a breadth-first search algorithm to determine the distribution level of the nodes. The algorithm starts from a source node in the network and expands outward layer by layer until all nodes are traversed. In each layer, the algorithm records all the nodes in the next layer that are directly connected to the nodes in that layer, and builds the hierarchical structure of the network in turn. In this way, the security intrusion detection device can obtain the distribution level of each communication node in the communication topology network.
[0123] It is worth noting that the division of distribution levels is not absolute, and it may be affected by many factors, such as the scale, structure, and degree of nodes of the network. Therefore, in practical applications, the security intrusion detection device may adjust and optimize the level division algorithm according to the specific situation to ensure that the obtained distribution level can accurately reflect the importance of the communication node in the network.
[0124] Step S1312 is a step of further determining the focus of the communication node after obtaining the distribution level of the communication node. The focus is a quantitative indicator used to measure the attention and importance of the communication node in network traffic security detection. In step S1312, the security intrusion detection device calculates the focus of the communication node according to the distribution level of the communication node. Generally, the node with a higher distribution level has a higher focus.
[0125] The calculation of the focus can be based on a variety of factors, such as the node distribution level, node degree, node centrality, node position in the network, etc. Among them, the distribution level is an important consideration because it directly reflects the position and importance of the node in the network. In order to calculate the focus, the security intrusion detection device can use a weighted summation method to assign different weights to different factors according to their importance, and add the weighted values of each factor to obtain the final focus value.
[0126] For example, suppose the following formula is used to calculate the focus of a communication node:
[0127] ;
[0128] in, , , is a weight coefficient used to adjust the contribution of various factors in the focus calculation. These weight coefficients can be adjusted and optimized according to actual conditions to ensure that the focus can accurately reflect the importance of the communication node in network traffic security detection.
[0129] In the specific calculation process, the security intrusion detection device substitutes the communication node distribution level obtained in step S1311 into the focus calculation formula. At the same time, the security intrusion detection device also calculates the degree (i.e., the number of nodes directly connected to it) and centrality (such as betweenness centrality, closeness centrality, etc.) of the node, and substitutes these values into the formula. Finally, the security intrusion detection device performs a weighted summation of various factors according to the weight coefficient to obtain the focus value of the communication node.
[0130] It is worth noting that the calculation of focus is not static, it may change with factors such as changes in network traffic, the addition of new nodes or the removal of old nodes. Therefore, in practical applications, the security intrusion detection device regularly updates the focus value of the communication node to ensure that it can accurately reflect the needs and focus of the current network traffic security detection.
[0131] In addition, in order to further improve the accuracy and reliability of focus calculation, the security intrusion detection device can also introduce a machine learning model to assist in the calculation. For example, a supervised learning algorithm can be used to train a focus prediction model that can predict the focus value of a communication node based on its characteristic vector (such as distribution level, node degree, centrality, etc.). By continuously iterating and optimizing model parameters, the security intrusion detection device can obtain a focus prediction model with better performance and apply it to actual network traffic security detection.
[0132] In one implementation, if the communication topology network includes a communication node that is consistent with the distribution level of the selected communication node and is different from the source communication node of the selected communication node, the method further includes:
[0133] Step S13123: Obtain the number of first communication nodes corresponding to the same source communication node as the selected communication node, and the number of second communication nodes corresponding to different source communication nodes and having the same distribution level as the selected communication node, wherein both the first communication node and the second communication node are communication nodes corresponding to communication data packets that have not been analyzed. If the number of first communication nodes is greater than the number of second communication nodes, the focus of the selected communication node is greater than the focus of communication nodes corresponding to different source communication nodes and having the same distribution level as the selected communication node; if the number of first communication nodes is less than the number of second communication nodes, the focus of the selected communication node is less than the focus of communication nodes corresponding to different source communication nodes and having the same distribution level as the selected communication node; if the number of first communication nodes is equal to the number of second communication nodes, the focus of the selected communication node is equal to the focus of communication nodes corresponding to different source communication nodes and having the same distribution level as the selected communication node.
[0134] Step S13123 is performed when there are communication nodes in the communication topology relationship network that are consistent with the distribution level of the selected communication node but different from the source communication node. Its purpose is to further adjust the focus of the selected communication node by comparing the number of these communication nodes. The execution of this step depends on the in-depth analysis of the communication topology relationship network and the accurate understanding of the relationship between the communication nodes.
[0135] When executing step S13123, the security intrusion detection device determines all communication nodes with the same distribution level as the selected communication node. These nodes are in the same position or level in the network structure as the selected communication node, so their importance in network traffic security detection may be similar. However, since these nodes may have different source communication nodes, that is, they may come from different network paths or session fragments, their focus may be adjusted according to other factors.
[0136] Next, the security intrusion detection device divides these communication nodes that are in the same distribution level as the selected communication node into two categories: the first category is the communication nodes that correspond to the same source communication node as the selected communication node (i.e., the first communication node), and the second category is the communication nodes that correspond to different source communication nodes as the selected communication node (i.e., the second communication node). The difference between these two types of communication nodes is whether their source communication nodes are the same, that is, whether they come from the same network path or session segment.
[0137] Then, the security intrusion detection device calculates the number of the first communication node and the second communication node respectively. These node numbers reflect the number of communication nodes that are consistent with the selected communication node in the distribution level but have different source communication nodes. By comparing the number of these two types of nodes, the security intrusion detection device can further determine the importance of the selected communication node in the network traffic security detection.
[0138] Specifically, if the number of the first communication nodes is greater than the number of the second communication nodes, it means that there are more communication nodes that have the same source communication node as the selected communication node. This means that the selected communication node may play a more important role in the network because it shares the same source communication node with more communication nodes. Therefore, in this case, the security intrusion detection device may increase the focus of the selected communication node to reflect its higher importance in network traffic security detection.
[0139] On the contrary, if the number of the first communication nodes is less than the number of the second communication nodes, it means that there are more communication nodes having different source communication nodes from the selected communication node. This means that the selected communication node may not be so important in the network because it shares the same source communication node with fewer communication nodes. Therefore, in this case, the security intrusion detection device may reduce the focus of the selected communication node.
[0140] If the number of the first communication nodes is equal to the number of the second communication nodes, it means that the number of communication nodes having the same source communication node and different source communication nodes as the selected communication node is equal. In this case, the security intrusion detection device may keep the focus of the selected communication node unchanged because the number of the two types of nodes is balanced and there is not enough information to support the increase or decrease of the focus.
[0141] It is worth noting that the focus adjustment in step S13123 is a relative process, which relies on the comparison of the number of communication nodes with the same distribution level as the selected communication node but with different source communication nodes. This comparison provides an additional dimension to evaluate the importance of the selected communication node, but it is not the only factor in determining the focus. In practical applications, the security intrusion detection device may combine other factors (such as the degree, centrality, data packet content, etc. of the node) to comprehensively determine the focus of the communication node.
[0142] In addition, the execution of step S13123 is not limited to the case of a single selected communication node. In actual network traffic security detection, the security intrusion detection device may process multiple selected communication nodes at the same time and execute step S13123 on them to refine the determination process of the focus. In this case, the security intrusion detection device may use an iterative or parallel processing method to efficiently execute step S13123.
[0143] The focus adjustment in step S13123 is a dynamic process. The focus of the communication node may change with the influence of factors such as changes in network traffic, the addition of new nodes or the removal of old nodes. Therefore, the security intrusion detection device regularly updates the focus value of the communication node to ensure that it can accurately reflect the needs and priorities of the current network traffic security detection.
[0144] In summary, step S13123 further refines the process of determining the focus by comparing the number of communication nodes with the same distribution level as the selected communication node but with different source communication nodes. The execution of this step depends on the in-depth analysis of the communication topology relationship network and the accurate understanding of the relationship between the communication nodes. Through the execution of step S13123, the security intrusion detection device can more accurately evaluate the importance of the communication node in the network traffic security detection, and provide strong support for the subsequent selection of reference communication nodes that support concurrent analysis.
[0145] In one implementation, the traffic security analysis result of the network traffic includes the traffic security analysis result of each session segment in the network traffic; and the method further includes:
[0146] Step SA: according to the analysis status of network traffic security detection for each session segment in the network traffic, an analysis tag is assigned to the communication node corresponding to the corresponding session segment in the communication topology relationship network; the analysis tag includes an analysis end tag, an unanalyzed tag and an analysis ongoing tag;
[0147] Step SB: According to the analysis tags added to each communication node in the communication topology relationship network, the unique identifier of the session segment in the communication node assigned with the analysis end tag and the corresponding traffic security analysis result are obtained;
[0148] Step SC: associate the obtained unique identifier with the corresponding traffic security analysis result and add them to the historical data set.
[0149] The purpose of step SA is to assign analysis tags to the communication nodes corresponding to the corresponding session segments in the communication topology network according to the analysis status of the network traffic security detection of each session segment in the network traffic. The analysis tag is a mark used to identify the current analysis status of the communication node, which can help the security intrusion detection device quickly understand the processing progress and results of each communication node. There are three main types of analysis tags: analysis end tag, unanalyzed tag and analysis in progress tag.
[0150] When executing step SA, the security intrusion detection device first performs security detection on each session segment in the network traffic. The detection process may involve content analysis, pattern matching, statistical analysis, or application of machine learning models to communication data packets. Based on the detection results, the security intrusion detection device determines whether each session segment has a security risk and generates a corresponding traffic security analysis result.
[0151] Next, the security intrusion detection device will find the corresponding communication node in the communication topology network according to the security detection results of the session segment and assign it a corresponding analysis tag. If the security detection of the session segment has been completed and no security risks have been found, the corresponding communication node will be assigned an analysis end tag. If the session segment has not been analyzed, or is being analyzed but not completed, the corresponding communication node will be assigned an unanalyzed tag or an analysis in progress tag.
[0152] For example, suppose that the network traffic contains an HTTP session segment, and the security intrusion detection device performs a security check on the session segment and finds that it contains a malware download link. In this case, the security intrusion detection device will find the communication node corresponding to the HTTP session segment in the communication topology relationship network, assign an analysis end tag to it, and record the security analysis result that the session segment has a security risk.
[0153] The process of assigning analysis tags not only helps the security intrusion detection device track the analysis progress and results of each session segment, but also provides a basis for subsequent optimization and scheduling. For example, the security intrusion detection device can prioritize those session segments that have not been analyzed based on the analysis tags, or allocate resources and monitor the progress of the session segments being analyzed.
[0154] Step SB is the second step of this derivative implementation scheme, and its purpose is to obtain the unique identifier of the session fragment in the communication node assigned the analysis end label and the corresponding traffic security analysis result based on the analysis label added to each communication node in the communication topology relationship network. The purpose of this step is to extract the session fragment information that has been analyzed and confirmed to be safe for subsequent processing and storage. When executing step SB, the security intrusion detection device will traverse all communication nodes in the communication topology relationship network and check their analysis labels. For those communication nodes that are assigned the analysis end label, the security intrusion detection device will further extract their corresponding session fragment information. This information usually includes a unique identifier of the session fragment (such as a hash value generated based on the session content) and the corresponding traffic security analysis result (such as security, risk, etc.).
[0155] The traffic security analysis result is the conclusion drawn after security testing of the session fragment, which usually includes information such as whether there is a security risk in the session fragment, the specific type of security risk, and the risk level. This information is of great significance for subsequent security policy formulation, incident response, and threat intelligence sharing. For example, suppose that a communication node in the communication topology network is assigned an analysis end tag, and the corresponding session fragment is an FTP session fragment. The security intrusion detection device will extract the session fragment information of the communication node, including its unique identifier (such as a hash value generated based on the FTP session content) and the corresponding traffic security analysis result (such as "safe"). Then, this information will be used for subsequent processing and storage.
[0156] The purpose of step SC is to associate the unique identifier obtained in step SB with the corresponding traffic security analysis result and add it to the historical data set. The historical data set is a database containing known traffic security analysis results, which records the previously analyzed session segments and their corresponding security analysis results. By associating and adding new analysis results to the historical data set, the security intrusion detection device can build a comprehensive security knowledge base to provide strong support for subsequent security detection and analysis.
[0157] When executing step SC, the security intrusion detection device first creates a new record item for storing the unique identifier obtained in step SB and the corresponding traffic security analysis result. Then, the security intrusion detection device associates and adds this new record item to the historical data set. The association process may involve operations such as index update and data table insertion of the historical data set to ensure that the new analysis result can be quickly retrieved and queried.
[0158] The construction and maintenance of historical data sets are of great significance to the long-term operation and continuous improvement of network security intrusion detection methods. By continuously accumulating and analyzing historical data, security intrusion detection devices can learn the normal and abnormal patterns of network traffic and improve the accuracy and efficiency of security detection. At the same time, historical data sets can also be used to generate security reports, share threat intelligence, and support the formulation of security policies.
[0159] For example, assume that in step SB, the security intrusion detection device extracts a unique identifier of an HTTP session segment and the corresponding traffic security analysis result (such as "there is a risk of SQL injection"). In step SC, the security intrusion detection device will associate this unique identifier with the traffic security analysis result and add it to the historical data set. Subsequently, when an HTTP session segment matching the unique identifier appears in the network traffic, the security intrusion detection device can quickly retrieve its previous security analysis results and take corresponding security measures accordingly.
[0160] In one implementation, the traffic security analysis result of the network traffic includes the traffic security analysis result of the concurrent analysis and the traffic security analysis result of the sequential analysis, wherein the traffic security analysis result is obtained by performing network traffic security detection based on multiple analysis modules of the analysis platform. When the analysis platform is used to perform network traffic security detection, it includes:
[0161] Step S11: obtaining the resource availability of each analysis module in the analysis platform, and determining the target communication node that needs to be currently subjected to network traffic security detection from the communication topology relationship network according to the focus of each communication node in the communication topology relationship network;
[0162] Step S12: determining a target analysis module from each analysis module according to resource availability of each analysis module and complexity of communication data packets contained in the target communication node;
[0163] Step S13: Use the target analysis module to perform network traffic security detection on the communication data packets contained in the target communication node. After obtaining the corresponding traffic security analysis results, reset the resource availability of the target analysis module to obtain the reset resource availability.
[0164] Step S11 is the first step of the analysis platform to perform network traffic security detection. Its core goal is to determine the target communication node for the current network traffic security detection from the communication topology network based on the resource availability of each analysis module in the analysis platform and the focus of each communication node in the communication topology network. The implementation of this step is of great significance for optimizing resource allocation and improving detection efficiency.
[0165] When executing step S11, the security intrusion detection device obtains resource availability information of each analysis module in the analysis platform. Resource availability usually includes key indicators such as the current number of tasks, computing usage ratio, and computing remaining ratio of the analysis module. These indicators reflect the current workload and computing power of the analysis module and are an important basis for determining the target communication node.
[0166] For example, suppose the analysis platform includes two analysis modules A and B. The current number of tasks of module A is 5, the computing utilization ratio is 80%, and the computing remaining ratio is 20%; while the current number of tasks of module B is 3, the computing utilization ratio is 60%, and the computing remaining ratio is 40%. Based on this information, the security intrusion detection device can preliminarily determine that module B currently has more computing resources available.
[0167] Next, the security intrusion detection device will further screen the target communication nodes according to the focus of each communication node in the communication topology network. Focus is an indicator to measure the importance and attention of a communication node in network traffic security detection. It is usually calculated based on factors such as the degree, centrality, and location of the communication node. The higher the focus of a communication node, the more important the communication data packets it contains are in network traffic security detection, so it should be detected first.
[0168] For example, in the communication topology network, suppose there are three communication nodes C, D and E, and their focus degrees are 0.9, 0.7 and 0.6 respectively. Based on this information, the security intrusion detection device can determine that node C is the most important communication node at present and should be detected first.
[0169] After comprehensively considering the resource availability of the analysis module and the focus of the communication nodes, the security intrusion detection device will use a heuristic algorithm or optimization strategy to determine the target communication node. This algorithm or strategy may be dynamically adjusted according to factors such as the resource load of the analysis module, the focus of the communication nodes, and the real-time changes in network traffic to ensure the efficient execution of the detection task.
[0170] For example, in the above example, if the computing resources of analysis module B are relatively abundant and communication node C has the highest focus, then the security intrusion detection device may choose communication node C as the target communication node and assign it to analysis module B for detection.
[0171] The purpose of step S12 is to select the most suitable target analysis module for security detection from multiple analysis modules in the analysis platform according to the resource availability of each analysis module and the complexity of the communication data packets contained in the target communication node. The implementation of this step is also crucial to improving the accuracy and efficiency of detection.
[0172] When executing step S12, the security intrusion detection device first evaluates the complexity of the communication data packet contained in the target communication node. The complexity of the communication data packet may be determined based on a variety of factors, such as the size of the data packet, the protocol type, the number of fields contained, etc. The more complex the data packet, the more difficult it is to process and analyze, so more computing resources are allocated.
[0173] For example, suppose the target communication node contains a large TCP data packet, which contains multiple complex protocol fields and a large amount of application layer data. Based on this information, the security intrusion detection device can determine that the complexity of the data packet is high and allocate more computing resources for analysis.
[0174] Next, the security intrusion detection device will select the appropriate analysis module based on the resource availability of each analysis module. Resource availability not only includes the usage of computing resources, but may also involve the usage of other resources such as memory, storage, and network bandwidth. When selecting an analysis module, the security intrusion detection device comprehensively considers factors such as the data packet complexity of the target communication node, the resource load of the analysis module, and the load balancing between modules.
[0175] For example, in the above example, if the computing resources of analysis module B are relatively abundant and it is good at processing TCP data packets with higher complexity, then the security intrusion detection device may choose analysis module B as the target analysis module to perform security detection on the communication data packets contained in the target communication node.
[0176] In addition, security intrusion detection devices can also consider introducing machine learning models to assist in selecting target analysis modules. For example, a classification model can be used to predict the effectiveness and efficiency of different analysis modules in processing specific types of data packets, thereby selecting the optimal analysis module. This model can continuously optimize its prediction accuracy through training and learning from historical data.
[0177] Step S13 is performed after the target communication node and the target analysis module are determined. Its core goal is to use the target analysis module to perform network traffic security detection on the communication data packets contained in the target communication node, and reset the resource availability of the target analysis module after the detection is completed. The implementation of this step is the key to completing the entire detection process.
[0178] When executing step S13, the security intrusion detection device will send the communication data packet contained in the target communication node to the target analysis module for processing. The analysis module will perform in-depth analysis on the data packet based on its built-in security detection algorithm or model to identify possible security threats or abnormal behaviors. The analysis process may involve a variety of technical means such as parsing the content of the data packet, pattern matching, statistical analysis, machine learning prediction, etc.
[0179] For example, assuming that the target analysis module is an intrusion detection system based on deep learning, it will perform in-depth analysis on the communication data packets contained in the target communication node to identify possible security threats such as attack fields and abnormal traffic patterns.
[0180] After the detection is completed, the security intrusion detection device will obtain the traffic security analysis results generated by the target analysis module and store them for subsequent use. At the same time, in order to maintain the continuous and efficient operation of the analysis platform, the security intrusion detection device also resets the resource availability of the target analysis module. The reset process usually includes updating the key indicators such as the current number of tasks and the calculation usage ratio of the analysis module to reflect its latest resource usage. For example, in the above example, if the target analysis module successfully completes the security detection task of the target communication node, the security intrusion detection device will reduce its current number of tasks by one and update its calculation usage ratio and calculation remaining ratio. In this way, when the analysis platform receives a new detection task, it can reallocate tasks and analysis modules according to the latest resource availability information.
[0181] In one implementation, step S12, according to the resource availability of each analysis module and the complexity of the communication data packet contained in the target communication node, determines a target analysis module from each analysis module, including:
[0182] Step S12A: Obtain the complexity of the communication data packet contained in the target communication node;
[0183] Step S12B: according to the resource availability of each analysis module, determining from the analysis platform a reference analysis module whose resource availability can perform network traffic security analysis on communication data packets of corresponding complexity;
[0184] Step S12C: Obtain the task allocation confidence of the reference analysis module, and determine the target analysis module according to the task allocation confidence of each analysis module.
[0185] The purpose of step S12A is to obtain the complexity of the communication data packet contained in the target communication node. The complexity of the communication data packet is an important indicator to measure the difficulty of data packet processing and the required computing resources. It is usually determined based on factors such as the size of the data packet, the protocol type, the number of fields contained, the complexity of the data content, etc. The more complex the data packet is, the more computing resources are usually required for its processing and analysis.
[0186] When executing step S12A, the security intrusion detection device first parses the communication data packet contained in the target communication node and extracts key information of the data packet, such as protocol type, data packet length, included fields, etc. Then, the security intrusion detection device calculates the complexity of the data packet based on this information and predefined complexity evaluation rules or models. The complexity evaluation rules or models may be constructed based on historical data analysis, expert experience, or machine learning algorithms, and can accurately reflect the processing difficulty and resource requirements of the data packet.
[0187] For example, assume that the target communication node contains an HTTP response data packet, which contains a large number of header fields and response body data. When executing step S12A, the security intrusion detection device will parse the HTTP response data packet and extract information such as the protocol type (HTTP), data packet length, number of header fields, and response body size. Then, the security intrusion detection device will calculate the complexity of the HTTP response data packet based on this information and a predefined complexity evaluation model. If the data packet length is long, the number of header fields is large, and the response body data is complex, the calculated complexity value may be high.
[0188] The purpose of step S12B is to determine, from the analysis platform, reference analysis modules that can perform network traffic security analysis on communication packets of corresponding complexity based on the resource availability of each analysis module. Reference analysis modules refer to those analysis modules that have high current resource availability and can process communication packets contained in the target communication node.
[0189] When executing step S12B, the security intrusion detection device will first obtain the resource availability information of each analysis module in the analysis platform, including the current number of tasks, the calculation usage ratio, the calculation remaining ratio, etc. Then, the security intrusion detection device will screen out the reference analysis module that can process the data packet based on this information and the complexity of the communication data packet contained in the target communication node. The screening process may involve comprehensive consideration of factors such as the computing power, processing speed, and protocol types that the analysis module is good at processing.
[0190] For example, assume that the analysis platform includes three analysis modules A, B, and C, and their resource availability information is as follows: the current number of tasks of module A is 2, the calculation usage ratio is 60%, and the calculation remaining ratio is 40%; the current number of tasks of module B is 1, the calculation usage ratio is 30%, and the calculation remaining ratio is 70%; the current number of tasks of module C is 3, the calculation usage ratio is 80%, and the calculation remaining ratio is 20%. At the same time, assume that the communication data packet contained in the target communication node is a TCP data packet with high complexity. When executing step S12B, the security intrusion detection device will screen out the reference analysis module that can process this TCP data packet based on this information and the processing capabilities of each analysis module. If module B is good at processing TCP data packets and has a high current resource availability (the calculation remaining ratio is 70%), then it may be selected as a reference analysis module.
[0191] The purpose of step S12C is to determine the final target analysis module based on the task allocation confidence of each analysis module after determining the reference analysis module. Task allocation confidence is an indicator to measure the ability and willingness of an analysis module to process tasks. It is usually determined based on factors such as the resource availability, processing speed, and historical task completion of the analysis module. The analysis module with higher task allocation confidence is more likely to be selected as the target analysis module to process the communication data packets contained in the target communication node.
[0192] When executing step S12C, the security intrusion detection device will first obtain the task allocation confidence of the reference analysis module. The task allocation confidence can be calculated in a variety of ways, one of which is based on the scheduling selection probability. The scheduling selection probability refers to the probability that the analysis module is selected by the scheduling system to process the task, which is usually proportional to the resource availability of the analysis module and inversely proportional to the total remaining load of the analysis platform. The greater the scheduling selection probability, the higher the current resource availability of the analysis module and the stronger its ability to process tasks, so the higher its task allocation confidence.
[0193] For example, suppose that two reference analysis modules B and C are determined in step S12B, and their resource availability information is respectively: the current number of tasks of module B is 1, the calculation usage ratio is 30%, and the calculation remaining ratio is 70%; the current number of tasks of module C is 3, the calculation usage ratio is 80%, and the calculation remaining ratio is 20%. When executing step S12C, the security intrusion detection device will calculate the scheduling selection probability of modules B and C based on this information and the predefined scheduling selection probability calculation formula. Assuming that the calculation formula of the scheduling selection probability is "resource availability of the analysis module / total remaining load of the analysis platform", the scheduling selection probability of module B is "70% / (70%+20%) = 7 / 9", and the scheduling selection probability of module C is "20% / (70%+20%) = 2 / 9". Since the scheduling selection probability of module B is higher, it means that its current resource availability is higher and its ability to process tasks is stronger, so it is more likely to be selected as the target analysis module to process the communication data packets contained in the target communication node.
[0194] If there is only one reference analysis module, then this module can be directly selected as the target analysis module without further comparison and selection. If there are multiple reference analysis modules, the security intrusion detection device will determine the final target analysis module based on the task allocation confidence of each analysis module (such as the scheduling selection probability). The analysis module with the highest task allocation confidence will be selected as the target analysis module to process the communication data packets contained in the target communication node.
[0195] In one implementation, in step S11, obtaining resource availability of each analysis module in the analysis platform includes:
[0196] Step S111: for the selected analysis module in the analysis platform, the current task quantity and the computing usage ratio of the selected analysis module are obtained; the current task quantity represents the number of communication nodes for which the selected analysis module is used to perform network traffic security detection, and the computing usage ratio represents the usage information of the computing resources of the selected analysis module when performing network traffic security detection on the corresponding number of communication nodes;
[0197] Step S112: determining the average computing usage ratio required by the selected analysis module when performing network traffic security analysis on a communication node according to the computing usage ratio and the current number of tasks;
[0198] Step S113: Obtain the remaining calculation ratio of the selected analysis module, and obtain the resource availability of the selected analysis module by combining the average calculation usage ratio and the remaining calculation ratio.
[0199] The purpose of step S111 is to obtain the current number of tasks and the calculation usage ratio for each selected analysis module in the analysis platform. The current number of tasks indicates the number of communication nodes that are currently using the analysis module to perform network traffic security detection, which reflects the current workload of the analysis module. The calculation usage ratio indicates the usage of the computing resources of the analysis module when performing network traffic security detection on the corresponding number of communication nodes, which is usually expressed in the form of a percentage. When executing step S111, the security intrusion detection device will traverse all analysis modules in the analysis platform. For each selected analysis module, it will query the current task queue of the module to obtain the number of communication nodes being processed, that is, the current number of tasks. At the same time, the security intrusion detection device will also monitor the computing resource usage of the module, such as CPU usage, memory occupancy, etc., and calculate the computing usage ratio based on this information. The calculation of the computing usage ratio may involve weighted summation or averaging of the usage of multiple computing resources to obtain an indicator that comprehensively reflects the computing resource utilization of the analysis module.
[0200] For example, suppose the analysis platform contains two analysis modules A and B. Currently, module A is processing the network traffic security detection task of 3 communication nodes, with a CPU usage rate of 70% and a memory usage rate of 60%; while module B is processing the task of 2 communication nodes, with a CPU usage rate of 50% and a memory usage rate of 40%. When executing step S111, the security intrusion detection device will obtain that the current number of tasks of module A is 3, and the calculated usage ratio is (70%+60%) / 2 = 65%; the current number of tasks of module B is 2, and the calculated usage ratio is (50%+40%) / 2 = 45%.
[0201] The purpose of step S112 is to determine the average computing usage ratio required by the module when performing network traffic security analysis on a communication node based on the current number of tasks and computing usage ratio of the analysis module. The average computing usage ratio is an indicator that measures the computing resources required by the analysis module to process a single communication node task, which helps the security intrusion detection device to more accurately evaluate the resource availability of the analysis module.
[0202] When executing step S112, the security intrusion detection device will first calculate the total computing usage ratio of the analysis module, that is, the product of the current number of tasks and the computing usage ratio will be accumulated. Then, the security intrusion detection device will divide the total computing usage ratio by the current number of tasks to obtain the average computing usage ratio. This ratio reflects the average amount of computing resources used by the analysis module when processing a single communication node task. For example, continue to take modules A and B as an example. Assume that the current number of tasks of module A is 3, and the computing usage ratio is 65%; the current number of tasks of module B is 2, and the computing usage ratio is 45%. When executing step S112, the security intrusion detection device will first calculate the total computing usage ratio of modules A and B, which are 3*65% =195% and 2*45% = 90% respectively. Then, the security intrusion detection device will divide these total computing usage ratios by their respective current number of tasks to obtain the average computing usage ratio. For module A, the average computing usage ratio is 195% / 3 = 65%; for module B, the average computing usage ratio is 90% / 2 = 45%.
[0203] The purpose of step S113 is to obtain the calculation surplus ratio of the analysis module, and to obtain the resource availability of the module by combining the average calculation usage ratio and the calculation surplus ratio. The calculation surplus ratio refers to the proportion of the calculation resources currently available to the analysis module, which reflects the potential of the analysis module in processing new tasks. Resource availability is a comprehensive indicator, which is calculated based on the average calculation usage ratio and the calculation surplus ratio of the analysis module, and is used to measure the ability of the analysis module to process new tasks in the current state.
[0204] When executing step S113, the security intrusion detection device first obtains the calculation remaining ratio of the analysis module. The calculation of the calculation remaining ratio may involve operations such as normalizing the difference between the total computing resources of the analysis module and the currently used computing resources. Then, the security intrusion detection device combines the average computing usage ratio with the calculation remaining ratio to obtain an indicator that comprehensively reflects the resource availability of the analysis module. This indicator may be a ratio, percentage or other form of quantitative value, which helps the security intrusion detection device to more intuitively understand the resource status of the analysis module.
[0205] For example, let's continue to take modules A and B as an example. Assume that the remaining calculation ratios of modules A and B are 35% and 55% respectively. When executing step S113, the security intrusion detection device will calculate the resource availability by combining the average calculation usage ratio (65% for module A and 45% for module B) and the remaining calculation ratio (35% for module A and 55% for module B) calculated in step S112. One possible calculation method is to use the ratio method, that is, to divide the remaining calculation ratio by (1+average calculation usage ratio). For module A, the resource availability is 35% / (1+65%)≈21%; for module B, the resource availability is 55% / (1+45%)≈38%. These resource availability indicators reflect the ability of modules A and B to process new tasks in the current state, among which module B has a higher resource availability and is therefore more likely to be selected as the target analysis module to process new network traffic security detection tasks.
[0206] In one implementation, in step S11, according to the focus of each communication node in the communication topology relationship network, a target communication node that needs to be currently subjected to network traffic security detection is determined from the communication topology relationship network, including:
[0207] Step S114: determining at least one first communication node at the analysis end tag and a source communication node of any first communication node from the communication topology relationship network;
[0208] Step S115: Obtain the focus of each source communication node in the communication topology relationship network, and simultaneously obtain a set of task nodes that require network traffic security detection;
[0209] Step S116: determining the focus degree of each second communication node in the task node set, and adding each source communication node to the task node set one by one according to the focus degree of each source communication node and the focus degree of each second communication node, so that the communication nodes in the task node set of each source communication node are arranged in descending order according to the corresponding focus degree;
[0210] Step S117: Determine the communication node with the highest focus from the task node set as the target communication node for current network traffic security detection.
[0211] The purpose of step S114 is to identify those communication nodes (i.e., first communication nodes) that have completed security analysis and marked as "analysis completed" from the communication topology relationship network, and further determine the source communication nodes of these nodes. The source communication node refers to the communication node that sends data to the first communication node in the network traffic. They are usually located upstream of the first communication node and are the starting point or intermediate node of the network traffic transmission.
[0212] When executing step S114, the security intrusion detection device traverses all communication nodes in the communication topology relationship network and checks their analysis tags. The analysis tag is a mechanism for marking the analysis status of the communication node, which usually includes states such as "analysis completed", "unanalyzed" and "analyzing". The security intrusion detection device will filter out all communication nodes marked as "analysis completed", which are the first communication nodes. Then, for each first communication node, the security intrusion detection device will trace its upstream path in the communication topology relationship network to find the source communication node that sends data to it.
[0213] For example, assume that the communication topology network includes three communication nodes A, B, and C, where nodes A and B have completed security analysis and are marked as "analysis completed", while node C has not yet been analyzed. The source communication node of node A may be node D, and the source communication node of node B may be node E or F. When executing step S114, the security intrusion detection device will identify nodes A and B as the first communication nodes, and further determine nodes D, E, and F as their source communication nodes.
[0214] The purpose of step S115 is to obtain the focus of each source communication node determined in step S114 in the communication topology relationship network, and initialize a task node set for storing the communication nodes to be analyzed later. Focus is an indicator to measure the importance and attention of communication nodes in network traffic security detection, which is usually calculated based on factors such as the degree, centrality, and location of the communication node.
[0215] When executing step S115, the security intrusion detection device will query the focus information of each source communication node in the communication topology relationship network. This information may be calculated in advance by an algorithm or model and stored in the data structure of the communication topology relationship network. At the same time, the security intrusion detection device will initialize an empty task node set for subsequent storage of communication nodes to be analyzed. The task node set is a dynamic data structure that is continuously updated as the detection process proceeds.
[0216] The purpose of step S116 is to sort the task node set according to the focus of the source communication node and the second communication node. The second communication node refers to all other communication nodes to be analyzed except the source communication node. The purpose of sorting is to ensure that the communication nodes in the task node set are arranged in descending order according to their focus, so as to give priority to those communication nodes that are more important and critical in network traffic security detection.
[0217] When executing step S116, the security intrusion detection device will first determine all the second communication nodes in the task node set and calculate their focus. Then, the security intrusion detection device will add each source communication node to the task node set one by one, and sort the task node set according to the focus of the source communication node and the second communication node. The sorting algorithm may use a simple bubble sort, selection sort or a more efficient sorting algorithm, such as quick sort, merge sort, etc. The basis for sorting is the focus value of the communication node, and the higher the focus, the closer the communication node is to the front of the task node set.
[0218] For example, assume that the task node set initially includes three second communication nodes X, Y, and Z, whose focusing degrees are 0.8, 0.6, and 0.7, respectively. At the same time, assume that the focusing degrees of the source communication nodes D, E, and F determined in step S114 are 0.9, 0.5, and 0.4, respectively. When executing step S116, the security intrusion detection device will add the source communication nodes D, E, and F to the task node set one by one, and sort the task node set according to their focusing degrees. The sorted task node set may become [D, X, Z, Y, E, F], where D has the highest focusing degree (0.9), X is second (0.8), and so on.
[0219] The purpose of step S117 is to select the communication node with the highest focus from the sorted task node set as the target communication node for network traffic security detection. The selection of the target communication node is based on its importance and attention in network traffic security detection, ensuring that detection resources can be allocated to the most valuable communication nodes first.
[0220] When executing step S117, the security intrusion detection device will traverse the sorted task node set and find the communication node with the highest focus. This node is the target communication node that currently needs to perform network traffic security detection. Once the target communication node is determined, the security intrusion detection device can pass it to the subsequent analysis module for processing and start the corresponding security detection process.
[0221] For example, let's continue to take the sorted task node set [D, X, Z, Y, E, F] in step S116 as an example. When executing step S117, the security intrusion detection device will traverse this task node set and find that node D has the highest focus (0.9). Therefore, node D will be selected as the target communication node for the current network traffic security detection. Subsequently, the security intrusion detection device will pass the relevant information of node D (such as communication data packets, source addresses, destination addresses, etc.) to the subsequent analysis module and start the security detection process.
[0222] The embodiment of the present invention provides a security intrusion detection device, such as Figure 2As shown, the security intrusion detection device 100 includes: a processor 101 and a memory 103. The processor 101 and the memory 103 are connected, such as through a bus 102. Optionally, the security intrusion detection device 100 may also include a transceiver 104. It should be noted that in actual applications, the transceiver 104 is not limited to one, and the structure of the security intrusion detection device 100 does not constitute a limitation on the embodiments of the present invention.
[0223] An embodiment of the present invention provides a security intrusion detection device. The security intrusion detection device in the embodiment of the present invention includes: one or more processors; a memory; one or more computer programs, wherein the one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more programs are executed by the processor, the method of the above embodiment is implemented.
Claims
1. A network security intrusion detection method, characterized in that: Applied to a security intrusion detection device, the method comprises: Obtaining network traffic to be analyzed; the network traffic covers several session segments, and a communication data packet in a session segment describes a complete network interaction process; A communication topology relationship network is established based on the communication data packets of each session segment in the network traffic and the association between different communication data packets; the communication topology relationship network includes a number of communication nodes and traffic pointers, one communication node includes a communication data packet of a session segment, and one traffic pointer indicates the association between the communication data packets contained in the connected communication nodes; When performing network traffic security detection on the network traffic according to the communication topology relationship network, determining a reference communication node supporting concurrent analysis from the communication topology relationship network; Concurrently analyzing the communication data packets contained in the reference communication node to obtain a traffic security analysis result of the network traffic; Wherein, determining a reference communication node supporting concurrent analysis from the communication topology relationship network includes: Determining the focus of each communication node based on the communication topology relationship network, wherein the focus represents the degree of attention paid to the communication data packets contained in the corresponding communication node when performing network traffic security detection, and at the same time, the communication data packets contained in the communication nodes with the same focus have the same degree of attention paid to the network traffic security detection; Communication nodes corresponding to the same focus degree are determined as reference communication nodes supporting concurrent analysis.
2. The method according to claim 1, characterized in that The communication topology relationship network established is an updated communication topology relationship network, and the traffic pointer contained in the updated communication topology relationship network has directionality; the communication topology relationship network is established according to the communication data packets of each session fragment in the network traffic, and the association between different communication data packets, including: An initial communication topology relationship network is established based on the communication data packets of each session segment in the network traffic and the association between different communication data packets, wherein the communication nodes included in the initial communication topology relationship network contain communication data packets, and the direction of pointers between the communication nodes is that the communication data packet that initiates the communication points to the communication data packet that receives the communication; According to the communication nodes and pointers included in the initial communication topology relationship network, the initial communication topology relationship network is updated to obtain an updated communication topology relationship network.
3. The method according to claim 2, characterized in that The updating of the initial communication topology relationship network according to the communication nodes and pointers included in the initial communication topology relationship network to obtain an updated communication topology relationship network includes: Retrieving the traffic security analysis results of the communication data packets contained in the selected communication nodes in the initial communication topology relationship network from the historical data set; If the traffic security analysis result of the communication data packet contained in the selected communication node is retrieved, then according to the distribution position of the selected communication node in the initial communication topology relationship network and the pointing of the pointer included in the initial communication topology relationship network, the associated communication node having a communication association with the selected communication node is obtained; Removing redundancy from the selected communication nodes and the associated communication nodes to obtain an updated communication topology relationship network; or; The updating of the initial communication topology relationship network according to the communication nodes and pointers included in the initial communication topology relationship network to obtain an updated communication topology relationship network includes: Based on the communication data packets contained in the communication nodes included in the initial communication topology relationship network, a plurality of different communication nodes having the same communication data packet content are determined in the initial communication topology relationship network; Removing the multiple different communication nodes and retaining one of the multiple different communication nodes; Based on the distribution of the multiple different communication nodes in the initial communication topology relationship network and the pointing of the pointers included in the initial communication topology relationship network, the communication nodes that have communication associations with the removed communication nodes and the remaining communication nodes are connected to obtain an updated communication topology relationship network.
4. The method according to claim 3, characterized in that The historical data set includes a unique identifier of a communication data packet of which the traffic security analysis result is known; the process of retrieving the traffic security analysis result of the communication data packet included in the communication node selected in the initial communication topology relationship network from the historical data set includes: Obtaining a reference unique identifier of a communication data packet contained in a selected communication node in the initial communication topology relationship network; In the historical data set, a unique identifier that is consistent with the reference unique identifier is retrieved. If a unique identifier that is consistent with the reference unique identifier is retrieved in the historical data set, it is determined that the traffic security analysis results of the communication data packets contained in the selected communication node are retrieved; if a unique identifier that is consistent with the reference unique identifier is not retrieved in the historical data set, it is determined that the traffic security analysis results of the communication data packets contained in the selected communication node are not retrieved.
5. The method according to claim 4, characterized in that The determining the focusing degree of each communication node based on the communication topology relationship network includes: Obtaining the distribution level of the selected communication node in the communication topology relationship network; Determining the focus degree of the selected communication node according to the distribution level of the selected communication node; If the communication topology relationship network includes a communication node that is consistent with the distribution level of the selected communication node and is different from the source communication node of the selected communication node, the method further includes: Acquire the number of first communication nodes corresponding to the same source communication node as the selected communication node, and the number of second communication nodes corresponding to different source communication nodes as the selected communication node and having the same distribution level, wherein both the first communication node and the second communication node are communication nodes corresponding to communication data packets that have not been analyzed; If the number of the first communication nodes is greater than the number of the second communication nodes, the focus of the selected communication node is greater than the focus of the communication nodes with different sources corresponding to the selected communication node and with the same distribution level; If the number of the first communication nodes is less than the number of the second communication nodes, the focus of the selected communication node is less than the focus of the communication nodes with different sources corresponding to the selected communication node and with the same distribution level; If the number of the first communication nodes is equal to the number of the second communication nodes, the focus degree of the selected communication node is equal to the focus degree of the communication nodes with different source nodes corresponding to the selected communication node and with the same distribution level.
6. The method according to claim 1, characterized in that The traffic security analysis result of the network traffic includes the traffic security analysis result of each session segment in the network traffic; the method further includes: According to the analysis status of network traffic security detection for each session segment in the network traffic, an analysis tag is assigned to the communication node corresponding to the corresponding session segment in the communication topology relationship network; the analysis tag includes an analysis end tag, an unanalyzed tag and an analysis ongoing tag; According to the analysis tags added to each communication node in the communication topology relationship network, a unique identifier of the session segment in the communication node assigned with the analysis end tag and a corresponding traffic security analysis result are obtained; The obtained unique identifier and the corresponding traffic security analysis result are associated and added to the historical data set.
7. The method according to claim 1, characterized in that The traffic security analysis result of the network traffic includes the traffic security analysis result of concurrent analysis and the traffic security analysis result of sequential analysis, wherein the traffic security analysis result is obtained by performing network traffic security detection based on multiple analysis modules of the analysis platform. When the network traffic security detection is performed using the analysis platform, it includes: Obtaining the resource availability of each analysis module in the analysis platform, and determining the target communication node that currently needs to be subjected to network traffic security detection from the communication topology relationship network according to the focus of each communication node in the communication topology relationship network; Determining a target analysis module from the analysis modules according to resource availability of each analysis module and complexity of communication data packets contained in the target communication node; The target analysis module is used to perform network traffic security detection on the communication data packets contained in the target communication node. After obtaining the corresponding traffic security analysis results, the resource availability of the target analysis module is reset to obtain the reset resource availability.
8. The method according to claim 7, characterized in that The step of determining a target analysis module from the analysis modules according to the resource availability of the analysis modules and the complexity of the communication data packets contained in the target communication node comprises: Obtaining the complexity of the communication data packet contained in the target communication node; According to the resource availability of each analysis module, a reference analysis module whose resource availability is capable of performing network traffic security analysis on communication data packets of corresponding complexity is determined from the analysis platform; Obtaining the task allocation confidence of the reference analysis module, and determining the target analysis module according to the task allocation confidence of each analysis module; The obtaining of resource availability of each analysis module in the analysis platform includes: For the selected analysis module in the analysis platform, the current task quantity and computing usage ratio of the selected analysis module are obtained; the current task quantity represents the number of communication nodes that use the selected analysis module to perform network traffic security detection, and the computing usage ratio represents the usage information of computing resources of the selected analysis module when performing network traffic security detection on the corresponding number of communication nodes; Determine, according to the computing usage ratio and the current number of tasks, an average computing usage ratio required by the selected analysis module when performing network traffic security analysis on a communication node; Obtaining the remaining calculation ratio of the selected analysis module, and combining the average calculation usage ratio and the remaining calculation ratio to obtain the resource availability of the selected analysis module; The step of determining a target communication node that needs to be currently subjected to network traffic security detection from the communication topology relationship network according to the focus degree of each communication node in the communication topology relationship network includes: Determine at least one first communication node at an analysis end tag and a source communication node of any first communication node from the communication topology relationship network; Obtaining the focus of each source communication node in the communication topology relationship network, and obtaining a set of task nodes that need to perform network traffic security detection; Determine the focusing degree of each second communication node in the task node set, and add each source communication node to the task node set one by one according to the focusing degree of each source communication node and the focusing degree of each second communication node, so that the communication nodes in the task node set of each source communication node are arranged in descending order according to the corresponding focusing degree; A communication node with the greatest focus is determined from the task nodes as a target communication node for current network traffic security detection.
9. A security intrusion detection device, characterized in that: include: one or more processors; Memory; one or more computer programs; The one or more computer programs are stored in the memory and configured to be executed by the one or more processors, and when the one or more computer programs are executed by the processors, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Fiber channel flow analyzing and recording method under avionic environment and fiber channel flow analyzing and recording device thereof
CN107276834A
5G slice network anomaly detection method based on virtual network flow analysis
CN114401516A