A P2P botnet detection method based on communication topology and network traffic
By combining the intelligent agent framework of graph neural networks and reinforcement learning, and combining traffic and topology detection methods, we solve the problems of high resource consumption and insufficient precision in large-scale P2P botnet detection, and achieve efficient and accurate botnet detection.
Patent Information
- Application Number
- CN202310191096.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-22
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2043-02-22
AI Technical Summary
Existing botnet detection technologies are difficult to achieve efficient and accurate detection in large-scale networks. Traffic-based methods consume a lot of resources, while topology-based methods lack detection accuracy and cannot effectively mitigate the threat of P2P botnets.
Combining traffic-based and topology-based detection methods, an intelligent agent framework based on graph neural networks and reinforcement learning is adopted. Node confidence is predicted through graph convolutional networks, and reinforcement learning is used to optimize detection strategies. False positives and false negatives are iteratively corrected to improve detection accuracy.
It achieves efficient and accurate detection in large-scale P2P botnets, combines the advantages of traffic and topology detection, reduces computing resource consumption, and improves detection efficiency and accuracy.
Smart Images

Figure CN116232721B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer network security technology, and belongs to the field of intrusion detection system (IDS), and in particular to a P2P botnet detection method based on communication topology and network traffic, and adopts graph neural network and reinforcement learning methods. Background Art
[0002] Botnets remain a serious threat to the internet. Adversaries can control hosts to conduct malicious activities such as distributed denial of service (DDoS) attacks, malware, spam, and market manipulation. Over the past few decades, botnet operators have maintained a competitive advantage through technological advancements, particularly the use of peer-to-peer (P2P) infrastructure to spread botnets. Research and industry efforts have devoted significant effort to detecting botnets, but these efforts have proven ineffective in mitigating the harm they cause. Therefore, research on botnet detection technology is of great practical significance.
[0003] Based on the detection approach, existing botnet detection techniques can be categorized as flow-based and topology-based. In flow-based approaches, security practitioners aim to detect suspicious communication behaviors indicative of malicious activity, such as command and control (C&C) communications and Domain Name System (DNS) queries. While flow-based approaches can provide fine-grained detection results, with the advancement of network infrastructure and the widespread adoption of 100Gbps and even terabit bandwidth, botnet traffic can be drowned out by a significant amount of background traffic (large-scale overlay topology). For example, in a wide area network (WAN) with 100,000 nodes, if any communication occurs between two hosts, a significant amount of traffic will be detected. Therefore, while flow-based approaches can provide precise behavioral detection, their significant overhead limits their practical application in the real world. Leveraging communication topology, on the other hand, holds great promise for large-scale botnet detection. This is because botnets typically exhibit specific communication patterns, whether distributed P2P or centralized C&C. For example, a series of popular structured botnets exhibit a rapid mixing process, representing the convergence trajectory of a random walk to a stationary distribution. Furthermore, the detection difficulty of various protocols varies to varying degrees because their graphs have different average path lengths. Generally speaking, leveraging topology is promising and effective for large-scale botnet detection because communication graph data is much smaller than traffic data. Topology-based methods aim to discover communication patterns from overlay topologies. Such schemes do not require analyzing the traffic of each host node, reducing the consumption of computational resources. However, using only communication topology to identify botnets is relatively coarse-grained.
[0004] Combining traffic-based and topology-based solutions to detect botnets with specific structures can avoid the high resource consumption brought about by the use of neural networks in large-scale traffic analysis, while also maintaining fine-grained detection of botnets. Therefore, weighing the pros and cons of the two methods to maximize the efficiency of botnet detection is the focus of research. The present invention jointly analyzes and deploys reinforcement learning-based intelligent agents to plan intelligent detection strategies. In this process, the intelligent agent will determine the nodes that need to be detected first (usually false positives or false negatives) based on the topology and current environmental information, thereby significantly improving accuracy, but only analyzing the traffic from a small part of the node. Summary of the Invention
[0005] This invention addresses the shortcomings of existing technologies by proposing a P2P botnet detection method based on communication topology and network traffic. This method combines traffic-based and topology-based detection schemes. Suspicious botnet nodes are located based on the communication topology. tCommander uses tScouter's output as the initial context to determine the next node to detect. tPatroller then analyzes the corresponding network traffic. tPatroller's detection results update the context, and tCommander continues to make decisions based on the new state. This iterative process corrects for false positives and false negatives, improving accuracy.
[0006] According to a first aspect of an embodiment of the present application, a method for detecting a P2P botnet based on communication topology and network traffic is provided, comprising:
[0007] (1) obtaining a P2P botnet, and synthesizing botnet communication traffic data by superimposing the network topology of the P2P botnet, traffic based on the P2P protocol, and real-world botnet topology onto background traffic topology;
[0008] (2) topologically representing the botnet communication traffic data to obtain a topological graph;
[0009] (3) using a graph convolutional network to predict the confidence that each node in the topological graph is a robot node, where the robot node is a node in the botnet;
[0010] (4) using the confidence of each node in the topological graph to initialize the environment and state and construct a neural network as an intelligent agent, the intelligent agent determines the next node to be detected by class distribution sampling;
[0011] (5) detecting the traffic of the detected node, updating the environment and state according to the detection results, and calculating the reward value to update the model parameters of the intelligent agent;
[0012] (6) Repeat steps (4) and (5) above until all nodes are detected or the rewards after a certain number of consecutive searches are close to 0;
[0013] (7) repeating steps (4) to (6) a predetermined number of times to complete the training of the agent;
[0014] (8) Replace the P2P botnet with the P2P network to be detected, repeat steps (1) to (3), use the node confidence as input, and use the trained neural network to obtain the optimal detection path.
[0015] Furthermore, the environment includes a node list, an edge list, and a label set of the node, and the state includes the confidence of the node.
[0016] Furthermore, the state adopts a state table structure, which includes the following information for each node: whether the node is a robot node, the confidence of the node, the degree of the node, the number of the node's viewed neighbors that are benign, the number of the node's viewed neighbors that are robot nodes, the number of the node's unviewed neighbors, and additional node features for constructing heterogeneous scenarios.
[0017] Furthermore, the neural network includes four linear layers, three ReLU activation layers and one softmax layer.
[0018] Furthermore, the reward value includes a reward value for node correction and a reward value for boundary discovery.
[0019] Furthermore, the process of detecting the traffic of a node includes:
[0020] Collect all communication traffic related to the detected node, including senders and receivers;
[0021] The communication traffic is divided according to sessions, and features thereof are extracted to form an identified sample set, and the sample set is detected to obtain a detection result thereof.
[0022] Furthermore, the environment and status are updated according to the test results, specifically:
[0023] If the detection result is different from the category corresponding to the confidence of the detected node, the detected node is determined to be an erroneous node and its status is corrected.
[0024] According to a second aspect of an embodiment of the present application, a P2P botnet detection device based on communication topology and network traffic is provided, comprising:
[0025] a synthesis module for obtaining a P2P botnet and synthesizing botnet communication traffic data by superimposing a network topology of the P2P botnet, traffic based on a P2P protocol, and a real-world botnet topology onto a background traffic topology;
[0026] A topology representation module, configured to perform topological representation on the botnet communication traffic data to obtain a topology map;
[0027] a prediction module, configured to use a graph convolutional network to predict the confidence that each node in the topological graph is a robot node, wherein the robot node is a node in the botnet;
[0028] An initialization module is used to initialize the environment and state using the confidence of each node in the topology map and to construct a neural network as an intelligent agent, and the intelligent agent determines the next node to be detected by class distribution sampling;
[0029] An updating module, configured to detect the flow of the detected node, update the environment and state according to the detection result, and calculate a reward value to update the model parameters of the intelligent agent;
[0030] The repetition module is used to repeat the above initialization module and update module process until all nodes are detected or the rewards after a certain number of consecutive searches are close to 0;
[0031] A training module, configured to repeat the above steps from the initialization module to the repetition module for a predetermined number of times to complete the training of the agent;
[0032] The detection module is used to replace the P2P botnet with the P2P network to be detected, repeat the process from the synthesis module to the prediction module, use the node confidence as input, and use the trained neural network to obtain the optimal detection path.
[0033] According to a third aspect of the embodiments of the present application, there is provided an electronic device, including:
[0034] one or more processors;
[0035] a memory for storing one or more programs;
[0036] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the first aspect.
[0037] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in the first aspect are implemented.
[0038] The technical solutions provided by the embodiments of the present application may have the following beneficial effects:
[0039] As can be seen from the above examples, this application proposes a large-scale P2P botnet detection framework based on communication topology and network traffic. This is a new framework based on an intelligent planning process that combines traffic-based and topology-based botnet detection schemes to achieve a trade-off between accuracy and efficiency. The present invention utilizes a graph neural network model, using a topology-based approach to locate suspicious bot nodes, while employing reinforcement learning to optimize the steps of the traffic-based approach. This allows the present invention to combine the two approaches for botnet detection while retaining the detection advantages of both approaches, thereby achieving efficient detection of large-scale P2P botnets.
[0040] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0042] Figure 1 The present invention is a flowchart showing a method for detecting a P2P botnet based on communication topology and network traffic according to an exemplary embodiment.
[0043] Figure 2 The figure is a flowchart of a method for reducing the dimension of a feature vector using LSTM according to an exemplary embodiment.
[0044] Figure 3 FIG. 4 is a diagram showing a framework structure for analyzing the substitutability between different types of DDoS attacks according to an exemplary embodiment.
[0045] Figure 4 The present invention is a block diagram showing a P2P botnet detection device based on communication topology and network traffic according to an exemplary embodiment.
[0046] Figure 5 is a schematic diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0047] Exemplary embodiments are described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different drawings represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this application.
[0048] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0049] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0050] Figure 1 FIG. 1 is a flow chart showing a method for detecting a P2P botnet based on communication topology and network traffic according to an exemplary embodiment. Figure 1 As shown, this method can be applied to the detection of any P2P botnet, but is more targeted at the detection of large-scale P2P botnets (large-scale in this application refers to the scale of the network. A network with thousands or tens of thousands of nodes is considered a large-scale network). The method may include the following steps:
[0051] (1) obtaining a P2P botnet, and synthesizing botnet communication traffic data by superimposing the network topology of the P2P botnet, traffic based on the P2P protocol, and real-world botnet topology onto background traffic topology;
[0052] (2) topologically representing the botnet communication traffic data to obtain a topological graph;
[0053] (3) using a graph convolutional network to predict the confidence that each node in the topological graph is a robot node, where the robot node is a node in the botnet;
[0054] (4) using the confidence of each node in the topology graph to initialize the environment and state table and construct a neural network as an intelligent agent, the intelligent agent determines the next node to be detected by class distribution sampling;
[0055] (5) detecting the traffic of the detected node, updating the environment and state according to the detection results, and calculating the reward value to update the model parameters of the intelligent agent;
[0056] (6) Repeat steps (4) and (5) above until all nodes are detected or the rewards of a certain number of consecutive searches are close to 0, and a detection path of the P2P botnet is obtained;
[0057] (7) repeating steps (4) to (6) a predetermined number of times to complete the training of the agent;
[0058] (8) Replace the P2P botnet with the P2P network to be detected, repeat steps (1) to (3), use the node confidence as input, and use the trained neural network to obtain the optimal detection path.
[0059] As can be seen from the above examples, this application proposes a large-scale P2P botnet detection framework based on communication topology and network traffic. This is a new framework based on an intelligent planning process that combines traffic-based and topology-based botnet detection schemes to achieve a trade-off between accuracy and efficiency. The present invention utilizes a graph neural network model, using a topology-based approach to locate suspicious bot nodes, while employing reinforcement learning to optimize the steps of the traffic-based approach. This allows the present invention to combine the two approaches for botnet detection while retaining the detection advantages of both approaches, thereby achieving efficient detection of large-scale P2P botnets.
[0060] In practice, this method is implemented through three tightly coupled components: tScouter, tCommander, and tPatroller. tScouter locates suspected botnet nodes based on the communication topology. tCommander uses tScouter's output as the initial context to determine the next node to detect. tPatroller then analyzes the corresponding network traffic. tPatroller's detection results update the context, and tCommander continues to make decisions based on the new state. This iterative process corrects for false positives and false negatives, improving accuracy.
[0061] First, tScouter is a graph convolutional network (GCN) model used to locate suspicious robot nodes based solely on communication topology. The output of tScouter is the malicious probability of each node, and these results and the connection relationships of the graph constitute the initial environment.
[0062] tCommander and tPatroller then work together to detect traffic flow at a few key nodes. tCommander is responsible for locating the next node to be detected, while tPatroller performs traffic analysis. tCommander is designed as a reinforcement learning (RL)-based agent, designed to prioritize nodes that are likely to be false positives or false negatives. Once tPatroller completes its traffic analysis, the results update the environment and state to determine the next agent's actions. Through this iteration, the framework can achieve efficient detection performance while only detecting traffic flow from a small number of nodes.
[0063] In a specific implementation of step (1), a P2P botnet is obtained, and the botnet communication traffic is synthesized by superimposing the network topology of the P2P botnet, traffic based on the P2P protocol, and the real-world botnet topology onto the background traffic topology;
[0064] Specifically, this step first generates a synthetic dataset for model training. In one embodiment, this dataset can include a large-scale background traffic topology, four P2P-based botnet topologies (such as Koorde, Kademlia, Chord, and LeetChord), two real-world botnets (P2P and centralized), and traffic based on seven common legal P2P protocols (such as Bamboo, Broose, Chord, Gia, Kademlia, Koorde, and Nice). The four P2P-based botnet topologies, two real-world botnets, and traffic based on seven common legal P2P protocols are superimposed on the large-scale background traffic to synthesize large-scale botnet communication traffic. The synthesis method uses node mapping and node embedding. Node mapping refers to mapping different IP addresses to nodes on a graph, and mapping communication traffic between IP addresses to edges on the graph. Node embedding involves cross-mixing traffic data from the seven common legal P2P protocols with data from the six bot traffic types. It should be noted that nodes in the P2P botnet are labeled to indicate whether they are bot nodes.
[0065] In the specific implementation of step (2), the communication traffic of the botnet is topologically represented to obtain a topological map;
[0066] Specifically, a communication topology can be defined as G = {V, A}, where V represents the host set, including n nodes {v1, ..., v n}, A is an n×n matrix, a symmetric matrix used to represent the node adjacent relationship. In the matrix, a ij =1 indicates node v i and v j The edge between them, otherwise it is a ij= 0, in order to record the number of edges of the node, the node degree matrix is defined as D = diag (d1, ..., d n )=0, where This step is to vectorize the traffic data in step (1) as the input of the model in step (3).
[0067] In the specific implementation of step (3), a graph convolutional network is used to predict the confidence that each node in the topological graph is a robot node, where the robot node is a node in the botnet;
[0068] Specifically, suspicious nodes are identified by using the graph convolutional network model of the tScouter component. The row feature vectors of the nodes in the graph convolutional network are updated by a learning matrix W. At the same time, the feature vectors of each node are averaged with the feature vectors of its direct neighbors, so that each node contains its neighbor information as a method to effectively explore the topological structure. The graph convolutional network can represent the vector size of each node as a fixed value (such as category) through multiple graph convolutional layers. In the last layer, the node vector is regarded as the attribute of the node, which is used to determine the confidence of whether the node is a robot node. For each convolutional layer, the feature vector expression of its node is:
[0069]
[0070] Among them, X is the feature matrix of the node, A is the symmetric matrix describing the adjacent relationship of the nodes, D is the node degree matrix, recording the number of edges of the node, and W is the learning matrix. is the row feature vector of the i-th node in the l-th layer, or D -1 A is the random procession normalization, a ij The i-th node is adjacent to the j-th node, d i is the degree of the i-th node. In addition, a nonlinear function σ is used to complete the update of the node feature vector of the current layer, which is expressed in the form of a matrix:
[0071]
[0072] Among them, U (l) is a learnable transformation matrix at layer l. Finally, the representation of the top node X (L) Cascading with the softmax function, the prediction vector of the i-th node can be expressed as:
[0073]
[0074] Among them, C i refers to the confidence of the classification, and ∑C i= 1, and the subsequent tCommander state table will be initialized based on the confidence level. At the same time, cross entropy is set as the loss function during training. When inputting the first layer, a uniform input of "1" is used, so that the feature initialization is independent of any node order.
[0075] In the specific implementation of step (4), the confidence of each node in the topological graph is used to initialize the environment and state table and construct a neural network as an intelligent agent, and the intelligent agent determines the next node to be detected by categorical distribution sampling;
[0076] Specifically, the environment refers to a series of information sets, and the state as input will be provided to the agent. The environment mainly consists of a node list, an edge list, and a label set. Among them, the node list records the host index, and the label set stores the corresponding ground truth ("benign" is "0", "botnet" is "1") confidence-label. In particular, we use the neighbor list to store edge information. The state serves as the input of the subsequent model, represents the current confidence of each node, and will be updated during the iteration. To this end, we developed a state table structure.
[0077] Figure 2This is the state table structure and update process diagram designed in the tCommander component of this invention. It consists of seven dimensional attributes: "Flag": indicates whether the node is a bot node ("0" for no, "1" for yes); "Probability": refers to the confidence level (initialized by tScouter and updated with detection results); "Degree": indicates the degree of the node (no direction); "D_benign": indicates the number of checked benign neighbors of this node; "D_botnet": indicates the number of checked botnet neighbors of this node; "D_none": indicates the number of unchecked neighbors. "Degree" = "D_benign" + "D_botnet" + "D_none". The last dimension stores some additional node features used to construct heterogeneous scenarios, such as the node's traffic volume. The state table update process is as follows: first, the environment is initialized, that is, the state table is initialized according to the output of the tScouter component. The tCommnader component will select the next node to be detected based on the initialized environment, such as detecting Node1. Then the tPatroller component will evaluate the actions performed by the tCommnader component based on the actual results of the sample, such as the "Flag" label of Node1 changes from "0" to "1", and the "D_botnet" label of Node2 changes from "0" to "1", and make corresponding rewards. At the same time, the environment state and the agent's policy gradient will be updated. The intelligent agent will use the new strategy to act on the updated environment next time. The tPatroller component will then evaluate the actions performed by the intelligent agent, and the cycle will be iterated continuously to train an intelligent agent that can intelligently optimize and identify zombie nodes.
[0078] To determine the next node to inspect, a custom neural network architecture, or agent, was designed within the tCommander component. The agent samples from the class distribution, expanding the range of action options and simultaneously determining the next node to inspect.
[0079] Figure 3 The tCommnader component is mainly composed of an agent agent and an environment. The agent agent is implemented by a customized neural network architecture. Specifically, the agent input (i.e., the state table) is N n ×N s Matrix, where N n Indicates the number of nodes, N s represents the dimension of the state table. The neural network consists of four linear layers (i.e., N s×32, 32×16, 16×4, 4×1), three ReLU activation layers and one softmax layer, the corresponding feature map is N n ×32→N n ×16→N n ×4→N n ×1. Then N n The product of the maximum value and the flag vector is subtracted from the ×1-dimensional output vector. This process aims to avoid the probability of selecting nodes that have already been detected. Finally, this difference vector is concatenated with the softmax layer and sampled using the class distribution to obtain the next node to be detected. The purpose of the class distribution is to enable the agent to discover more optimized strategies through random probability sampling.
[0080] In the specific implementation of step (5), the flow of the detected node is detected, the environment and state are updated according to the detection results, and the reward value is calculated to update the model parameters of the intelligent agent;
[0081] Specifically, when the agent makes a decision, the tPatroller component detects the traffic of the corresponding node. Specifically, the tPatroller component collects all communication traffic related to the detected node, including both senders and receivers. This traffic is then divided into sessions and its features are extracted into a set of recognized samples. The tPatroller component detects these samples and obtains its prediction results. If the recognition result differs from the class predicted by the tScouter component, the node is predicted to be an error node and its node state is corrected.
[0082] When the agent makes a decision, the tPatroller component monitors traffic at the corresponding node and provides feedback on whether the machine is engaging in botnet behavior. The detection results update the environment and state. Simultaneously, rewards are calculated to update model parameters during training. During training, the policy gradient method from reinforcement learning is utilized to optimize the agent. The agent's goal is to find a policy π that, based on the transition probability π(A|S), performs a series of actions A from state S to maximize the expected value of the immediate reward R. This expected value is formulated as follows:
[0083]
[0084] Among them, V πθ (s) represents the value function of the state, is πθ and The fixed distribution of the Markov chain. According to the strategy progression theory, the gradient calculation process of the expected value is expressed as:
[0085]
[0086]
[0087] Among them, the reward value s of checking a node t This reward consists of two parts: 1) a node correction reward; 2) a boundary discovery reward, which is the reward for discovering a robot node within its neighborhood. Since discovering a true predicted node does not generate a reward, the tCommander component can find more false nodes, earning a higher reward. Furthermore, a unit reward for node correction is set to balance the difference in their numbers.
[0088] In the specific implementation of step (6), the above steps (4) and (5) are repeated until all nodes are detected or the rewards of a certain number of consecutive searches are close to 0, thereby obtaining a detection path of the P2P botnet;
[0089] Specifically, the training process of the agent is an iterative process. tCommander searches the current state table for nodes that tScouter failed to predict. tCommander completes a training round when tPatroller has finally checked all nodes (i.e., the "Flag" value in the state table is all 1), or when the reward for a certain number of consecutive searches is close to 0 (i.e., the confidence level has not changed much). In practice, the degree of closeness to 0 is customized (e.g., confidence level greater than 0.97 or less than 0.03). If the confidence level of all nodes is greater than 0.97 or less than 0.03, the reward for subsequent searches may change very little, close to 0. Although not all nodes have been checked, the current node's state (benign or robotic) can be considered confirmed. The number of consecutive searches is the number of times the node was checked to meet the above criteria (confidence level greater than 0.97 or less than 0.03).
[0090] In the specific implementation of step (7), the above steps (4) to (6) are repeated a predetermined number of times to obtain a number of detection paths of the P2P botnet, and the shortest detection path among them is used as the optimal detection path;
[0091] Specifically, the number of repetitions is set according to your needs. The range is greater than or equal to 1. The more times you train, the easier it is to obtain an efficient detection path.
[0092] In the specific implementation of step (8), the P2P network to be detected is substituted for the P2P botnet, and steps (1) to (3) are repeated, with the node confidence as input, and the trained neural network is used to obtain the optimal detection path;
[0093] Specifically, the specific implementation of steps (1) to (3) has been described in detail above. "Using node confidence as input and using the trained neural network to obtain the optimal detection path" is the model reasoning process, which is a conventional technical means in this field and will not be repeated here.
[0094] Corresponding to the aforementioned embodiment of a P2P botnet detection method based on communication topology and network traffic, the present application also provides an embodiment of a P2P botnet detection device based on communication topology and network traffic.
[0095] Figure 4 This is a block diagram of a P2P botnet detection device based on communication topology and network traffic according to an exemplary embodiment. Figure 4 , the apparatus may include:
[0096] a synthesis module 21 for obtaining a P2P botnet and synthesizing botnet communication traffic data by superimposing the network topology of the P2P botnet, traffic based on the P2P protocol, and the real-world botnet topology onto a background traffic topology;
[0097] A topology representation module 22 is configured to perform topological representation on the botnet communication traffic data to obtain a topology graph;
[0098] A prediction module 23 is configured to use a graph convolutional network to predict the confidence level of each node in the topology graph as a robot node, wherein the robot node is a node in the botnet;
[0099] Initialization module 24, used to initialize the environment and state using the confidence of each node in the topology map and construct a neural network as an intelligent agent, and the intelligent agent determines the next node to be detected by class distribution sampling;
[0100] An updating module 25 is configured to detect the flow of the detected node, update the environment and state according to the detection result, and calculate the reward value to update the model parameters of the agent;
[0101] A repetition module 26 is used to repeat the process of the initialization module and the update module until all nodes are detected or the rewards after a certain number of consecutive searches are close to 0;
[0102] A training module 27 is used to repeat the above steps from the initialization module to the repetition module for a predetermined number of times to complete the training of the intelligent agent;
[0103] The detection module 28 is used to replace the P2P botnet with the P2P network to be detected, repeat the process from the synthesis module to the prediction module, use the node confidence as input, and use the trained neural network to obtain the optimal detection path.
[0104] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0105] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present application scheme. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0106] Accordingly, the present application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned P2P botnet detection method based on communication topology and network traffic. Figure 5 As shown in the figure, a hardware structure diagram of a device with data processing capability in which a P2P botnet detection method based on communication topology and network traffic is provided in an embodiment of the present invention, except Figure 5 In addition to the processor, memory, and network interface shown, any device with data processing capabilities in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0107] Accordingly, the present application also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement a P2P botnet detection method based on communication topology and network traffic as described above. The computer-readable storage medium can be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), an SD card, a flash card (FlashCard), etc. equipped on the device. Furthermore, the computer-readable storage medium can also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store data that has been output or is to be output.
[0108] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed in this application.
[0109] It will be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.
Claims
1. A P2P botnet detection method based on communication topology and network traffic, characterized in that: include: (1) obtaining a P2P botnet, and synthesizing botnet communication traffic data by superimposing the network topology of the P2P botnet, traffic based on the P2P protocol, and real-world botnet topology onto background traffic topology; (2) topologically representing the botnet communication traffic data to obtain a topological graph; (3) using a graph convolutional network to predict the confidence that each node in the topological graph is a robot node, where the robot node is a node in the botnet; (4) using the confidence of each node in the topological graph to initialize the environment and state and construct a neural network as an intelligent agent, the intelligent agent determines the next node to be detected by class distribution sampling; (5) detecting the traffic of the detected node, updating the environment and state according to the detection results, and calculating the reward value to update the model parameters of the intelligent agent; (6) Repeat steps (4) and (5) above until all nodes are detected or the rewards after a certain number of consecutive searches are close to 0; (7) repeating steps (4) to (6) a predetermined number of times to complete the training of the agent; (8) Replace the P2P botnet with the P2P network to be detected, repeat steps (1) to (3), use the node confidence as input, and use the trained neural network to obtain the optimal detection path.
2. The method according to claim 1, characterized in that The environment includes a node list, an edge list, and a label set of the node, and the state includes the confidence of the node.
3. The method according to claim 2, characterized in that The state adopts a state table structure, which includes the following information for each node: whether the node is a robot node, the confidence of the node, the degree of the node, the number of benign neighbors of the node that have been viewed, the number of robot nodes of the node that have been viewed, the number of unviewed neighbors of the node, and additional node features for constructing heterogeneous scenarios.
4. The method according to claim 1, wherein The neural network includes four linear layers, three ReLU activation layers and one softmax layer.
5. The method according to claim 1, characterized in that The reward value includes a reward value for node correction and a reward value for boundary discovery.
6. The method according to claim 1, wherein The process of detecting node traffic includes: Collect all communication traffic related to the detected node, including senders and receivers; The communication traffic is divided according to sessions, and features thereof are extracted to form an identified sample set, and the sample set is detected to obtain a detection result thereof.
7. The method according to claim 1, characterized in that Update the environment and status based on the test results, specifically: If the detection result is different from the category corresponding to the confidence of the detected node, the detected node is determined to be an erroneous node and its status is corrected.
8. A P2P botnet detection device based on communication topology and network traffic, characterized in that: include: a synthesis module for obtaining a P2P botnet and synthesizing botnet communication traffic data by superimposing a network topology of the P2P botnet, traffic based on a P2P protocol, and a real-world botnet topology onto a background traffic topology; A topology representation module, configured to perform topological representation on the botnet communication traffic data to obtain a topology map; a prediction module, configured to use a graph convolutional network to predict the confidence that each node in the topological graph is a robot node, wherein the robot node is a node in the botnet; An initialization module is used to initialize the environment and state using the confidence of each node in the topology map and to construct a neural network as an intelligent agent, and the intelligent agent determines the next node to be detected by class distribution sampling; An updating module, configured to detect the flow of the detected node, update the environment and state according to the detection result, and calculate a reward value to update the model parameters of the intelligent agent; The repetition module is used to repeat the above initialization module and update module process until all nodes are detected or the rewards after a certain number of consecutive searches are close to 0; A training module, configured to repeat the above steps from the initialization module to the repetition module for a predetermined number of times to complete the training of the agent; The detection module is used to replace the P2P botnet with the P2P network to be detected, repeat the process from the synthesis module to the prediction module, use the node confidence as input, and use the trained neural network to obtain the optimal detection path.
9. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
P2P botnet detection method and device and medium
CN110149331A
Network intrusion detection method, device and equipment
CN113179263A