Network attack near-source blocking method of hierarchical and domain-divided security control protocol
Through the combination of graph neural networks and deep reinforcement learning of multi-agents, cross-domain network data is collected and analyzed in real time, and cross-domain collaborative blocking strategies are generated, which solves the problems of low information interaction efficiency and slow strategy adjustment in cross-domain collaborative defense in the existing technology, and realizes accurate positioning and rapid response blocking of network attacks.
Patent Information
- Application Number
- CN202510621126.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-15
AI Technical Summary
In cross-domain collaborative defense, existing network security protection technologies have low information interaction efficiency and slow strategy adjustment, making it difficult to achieve accurate positioning and rapid response blocking of attack sources under complex network structures.
Using a method of combining graph neural networks with multi-agent deep reinforcement learning, we collect cross-domain network topology information and traffic data in real time, build a network topology structure diagram, extract node-level abnormal traffic characteristics and attack paths through graph neural network, activate multiple defense agents to generate initial defense blocking strategies, and realize cross-domain policy interaction and dynamic game iterative optimization through multi-agent deep reinforcement learning mechanism, generate cross-domain collaborative blocking strategies, and issue blocking instructions in real time.
It realizes cross-domain precise positioning and rapid response blocking of network attacks, improves the accuracy of attack positioning and real-time blocking, and reduces the overall impact of the network system.
Smart Images

Figure CN120498754A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network communication and information security, and in particular to a network attack near-source blocking method using a layered and domain-based security control protocol. Background Art
[0002] In recent years, with the rapid development of internet technology, various cyberattack techniques have become increasingly complex and covert. Cybersecurity issues are becoming increasingly cross-regional and cross-domain, placing higher demands on cybersecurity protection technologies. In particular, the destructive impact of cyberattacks is becoming increasingly prominent in important network systems such as critical infrastructure, large enterprise networks, and cloud service platforms. The importance of cross-domain coordinated defense and precise near-source blocking technologies is becoming increasingly prominent.
[0003] Traditional network security solutions often focus on single-domain network defense systems based on rules or traffic statistics, primarily employing static or rule-driven approaches such as firewalls, intrusion detection systems (IDS), and intrusion prevention systems (IPS). These solutions typically identify anomalous traffic based on predefined rule bases or historical traffic patterns to block known attacks. These technologies are simple to implement, easy to deploy, and low-cost, and are effective in general attack protection scenarios.
[0004] However, with the evolution of attack methods and the diversification of attack patterns, existing defense technologies have gradually exposed significant shortcomings, especially in accurately identifying attack sources and rapidly responding to and blocking them. Specifically, traditional technologies generally rely on traffic statistics and feature analysis within a single network domain, lacking the ability to deeply analyze network attack paths and topology structures, making it difficult to effectively capture the dynamic propagation paths of attacks. Furthermore, static rule base updates have a significant lag, making it unable to adapt to dynamic changes in attack patterns in a timely manner, severely restricting attack response speed and positioning accuracy.
[0005] To enhance cross-domain collaborative defense capabilities, the industry has recently begun exploring cross-domain joint defense and control technologies. By deploying collaborative defense systems across different network domains and leveraging cross-domain data sharing and linkage strategies, these technologies enhance overall protection against network attacks. Typical implementations of these technologies include cross-domain threat intelligence sharing mechanisms and distributed anomaly detection systems. However, due to significant differences in administrative permissions, data privacy, and security policies across different network domains, cross-domain collaboration suffers from inefficient information exchange and slow policy adjustments, making it difficult to meet the actual needs of efficient collaborative blocking. This is particularly true for complex advanced persistent threats (APTs) or large-scale distributed denial of service (DDoS) attacks, resulting in unsatisfactory responses.
[0006] On the other hand, in recent years, machine learning, especially deep learning technology, has begun to be widely used in the field of network security. For example, models such as convolutional neural networks (CNN) and recurrent neural networks (RNN) have been used in traffic feature analysis and attack behavior identification, and have made some progress. However, most traditional deep learning models focus on independent network traffic data, ignoring network topology and inter-node correlation information. Therefore, the ability to accurately locate the source of attacks in complex network structures remains relatively weak. Graph neural networks (GNNs), as an emerging deep learning technology, can effectively model the relationship characteristics between nodes and are currently being used in the field of network node anomaly detection. However, their application in the field of accurately locating cross-domain network attacks and blocking near-source attacks is still blank, and no substantial technological breakthroughs have yet been achieved.
[0007] Furthermore, multi-agent deep reinforcement learning (MADRL) has become a research hotspot in recent years. MADRL improves the flexibility and real-time performance of network security policy adjustments through collaborative game-playing and adaptive learning among multiple agents. However, existing MADRL technologies have primarily focused on theoretical research or optimization of local defense strategies in a single domain. MADRL technology has not yet been fully applied to practical cross-domain collaborative defense strategies, lacking a truly intelligent cross-domain linkage mechanism.
[0008] Therefore, how to provide a method for blocking network attacks near the source of a layered and domain-specific security control protocol is an urgent problem that those skilled in the art need to solve. Summary of the Invention
[0009] One purpose of the present invention is to propose a network attack near-source blocking method for a layered and domain-based security control protocol. The present invention adopts a method that integrates graph neural networks and multi-agent deep reinforcement learning. By real-time collection of network topology information and traffic data in a cross-domain network environment, a network topology structure diagram is constructed, and the graph neural network is used to dynamically extract node-level abnormal traffic characteristics and attack paths and accurately determine the near-source location of the network attack; then, multiple defense agents deployed in each network domain are activated, and an initial defense blocking strategy is generated based on the network attack characteristics and network topology information of the domain. Using this as the initial input, cross-domain strategy interaction and dynamic game iterative optimization are realized through the multi-agent deep reinforcement learning mechanism to form a cross-domain collaborative blocking strategy; finally, according to the cross-domain collaborative blocking strategy, precise blocking instructions are issued in real time to the near-source domain of the network attack to perform real-time blocking of network attack traffic; the present invention realizes cross-domain precise positioning and rapid response blocking of network attacks, significantly improves the accuracy of attack positioning and the real-time performance of blocking, and effectively reduces the overall impact of attack events on the network system.
[0010] A method for blocking network attacks near the source using a layered and domain-specific security control protocol according to an embodiment of the present invention includes the following steps:
[0011] S1. Collect network topology information and traffic data in a cross-domain network environment in real time and generate a network topology diagram;
[0012] S2. Using a graph neural network to dynamically extract features from nodes and edge relationships in the network topology graph to obtain node-level abnormal traffic features and topological paths of network attacks;
[0013] S3. Determine the near-source location of the network attack behavior based on the extracted topological path and abnormal traffic characteristics;
[0014] S4. After determining the near-source location, multiple defense agents deployed in each network domain are activated. Each defense agent generates an initial defense blocking strategy based on the network attack characteristics and network topology information of the domain.
[0015] S5. Multiple defense agents use the initial defense blocking strategy as initial input, conduct cross-domain strategy interaction and dynamic game iterative optimization through a multi-agent deep reinforcement learning mechanism, and generate a cross-domain collaborative blocking strategy.
[0016] S6. Based on the cross-domain collaborative blocking strategy, a blocking instruction is issued to the network attack near-source domain to perform real-time network attack traffic blocking.
[0017] Optionally, the S1 specifically includes:
[0018] S11. Traffic collection equipment deployed at the network domain boundary synchronously collects source IP address, destination IP address, packet size, transmission protocol type, and timestamp data with a collection cycle of 10 to 30 seconds;
[0019] S12. Using a routing device within the network domain to obtain routing status information in real time, the routing status information includes link connectivity status, node device IP address, port connection relationship, and link delay value;
[0020] S13. Based on the source IP address, destination IP address, and node device IP address data, and in combination with the port connection relationship, a correspondence between network nodes and links between nodes is generated to form an initial network topology structure;
[0021] S14, performing time series synchronization processing on the node traffic data in the network topology structure according to the timestamp data to obtain network traffic data under a unified time scale;
[0022] S15. Filtering out links with delays greater than 100 milliseconds based on the link delay values, removing links with abnormal delays in the initial network topology, and obtaining an updated network topology;
[0023] S16. Based on the updated network topology structure and combined with the network traffic data, a network topology structure diagram for graph neural network analysis is generated.
[0024] Optionally, the S2 specifically includes:
[0025] S21, taking the network topology graph as input, obtaining the initial topological feature vector of the target node by the adjacent node aggregation method;
[0026] S22, performing a multi-layer graph convolution operation on the initial topological feature vector, wherein each layer of convolution operation takes the features of the current node and adjacent nodes as input, updates the node topological feature vector, and the number of convolution layers is set to 3 to 5 layers;
[0027] S23, concatenating and fusing the convolved node topology feature vector and the node traffic data to obtain a node feature vector after topology and traffic fusion;
[0028] S24. Calculate the feature weight coefficient for the fused node feature vector using the attention mechanism, and perform weighted combination of the node feature vectors according to the weight coefficient to obtain a node-level dynamic embedding vector;
[0029] S25. Based on the node-level dynamic embedding vector, a density clustering method is used to identify abnormal traffic characteristics of the node, and the abnormality judgment threshold is set to an inter-cluster distance greater than 0.8;
[0030] S26. Based on the identified abnormal traffic features, reversely trace the embedding vector similarity of adjacent nodes starting from the abnormal traffic features. When the cosine similarity value of the embedding vectors between nodes is greater than 0.9, it is determined that there is a network attack topology path between the nodes, and the topology path of the network attack is determined.
[0031] Optionally, the S25 specifically includes:
[0032] S251, taking the node-level dynamic embedding vector as input, calculate the Euclidean distance between any two nodes and generate a node distance matrix D, where the matrix element d ij Represents the Euclidean distance between node i and node j;
[0033] S252, according to the node distance matrix D, the distance values between nodes are arranged in ascending order, and the distance values of the nodes in the top 15% of the distance ranking are selected as the distance determination basis of the local high-density area;
[0034] S253, calculate the local density value ρ for each node i :
[0035]
[0036] Where N is the total number of nodes, dc is the basis for determining the distance to the local high-density area, χ(x) is the cutoff function, when x<0, χ(x)=1, when x≥0, χ(x)=0;
[0037] S254. Calculate the node density difference δ for each node i , For the node with the highest density, set δ i =max(d ij );
[0038] S255, according to the local density value ρ i and density difference δ i , calculate the cluster center metric value γ of the node i ,γ i =ρ i ×δ i , select the cluster center metric value γ i Nodes with a score greater than 0.5 are used as cluster centers;
[0039] S256. Cluster nodes with the cluster center as the core node and the distance between clusters greater than 0.8 as the abnormality judgment threshold. Nodes that are not included in any cluster or have a distance from the cluster center greater than 0.8 are determined to be nodes with abnormal traffic characteristics.
[0040] Optionally, the S3 specifically includes:
[0041] S31. Based on the network attack topology path, select the node on the topology path that directly corresponds to the abnormal traffic characteristics as the node to be located;
[0042] S32, obtaining the data packet size, protocol type, and transmission timestamp data of the node to be located within three consecutive acquisition cycles;
[0043] S33, calculate the variance value of the data packet size of the node to be located in three consecutive acquisition cycles And set A packet size greater than 1500 bytes is considered abnormal;
[0044] S34, for the protocol type data of the node to be located, counting the ratio of packets of the Transmission Control Protocol (TCP) and the User Datagram Protocol (UDP), and determining that the protocol type is abnormal when the ratio of UDP packets exceeds 60%;
[0045] S35, based on the timestamp data transmitted by the node to be located, calculate the average value t of the data transmission interval of the node in three consecutive acquisition cycles avg , when t avg When it is less than 50 milliseconds, the node transmission frequency is determined to be abnormal;
[0046] S36. When the node to be located meets the three conditions of abnormal data packet size, abnormal protocol type, and abnormal node transmission frequency at the same time, the node to be located is determined to be the near source location of the network attack behavior.
[0047] Optionally, the S4 specifically includes:
[0048] S41. After determining the near-source location of the network attack behavior, activate the defense agent deployed in each network domain to obtain the source IP address, destination IP address, and number of data packets of the near-source location node in the network domain;
[0049] S42. Based on the number of data packets in the last three collection cycles, count the number of data packets sent by the nodes near the source location in the network domain to different destination IP addresses, and mark the destination IP addresses for which the number of data packets from a single destination IP address accounts for more than 30% as high-frequency destination IP addresses;
[0050] S43. Based on the high-frequency target IP address, determine the link load value of the node connected to the high-frequency target IP address in the current network domain;
[0051] S44. Identify links within the network domain with a link load value greater than 5000 bytes / second, and mark nodes associated with the links as key defense nodes;
[0052] S45. Determine a node blocking priority based on a link load value of the key defense node. When the link load value is greater than 10,000 bytes / second, the blocking priority is set to high. When the link load value is between 5,000 bytes / second and 10,000 bytes / second, the blocking priority is set to medium.
[0053] S46. Each defense agent generates an initial defense blocking strategy based on the blocking priority and link load value of the key defense nodes in the network domain. The initial defense blocking strategy includes the node blocking time interval and link blocking direction. The high-priority node blocking time interval is set to 5 seconds, and the medium-priority node blocking time interval is set to 10 seconds.
[0054] Optionally, the S5 specifically includes:
[0055] S51. Take the initial defense blocking strategy of each defense agent as the initial strategy set, and define the strategy space of each defense agent as A. Each strategy in the strategy space A is denoted as a i ;
[0056] S52, executing strategy interaction based on the initial strategy set, where each defense agent calculates the reward value of its own strategy in the current state according to the strategy states of other network domain defense agents;
[0057] S53. Each defensive agent updates its own strategy value function Q based on the reward value. i (s t ,a i ):
[0058]
[0059] Where η represents the learning rate, ranging from 0.01 to 0.1, and γ represents the policy value discount factor, ranging from 0.8 to 0.95;
[0060] S54, adjusting the strategy selection probability in the strategy space A of each defense agent based on the updated strategy value function;
[0061] S55, executing the strategy selection probability update of the defense agent, with the number of iterations set to 10 to 20 times;
[0062] S56. After iterative optimization, the strategy with the highest probability of selection of the defense agent strategy of each network domain is selected to form a cross-domain collaborative blocking strategy.
[0063] Optionally, the S52 specifically includes:
[0064] S521, record the current state of the defense agent as s t , current state s t It is defined as the set of link load values and blocking priorities of key defense nodes within the network domain;
[0065] S522, obtain the current state s of other network domain defense agents t Strategy a j , strategy a j Including node blocking time interval and link blocking direction;
[0066] S523, calculate other network domain defense agent execution strategy a j The number of data packets blocked after pj , the statistical time is set to the latest three collection cycles;
[0067] S524, calculate other network domain defense agent execution strategy a j The resulting link blocking cost C j ;
[0068] S525, based on the number of data packets blocked N pj and link blocking cost C j , calculate the network domain defense agent to execute its own strategy a i The return value r i (s t ,a i ):
[0069] r i (s t ,a i )=α×N pi -β×C i ;
[0070] Among them, α is the reward weight coefficient of the number of packet blocking, ranging from 0.1 to 1.0, β is the reward penalty coefficient of the link blocking cost, ranging from 0.1 to 1.0, N pi and C i Respectively represent the network domain defense agent strategy a i The number of blocked packets and the link blocking cost.
[0071] Optionally, the S6 specifically includes:
[0072] S61. Based on the cross-domain collaborative blocking strategy, determine the blocking time interval and link blocking direction of key defense nodes in the network attack near-source domain;
[0073] S62. Generate a blocking instruction for the key defense node based on the blocking time interval;
[0074] S63. Send the blocking instruction to the network edge device in the domain close to the network attack source in an encrypted and authenticated manner. The authentication method adopts 128-bit symmetric key encryption, and the key update period is set to 12 hours.
[0075] S64, after the network edge device in the network attack near source domain receives the blocking instruction, at the starting blocking time T s Modify the packet forwarding rules for the specified port to discard incoming or outgoing packets of the specified port;
[0076] S65. The network edge device records the number of data packets blocked by the designated port in real time during the execution of the blocking instruction, with the time resolution of the data packet number statistics set to 1 second.
[0077] S66, blocking duration T d After the end, the network edge device automatically restores the designated port packet forwarding rules to the state before blocking, and feeds back the number of packets during the blocking period to the defense agent of the corresponding network domain.
[0078] Optionally, the S62 specifically includes:
[0079] S621, based on the blocking time interval of the key defense node, calculate the starting blocking time T of the blocking instruction s , starting blocking time T s Defined as the next integer second after the end of the current acquisition cycle;
[0080] S622. Determine the blocking duration T of the key defense nodes based on the cross-domain collaborative blocking strategy. d When the blocking priority is high, the blocking duration is T d Set to 5 seconds. When the blocking priority is medium, the blocking duration is T d Set to 10 seconds;
[0081] S623. Obtain the port identifier of the link corresponding to the key defense node. The port identifier consists of the node IP address and the port number. The node IP address uses the IPv4 format, and the port number ranges from 1024 to 65535.
[0082] S624, based on the starting blocking time T s , blocking duration T d and the port identifier to generate a port-level blocking instruction, which includes a start timestamp, an end timestamp, and a specific port number;
[0083] S625: Verify the compliance of the port-level blocking instruction, including the duration of the blocking, T d Whether the time is within the range of 5 to 10 seconds and whether the port number is within the range of 1024 to 65535. After verification, the blocking instruction is confirmed to be valid;
[0084] S626: Store the verified port-level blocking instruction into the blocking instruction queue to be sent, and s They will be issued in the order of time.
[0085] The beneficial effects of the present invention are:
[0086] (1) The present invention uses a graph neural network to dynamically extract features from the nodes and edge relationships of the network topology diagram, clearly obtains node-level abnormal traffic features and network attack topology paths, and achieves accurate positioning of the near-source location of network attacks, effectively improving the identification accuracy of the source of network attacks and enhancing the rapid response capability and accuracy of attack positioning in complex cross-domain network environments.
[0087] (2) The present invention activates the defense agents deployed in each network domain and generates an initial defense blocking strategy. It uses a multi-agent deep reinforcement learning mechanism to realize the interaction and dynamic game iterative optimization of cross-domain defense strategies. It can achieve collaborative defense and adaptive strategy adjustment among multiple network domains, significantly improve the defense response efficiency and the effectiveness of the blocking strategy, and show better real-time adaptability and coordination in the prevention and control of cross-domain network attacks.
[0088] (3) In terms of real-time response to network attack blocking, the present invention generates precise blocking instructions at the port level, verifies compliance, and encrypts and sends them to network edge devices in the near-source domain of the network attack. This effectively solves the problems of large response delays and coarse blocking granularity in traditional technologies, breaks through the bottleneck of precise blocking at the second level that is difficult to achieve with existing technologies, and realizes real-time and precise execution of blocking instructions, thereby effectively improving the overall defense performance and engineering application level of cross-domain network security prevention and control. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0090] Figure 1 This is an overall flow chart of a network attack near-source blocking method of a layered and domain-based security control protocol proposed by the present invention;
[0091] Figure 2 This is a flowchart of the dynamic feature extraction of a graph neural network for a network attack near-source blocking method of a layered and domain-based security control protocol proposed by the present invention;
[0092] Figure 3 This is a flowchart of the cross-domain multi-agent deep reinforcement learning strategy optimization for a network attack near-source blocking method for a layered and domain-based security control protocol proposed in the present invention. DETAILED DESCRIPTION
[0093] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0094] refer to Figure 1-Figure 3 A method for blocking network attacks near the source based on a layered and domain-specific security control protocol, characterized by comprising the following steps:
[0095] S1. Collect network topology information and traffic data in a cross-domain network environment in real time and generate a network topology diagram;
[0096] S2. Using a graph neural network to dynamically extract features from nodes and edge relationships in the network topology graph to obtain node-level abnormal traffic features and topological paths of network attacks;
[0097] S3. Determine the near-source location of the network attack behavior based on the extracted topological path and abnormal traffic characteristics;
[0098] S4. After determining the near-source location, multiple defense agents deployed in each network domain are activated. Each defense agent generates an initial defense blocking strategy based on the network attack characteristics and network topology information of the domain.
[0099] S5. Multiple defense agents use the initial defense blocking strategy as initial input, conduct cross-domain strategy interaction and dynamic game iterative optimization through a multi-agent deep reinforcement learning mechanism, and generate a cross-domain collaborative blocking strategy.
[0100] S6. Based on the cross-domain collaborative blocking strategy, a blocking instruction is issued to the network attack near-source domain to perform real-time network attack traffic blocking.
[0101] During the specific implementation of the present invention, the collection of network topology information and traffic data in a cross-domain network environment is specifically achieved through traffic collection equipment and intra-domain routing equipment deployed at the network domain boundary. The traffic collection equipment synchronously obtains the source IP address, destination IP address, data packet size, transmission protocol type and timestamp data with a collection period of 10 seconds to 30 seconds. The intra-domain routing equipment obtains the link connectivity status, node device IP address, port connection relationship and link delay value in real time. After obtaining the data, it forms a network node and link correspondence based on the port connection relationship, and then obtains an updated network topology structure diagram by eliminating abnormal links with link delays greater than 100 milliseconds. The graph neural network adopted in the present invention includes 3 layers of graph convolution operations, and uses the attention mechanism to calculate the feature weight coefficient to generate a dynamic embedding vector; the recognition of node-level abnormal traffic features adopts dense The method is implemented by using a degree clustering method, in which the cluster distance of 0.8 is used as the anomaly judgment threshold in the local density calculation; the initial defense blocking strategy of the defense agent is determined by the link load value of the defense key node in each network domain. When the link load value is greater than 10,000 bytes / second, the blocking priority is high and the blocking time interval is set to 5 seconds. When the link load value is between 5,000 bytes / second and 10,000 bytes / second, the blocking priority is medium and the blocking time interval is set to 10 seconds; the multi-agent reinforcement learning iterative optimization times are 10 to 20 times, and the strategy value discount factor is 0.8 to 0.95 and the learning rate is 0.01 to 0.1. After finally forming a cross-domain collaborative blocking strategy, it uses 128-bit symmetric key encryption authentication method, updates the key every 12 hours, and sends the blocking instruction to the network edge device in the near-source domain of the network attack to complete real-time attack traffic blocking.
[0102] The present invention constructs a layered and domain-based network attack near-source blocking method by integrating graph neural networks and multi-agent deep reinforcement learning methods. Compared with the existing technology, it can achieve accurate positioning of the source of network attacks, improve the accuracy of abnormal traffic feature recognition and response speed, and at the same time utilize the collaborative defense strategy of cross-domain multi-agents to realize dynamic real-time adjustment and cross-domain linkage of defense strategies, significantly improving the real-time blocking efficiency and blocking accuracy of network attacks, and reducing the potential impact range caused by network security incidents.
[0103] In this embodiment, S1 specifically includes:
[0104] S11. Traffic collection equipment deployed at the network domain boundary synchronously collects source IP address, destination IP address, packet size, transmission protocol type, and timestamp data with a collection cycle of 10 to 30 seconds;
[0105] S12. Using a routing device within the network domain to obtain routing status information in real time, the routing status information includes link connectivity status, node device IP address, port connection relationship, and link delay value;
[0106] S13. Based on the source IP address, destination IP address, and node device IP address data, and in combination with the port connection relationship, a correspondence between network nodes and links between nodes is generated to form an initial network topology structure;
[0107] S14, performing time series synchronization processing on the node traffic data in the network topology structure according to the timestamp data to obtain network traffic data under a unified time scale;
[0108] S15. Filtering out links with delays greater than 100 milliseconds based on the link delay values, removing links with abnormal delays in the initial network topology, and obtaining an updated network topology;
[0109] S16. Based on the updated network topology structure and combined with the network traffic data, a network topology structure diagram for graph neural network analysis is generated.
[0110] During the specific implementation of the present invention, the flow collection equipment deployed at the boundary of the network domain synchronizes the system clock once every 10 seconds to ensure the accuracy of the data packet collection cycle, and the duration of each collection cycle is uniformly set to 20 seconds. The routing device obtains link connectivity status information through real-time polling of the SNMP protocol. The link delay value is measured using the round-trip time (RTT) between routing devices in the network domain, and the average value of the RTT values of three consecutive samples is taken as the final link delay value. The generation process of the initial network topology structure uses an adjacency matrix to represent the node connection relationship, wherein the matrix elements are defined as the port connection relationship corresponding to the node device IP address, and the specific value range of the port connection relationship is 1024 to 65535. The time series synchronization processing specifically uses the NTP protocol to obtain a unified network clock reference to ensure that the time accuracy of the node flow data is controlled within the millisecond level. The link anomaly judgment threshold is set according to the network performance requirements. Links with delays exceeding 100 milliseconds are directly marked as abnormal links and eliminated. The updated network topology diagram is stored in a standard graph structure (GraphML) format, and clearly marks the node ID, edge relationship, link delay value and port connection relationship, which serves as the input data for the subsequent dynamic feature extraction of the graph neural network.
[0111] Through strict and standardized data collection and preprocessing processes, the present invention ensures the high accuracy and consistency of network topology information and traffic data, reduces the impact of data synchronization errors on subsequent analysis, effectively improves the accuracy of abnormal node identification and network attack source location, and enhances the real-time and accuracy of the overall network attack defense strategy formulation.
[0112] In this embodiment, S2 specifically includes:
[0113] S21, taking the network topology graph as input, obtaining the initial topological feature vector of the target node by the adjacent node aggregation method;
[0114] S22, performing a multi-layer graph convolution operation on the initial topological feature vector, wherein each layer of convolution operation takes the features of the current node and adjacent nodes as input, updates the node topological feature vector, and the number of convolution layers is set to 3 to 5 layers;
[0115] S23, concatenating and fusing the convolved node topology feature vector and the node traffic data to obtain a node feature vector after topology and traffic fusion;
[0116] S24. Calculate the feature weight coefficient for the fused node feature vector using the attention mechanism, and perform weighted combination of the node feature vectors according to the weight coefficient to obtain a node-level dynamic embedding vector;
[0117] S25. Based on the node-level dynamic embedding vector, a density clustering method is used to identify abnormal traffic characteristics of the node, and the abnormality judgment threshold is set to an inter-cluster distance greater than 0.8;
[0118] S26. Based on the identified abnormal traffic features, reversely trace the embedding vector similarity of adjacent nodes starting from the abnormal traffic features. When the cosine similarity value of the embedding vectors between nodes is greater than 0.9, it is determined that there is a network attack topology path between the nodes, and the topology path of the network attack is determined.
[0119] During the specific implementation of the present invention, the adjacent node aggregation method adopts the average aggregation method, and the dimension of the initial topological feature vector is specifically set to 128 dimensions; the graph convolution operation adopts the standard GraphSAGE algorithm, in which each layer of graph convolution uses the ReLU function as the nonlinear activation function, and the feature dimension after each graph convolution update is fixed to 128 dimensions; the node feature vectors after topology and traffic fusion are fused by a simple splicing method, and the dimension of the spliced node feature vector is uniformly expanded to 256 dimensions; the attention mechanism adopts the standard self-attention calculation method, the feature weight coefficient is normalized by the Softmax function, and the feature vector dimension remains unchanged at 256 dimensions during calculation; the density clustering method uses a clustering method based on local density peaks for abnormal feature identification, and the truncation distance dc of the local density value ρ selects the value in the top 15% after sorting the Euclidean distance between nodes from small to large as the truncation threshold; the embedding vector similarity adopts the standard cosine similarity calculation method, and when the calculated cosine similarity value is greater than 0.9, it is determined that there is an attack topological path between the two nodes.
[0120] Through a specific and rigorous graph neural network feature extraction and anomaly identification process, the present invention significantly improves the accuracy of abnormal traffic feature detection and the precision of network attack topology path identification, enhances the reliability and real-time performance of attack source tracking and network attack defense strategy formulation, and significantly improves the overall performance of network security protection.
[0121] In this embodiment, the S25 specifically includes:
[0122] S251, taking the node-level dynamic embedding vector as input, calculate the Euclidean distance between any two nodes and generate a node distance matrix D, where the matrix element d ij Represents the Euclidean distance between node i and node j;
[0123] S252, according to the node distance matrix D, the distance values between nodes are arranged in ascending order, and the distance values of the nodes in the top 15% of the distance ranking are selected as the distance determination basis of the local high-density area;
[0124] S253, calculate the local density value ρ for each node i :
[0125]
[0126] Where N is the total number of nodes, d c is the basis for determining the distance to the local high-density area, χ(x) is the cutoff function, when x<0, χ(x)=1, when x≥0, χ(x)=0;
[0127] S254. Calculate the node density difference δ for each node i , For the node with the highest density, set δ i =max(d ij );
[0128] S255, according to the local density value ρ i and density difference δ i , calculate the cluster center metric value γ of the node i ,γ i =ρ i ×δ i , select the cluster center metric value γ i Nodes with a score greater than 0.5 are used as cluster centers;
[0129] S256. Cluster nodes with the cluster center as the core node and the distance between clusters greater than 0.8 as the abnormality judgment threshold. Nodes that are not included in any cluster or have a distance from the cluster center greater than 0.8 are determined to be nodes with abnormal traffic characteristics.
[0130] In the specific implementation process of the present invention, the node distance matrix D is obtained by performing pairwise Euclidean distance calculation on the input node-level dynamic embedding vector, and the embedding vector dimension of each node is unified to 256 dimensions; the distance values between nodes are arranged in ascending order by a quick sorting algorithm; the truncation function χ(x) in the local density calculation adopts a unit step function, wherein the truncation distance dc is determined by calculating the total number of distances between all nodes and sorting them in ascending order, and selecting the specific distance value corresponding to the 15% position after sorting as dc; the calculation method of the density difference δ is clearly to traverse the node density The degree value ρ is used to traverse and compare the distances between all nodes with a higher density than the current node and the current node, and the minimum distance value is selected as the δ value of the current node. The δ value of the node with the highest density is the maximum distance in the distance matrix D. In the calculation process of the cluster center metric value γ, ρ and δ are normalized to the interval [0, 1] and then multiplied to obtain γ to ensure the consistency of cluster center selection in networks of different scales. When clustering, the distance between each node and the cluster center is obtained by calculating the Euclidean distance of the embedding vectors between nodes. Nodes with a distance greater than 0.8 are directly marked as abnormal nodes.
[0131] The present invention adopts the density clustering method to identify abnormal node features, clarifies the specific calculation methods of the distance matrix, local density value and cluster center measurement value, ensures the rigor and objectivity of the abnormal node judgment process, significantly improves the accuracy of node abnormal traffic feature identification, provides a reliable basis for the accurate tracking of network attack topology paths, and improves the effectiveness of the overall network security protection solution.
[0132] In this embodiment, S3 specifically includes:
[0133] S31. Based on the network attack topology path, select the node on the topology path that directly corresponds to the abnormal traffic characteristics as the node to be located;
[0134] S32, obtaining the data packet size, protocol type, and transmission timestamp data of the node to be located within three consecutive acquisition cycles;
[0135] S33, calculate the variance value of the data packet size of the node to be located in three consecutive acquisition cycles And set A packet size greater than 1500 bytes is considered abnormal;
[0136] S34, for the protocol type data of the node to be located, counting the ratio of packets of the Transmission Control Protocol (TCP) and the User Datagram Protocol (UDP), and determining that the protocol type is abnormal when the ratio of UDP packets exceeds 60%;
[0137] S35, based on the timestamp data transmitted by the node to be located, calculate the average value t of the data transmission interval of the node in three consecutive acquisition cycles avg , when t avg When it is less than 50 milliseconds, the node transmission frequency is determined to be abnormal;
[0138] S36. When the node to be located meets the three conditions of abnormal data packet size, abnormal protocol type, and abnormal node transmission frequency at the same time, the node to be located is determined to be the near source location of the network attack behavior.
[0139] In the specific implementation process of the present invention, the method for determining the node to be located on the topological path that directly corresponds to the abnormal traffic characteristics is as follows: the node determined as an abnormal node in the density clustering stage and having the highest topological centrality on the topological path is selected as the node to be located; the variance value σ_p of the packet size 2The standard variance formula is used for calculation. The number of data packets in three consecutive collection cycles shall not be less than 300. In the proportion statistics of protocol type data, TCP and UDP data packets are counted separately and the UDP protocol ratio is calculated by dividing the number of UDP packets by the total number of TCP and UDP data packets. The calculation method of the average value t_avg of the node data transmission interval is specifically as follows: the sum of the timestamp differences between all adjacent data packets in three consecutive collection cycles is divided by the total number of timestamp differences, where the statistical accuracy of the timestamp differences between adjacent data packets is at the millisecond level. For nodes that meet the three conditions of abnormal data packet size, abnormal protocol type and abnormal transmission frequency at the same time, the system automatically marks them as near-source nodes of network attacks, and records the IP address, abnormal judgment time and abnormal indicator value of the marked node for subsequent processing.
[0140] The present invention ensures the accuracy and objectivity of abnormal behavior judgment of the node to be located through clear and detailed data statistical analysis and abnormal node judgment methods, effectively avoiding the problems of strong subjectivity and insufficient accuracy in node anomaly judgment in the existing technology, and significantly improves the accuracy and reliability of identifying near-source nodes of network attacks, thereby improving the overall effect of network security defense strategies.
[0141] In this embodiment, the S4 specifically includes:
[0142] S41. After determining the near-source location of the network attack behavior, activate the defense agent deployed in each network domain to obtain the source IP address, destination IP address, and number of data packets of the near-source location node in the network domain;
[0143] S42. Based on the number of data packets in the last three collection cycles, count the number of data packets sent by the nodes near the source location in the network domain to different destination IP addresses, and mark the destination IP addresses for which the number of data packets from a single destination IP address accounts for more than 30% as high-frequency destination IP addresses;
[0144] S43. Based on the high-frequency target IP address, determine the link load value of the node connected to the high-frequency target IP address in the current network domain;
[0145] S44. Identify links within the network domain with a link load value greater than 5000 bytes / second, and mark nodes associated with the links as key defense nodes;
[0146] S45. Determine a node blocking priority based on a link load value of the key defense node. When the link load value is greater than 10,000 bytes / second, the blocking priority is set to high. When the link load value is between 5,000 bytes / second and 10,000 bytes / second, the blocking priority is set to medium.
[0147] S46. Each defense agent generates an initial defense blocking strategy based on the blocking priority and link load value of the key defense nodes in the network domain. The initial defense blocking strategy includes the node blocking time interval and link blocking direction. The high-priority node blocking time interval is set to 5 seconds, and the medium-priority node blocking time interval is set to 10 seconds.
[0148] During the specific implementation of the present invention, the activation of the defense intelligent body is completed by issuing an activation instruction through the control device in the network domain, and the activation process duration is controlled within 2 seconds; the statistical method of the high-frequency target IP address is implemented by a counting sorting algorithm, and the top 10 IP addresses whose number of single destination IP address data packets accounts for more than 30% of the total number of data packets in this domain are selected as high-frequency target IP addresses; the specific calculation method of the link load value is: link load value (bytes / second) = number of data packets transmitted on the link × average size of a single data packet ÷ total time of the acquisition cycle, and the average size of a single data packet is uniformly 512 bytes Section; Link identification with a link load value greater than 5000 bytes / second is determined by the threshold screening method, and the IP address and port number are used to mark the key defense nodes; the blocking priority is specifically divided into two levels: high and medium. The link load value greater than 10000 bytes / second is high priority, and the blocking policy execution time is uniformly 5 seconds. The link load value between 5000 bytes / second and 10000 bytes / second is medium priority, and the blocking policy execution time is uniformly 10 seconds; the blocking direction is determined by the data flow transmission direction during the link load value statistics process, and the link port consistent with the data flow direction is defined as the blocking port.
[0149] The present invention accurately identifies the links and nodes that require key defense within the network domain through a specific and rigorous method for calculating node link load values and marking key defense nodes, effectively reducing the defects of inaccurate defense strategy execution and blind blocking actions in the existing technology, significantly improving the effectiveness and pertinence of the defense strategy, and further enhancing the reliability and response speed of the network attack prevention and control solution.
[0150] In this embodiment, the S5 specifically includes:
[0151] S51. Take the initial defense blocking strategy of each defense agent as the initial strategy set, and define the strategy space of each defense agent as A. Each strategy in the strategy space A is denoted as a i ;
[0152] S52: Execute strategy interaction based on the initial strategy set. Each defense agent calculates the reward value of its own strategy in the current state according to the strategy state of other network domain defense agents.
[0153] S53. Each defensive agent updates its own strategy value function Q based on the reward value. i (st ,a i ):
[0154]
[0155] Where η represents the learning rate, ranging from 0.01 to 0.1, and γ represents the policy value discount factor, ranging from 0.8 to 0.95;
[0156] S54, adjusting the strategy selection probability in the strategy space A of each defense agent based on the updated strategy value function;
[0157] S55, executing the strategy selection probability update of the defense agent, with the number of iterations set to 10 to 20 times;
[0158] S56. After iterative optimization, the strategy with the highest probability of selection of the defense agent strategy of each network domain is selected to form a cross-domain collaborative blocking strategy.
[0159] During the specific implementation of the present invention, the strategy space A of each defense agent is established in a discretized manner. The specific strategy a_i in the strategy space A is formed by a combination of the node blocking time interval and the link blocking direction, and the specific number of strategies is controlled within 20; the strategy state is determined using the link load value and the node blocking priority of the defense key node as input parameters; the strategy reward value is calculated by multiplying the total number of successfully blocked data packets by the data packet blocking reward weight coefficient (ranging from 0.1 to 1.0) minus the link blocking cost, where the link blocking cost is the link load value multiplied by the blocking duration and then multiplied by the link blocking cost penalty coefficient (ranging from 0.1 to 1.0); when the strategy value function is updated, the learning rate takes a fixed value of 0.05, and the strategy value discount factor takes a fixed value of 0.9; the adjustment of the strategy selection probability is implemented by using the Softmax function calculation, where the temperature coefficient of the strategy selection probability is fixed at 0.8; after each iterative optimization process, the top five strategies with the highest strategy selection probability are retained, the strategy with the lowest probability is eliminated, and a new random strategy is added to ensure the diversity of the strategy set.
[0160] Through a specific and clear strategy space construction method and strategy iteration optimization method, the present invention enables multi-agent strategy interaction and collaborative defense to achieve more efficient strategy optimization and selection, significantly improving the adaptability and robustness of cross-domain collaborative defense strategies, effectively reducing network attack response delays and blocking costs, and enhancing the intelligence and practicality of the overall network security prevention and control solution.
[0161] In this embodiment, the S52 specifically includes:
[0162] S521, record the current state of the defense agent as st , current state s t It is defined as the set of link load values and blocking priorities of key defense nodes within the network domain;
[0163] S522, obtain the current state s of other network domain defense agents t Strategy a j , strategy a j Including node blocking time interval and link blocking direction;
[0164] S523, calculate other network domain defense agent execution strategy a j The number of data packets blocked after pj , the statistical time is set to the latest three collection cycles;
[0165] S524, calculate other network domain defense agent execution strategy a j The resulting link blocking cost C j ;
[0166] S525, based on the number of data packets blocked N pj and link blocking cost C j , calculate the network domain defense agent to execute its own strategy a i The return value r i (s t ,a i ):
[0167] r i (s t ,a i )=α×N pi -β×C i ;
[0168] Among them, α is the reward weight coefficient of the number of packet blocking, ranging from 0.1 to 1.0, β is the reward penalty coefficient of the link blocking cost, ranging from 0.1 to 1.0, N pi and C i Respectively represent the network domain defense agent strategy a i The number of blocked packets and the link blocking cost.
[0169] During the specific implementation of the present invention, the current state of the defense agent is obtained through real-time monitoring of the link load value and real-time updating of the blocking priority, wherein the update period of the link load value is once every 10 seconds, and the blocking priority is determined based on the real-time update of the link load value; the acquisition method of the defense agent strategy for other network domains is to obtain it through timed synchronization through a secure communication channel, with a synchronization period of once every 5 seconds, and the strategy synchronization communication adopts an encryption authentication mechanism to ensure communication security; the number of packet blocking is specifically based on the total number of incoming data packets at the designated port after the defense agent executes the blocking action. The statistical data packets are data packets with a uniform length of 512 bytes, and the statistical period is the last three consecutive collection periods; the link blocking cost is specifically obtained by multiplying the link load value by the duration of the strategy execution, wherein the link load value takes the average value of the link load value within the statistical period; the reward value weight coefficient and the penalty coefficient, the reward value calculation formula for the defense agent to execute its own strategy is strictly based on the provided formula, and the number of packet blocking and the link blocking cost are respectively counted and calculated in the same way.
[0170] The present invention accurately evaluates the actual effect of each defense agent strategy through clear and specific strategy state definitions and reward value calculation methods, significantly improving the accuracy of cross-domain defense strategy interactions and the rationality of defense actions, and effectively improving the shortcomings of non-objective strategy evaluation and blind defense decision-making in traditional cross-domain network attack defense, thereby improving the performance and effectiveness of the overall cross-domain collaborative defense system.
[0171] In this embodiment, S6 specifically includes:
[0172] S61. Based on the cross-domain collaborative blocking strategy, determine the blocking time interval and link blocking direction of key defense nodes in the network attack near-source domain;
[0173] S62. Generate a blocking instruction for the key defense node based on the blocking time interval;
[0174] S63. Send the blocking instruction to the network edge device in the domain close to the network attack source in an encrypted and authenticated manner. The authentication method adopts 128-bit symmetric key encryption, and the key update period is set to 12 hours.
[0175] S64, after the network edge device in the network attack near source domain receives the blocking instruction, at the starting blocking time T s Modify the packet forwarding rules for the specified port to discard incoming or outgoing packets of the specified port;
[0176] S65. The network edge device records the number of data packets blocked by the designated port in real time during the execution of the blocking instruction, with the time resolution of the data packet number statistics set to 1 second.
[0177] S66, blocking duration T d After the end, the network edge device automatically restores the designated port packet forwarding rules to the state before blocking, and feeds back the number of packets during the blocking period to the defense agent of the corresponding network domain.
[0178] During the specific implementation of the present invention, the blocking time interval and link blocking direction of the key defense nodes are determined based on the link load value and blocking priority given in the cross-domain collaborative blocking strategy. The blocking time interval is specifically fixed at 5 seconds for high-priority nodes and 10 seconds for medium-priority nodes. The blocking instruction generation process adopts the standard JSON data format, which clearly includes the starting blocking time, blocking duration, and port identification of the blocked link, wherein the port identification includes the node IP address and port number, and the port number value range is clearly 1024 to 65535. The 128-bit symmetric key encryption authentication method specifically adopts the AES encryption algorithm, and the key update process is performed by the network boundary device in the near-source domain of the network attack and the defense intelligent body using Diffie-H The ellman key exchange protocol is negotiated and generated, and the key update operation is automatically performed every 12 hours; the packet forwarding rule modification is implemented by adding blocking rules to the ACL access control list of the network edge device, and the accuracy of the rule modification is specific to the IP address and port level; the specific statistical method for the number of data packets recorded in real time is: the total amount of data packets discarded by the port of the network edge device is used as the statistical basis, and the statistics are counted every 1 second and recorded in the log file; after the blocking instruction is executed, the network edge device automatically restores the packet forwarding rule by deleting the corresponding rule in the ACL access control list, and feeds back the statistical results of the number of data packets during the blocking period to the defense intelligence of the corresponding network domain in JSON format through a secure communication channel.
[0179] The present invention ensures the real-time and accuracy of blocking operations by clarifying and rigorously defining the blocking instruction generation, transmission, and execution process, effectively reduces the network attack blocking response delay, and enhances the traceability and verifiability of data packet blocking effects, thereby significantly improving the overall performance and application reliability of cross-domain network attack defense strategies.
[0180] In this embodiment, the S62 specifically includes:
[0181] S621, based on the blocking time interval of the key defense node, calculate the starting blocking time T of the blocking instruction s , starting blocking time T s Defined as the next integer second after the end of the current acquisition cycle;
[0182] S622. Determine the blocking duration T of the key defense nodes based on the cross-domain collaborative blocking strategy. d When the blocking priority is high, the blocking duration is T dSet to 5 seconds. When the blocking priority is medium, the blocking duration is T d Set to 10 seconds;
[0183] S623. Obtain the port identifier of the link corresponding to the key defense node. The port identifier consists of the node IP address and the port number. The node IP address uses the IPv4 format, and the port number ranges from 1024 to 65535.
[0184] S624, based on the starting blocking time T s , blocking duration T d and the port identifier to generate a port-level blocking instruction, which includes a start timestamp, an end timestamp, and a specific port number;
[0185] S625: Verify the compliance of the port-level blocking instruction, including the duration of the blocking, T d Whether the time is within the range of 5 to 10 seconds and whether the port number is within the range of 1024 to 65535. After verification, the blocking instruction is confirmed to be valid;
[0186] S626: Store the verified port-level blocking instruction into the blocking instruction queue to be sent, and s They will be issued in the order of time.
[0187] During the specific implementation of the present invention, the starting blocking time specifically adopts the unified Network Time Protocol (NTP) standard in the network domain for time synchronization to ensure millisecond accuracy; the blocking duration is clearly distinguished between a fixed value of 5 seconds when the blocking priority is set to high and a fixed value of 10 seconds when the blocking priority is set to medium, and no other intermediate values are used; the node IP address in the port identifier clearly uses the IPv4 standard address, specifically in dotted decimal format, and the port number is clearly limited to the range of 1024 to 65535; the port-level blocking instruction is specifically stored and transmitted in the form of a JSON data structure. The data is input, including the start timestamp (accurate to milliseconds), the end timestamp (accurate to milliseconds) and the specific port number; the port-level blocking instruction compliance verification clearly stipulates that the blocking duration T_d can only be one of the two values 5 seconds or 10 seconds, and the port number range is strictly checked to be within the range of 1024 to 65535; the queue of blocking instructions to be sent is specifically managed in a first-in-first-out (FIFO) manner, and the blocking instructions are issued in ascending order according to the starting blocking time T_s. The interval between each instruction issuance is fixed at 1 second to avoid performance fluctuations that may be caused by the network device executing multiple instructions at the same time.
[0188] The present invention effectively ensures the accuracy and controllability of the blocking instruction execution process through precise blocking instruction generation and compliance verification mechanisms, reduces the risk of instruction misoperation or execution delay that may occur during the execution of network attack defense, and significantly improves the overall operational stability and security of the cross-domain network security prevention and control system.
[0189] Example 1:
[0190] In order to verify the feasibility of the present invention, the present invention is applied in a large cross-domain network environment of a provincial energy and power dispatching center. Since the center is responsible for the important task of power transmission and dispatching in the entire province, the amount of business data is huge and the network topology is complex, so network security protection has become a core requirement for daily operations. Recently, the network monitoring department of the center discovered during routine network inspections that there have been many suspected continuous network attack incidents targeting important business nodes in the power production dispatching network, which manifested as continuous high-frequency data requests, abnormal traffic fluctuations and abnormal increases in link loads, resulting in significant delays in normal business. Traditional protection measures are unable to effectively prevent the spread of attack incidents in a timely manner due to the lack of cross-domain collaborative response capabilities and attack source precise positioning mechanisms, which seriously affects the normal operation and stability of the business system.
[0191] In order to meet the above challenges, the energy and power dispatching center decided to adopt the layered and domain-based network attack near-source blocking method proposed by the present invention based on graph neural network fusion multi-agent deep reinforcement learning. The present invention first deployed multiple sets of traffic collection equipment at the boundaries of each major network domain of the center, and synchronously collected network traffic data of all devices in the domain every 10 seconds, including source IP address, destination IP address, packet size, transmission protocol type and timestamp information accurate to milliseconds, and combined with the routing equipment in the domain to obtain the link connectivity status, node device IP address, port connection relationship and link delay data in real time through the SNMP protocol, thereby constructing an accurate initial network topology structure. Through strict data synchronization and standardization processing, the center obtained network traffic data of a unified scale, providing a reliable foundation for subsequent attack analysis.
[0192] After preprocessing the network topology data, the center used the deployed graph neural network to dynamically extract and analyze the topology structure and traffic data. The specific process included using the multi-layer graph convolution operation and attention mechanism of the GraphSAGE algorithm to generate node-level dynamic embedding vectors. It then used density clustering analysis to identify abnormal traffic characteristics and accurately locate the network attack topology path. During the implementation period, the center's network monitoring platform clearly detected a remote measurement and control terminal (RTU) node located in the production domain exhibiting significant abnormal characteristics. The node's UDP protocol packet ratio exceeded the normal level for a short period of time, reaching over 78%, with a packet size variance of more than 1,800 bytes and an average data transmission interval of only 23 milliseconds, significantly lower than normal. These abnormal indicators clearly indicated that the node was under a malicious traffic attack.
[0193] The system then identifies the proximate source of the network attack based on abnormal node characteristics and attack path analysis, and immediately activates defense agents deployed within each network domain. These agents rapidly generate initial defense blocking strategies based on real-time collected link load values and blocking priorities. For critical link nodes with link load values greater than 10,000 bytes / second, the blocking priority is immediately set to high, with a blocking time of 5 seconds. For links with load values between 5,000 and 10,000 bytes / second, the blocking priority is set to medium, with a blocking time of 10 seconds. This process ensures efficient and targeted strategies.
[0194] After initially determining the defense strategy, the center utilized a multi-agent deep reinforcement learning mechanism to optimize the cross-domain collaborative strategy. This system evaluated the effectiveness of each defense agent's strategy in real time, dynamically adjusted the probability of strategy selection, and conducted more than ten iterations of strategy optimization. Ultimately, the center generated a cross-domain collaborative blocking strategy, which was then securely and accurately distributed to edge devices near the source of the network attack using AES-128 symmetric key encryption and authentication. The network edge devices precisely modified the port packet forwarding rules at the start of the blocking phase, rapidly blocking the malicious traffic and curbing the further spread of the attack.
[0195] Within three months of implementation, the center significantly improved its network security status, and related security indicators improved significantly. The specific data are shown in Table 1:
[0196] Table 1: Network security protection implementation effect of energy and power dispatching center
[0197]
[0198] As shown in Table 1, after implementing the method of the present invention, the number of network attack incidents in the center dropped from an average of 22 per month before implementation to 2 per month, and the accuracy of accurately locating the source of network attacks increased significantly from 79% to 98.5%. After implementing the present invention, the response speed of network attack incidents dropped significantly from the previous average response time of 17 minutes to less than 12 seconds. The duration of attack incidents also dropped from an average of 31 minutes before implementation to less than 2 minutes. The number of diffusion nodes of a single attack incident was significantly reduced, and the scope of impact of security incidents was significantly reduced. After implementation, the average number of affected nodes for a single incident was only 1 node, which was a reduction of more than 90% compared to before. In addition, because the defense mechanism is more efficient and intelligent, the input of manual labor hours is significantly reduced, and the monthly manual labor hours used for security incident processing have dropped from 89 hours to 16 hours.
[0199] The detailed implementation data above demonstrates that the method presented here fully addresses the challenges of existing network security technologies, such as the difficulty in accurately locating attack sources, slow blocking responses, and low cross-domain collaboration efficiency. By integrating graph neural networks with multi-agent deep reinforcement learning, the Energy and Power Dispatching Center has significantly improved the accuracy and real-time nature of network security defenses, significantly reducing the frequency of attacks and the negative impact on business operations, and providing effective security for the province's power dispatching network.
[0200] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A network attack near-source blocking method based on a layered and domain-specific security control protocol, characterized in that: The steps include: S1. Collect network topology information and traffic data in a cross-domain network environment in real time and generate a network topology diagram; S2. Using a graph neural network to dynamically extract features from nodes and edge relationships in the network topology graph to obtain node-level abnormal traffic features and topological paths of network attacks; S3. Determine the near-source location of the network attack behavior based on the extracted topological path and abnormal traffic characteristics; S4. After determining the near-source location, multiple defense agents deployed in each network domain are activated. Each defense agent generates an initial defense blocking strategy based on the network attack characteristics and network topology information of the domain. S5. Multiple defense agents use the initial defense blocking strategy as initial input, conduct cross-domain strategy interaction and dynamic game iterative optimization through a multi-agent deep reinforcement learning mechanism, and generate a cross-domain collaborative blocking strategy. S6. Based on the cross-domain collaborative blocking strategy, a blocking instruction is issued to the network attack near-source domain to perform real-time network attack traffic blocking.
2. The method for blocking network attacks near the source of a layered and domain-specific security control protocol according to claim 1, characterized in that: Said S1 specifically includes: S11. Traffic collection equipment deployed at the network domain boundary synchronously collects source IP address, destination IP address, packet size, transmission protocol type, and timestamp data with a collection cycle of 10 to 30 seconds; S12. Using a routing device within the network domain to obtain routing status information in real time, the routing status information includes link connectivity status, node device IP address, port connection relationship, and link delay value; S13. Based on the source IP address, destination IP address, and node device IP address data, and in combination with the port connection relationship, a correspondence between network nodes and links between nodes is generated to form an initial network topology structure; S14, performing time series synchronization processing on the node traffic data in the network topology structure according to the timestamp data to obtain network traffic data under a unified time scale; S15. Filtering out links with delays greater than 100 milliseconds based on the link delay values, removing links with abnormal delays in the initial network topology, and obtaining an updated network topology; S16. Based on the updated network topology structure and combined with the network traffic data, a network topology structure diagram for graph neural network analysis is generated.
3. The method for blocking network attacks near the source of a layered and domain-specific security control protocol according to claim 1, characterized in that: The S2 specifically includes: S21, taking the network topology graph as input, obtaining the initial topological feature vector of the target node by the adjacent node aggregation method; S22, performing a multi-layer graph convolution operation on the initial topological feature vector, wherein each layer of convolution operation takes the features of the current node and adjacent nodes as input, updates the node topological feature vector, and the number of convolution layers is set to 3 to 5 layers; S23, concatenating and fusing the convolved node topology feature vector and the node traffic data to obtain a node feature vector after topology and traffic fusion; S24. Calculate the feature weight coefficient for the fused node feature vector using the attention mechanism, and perform weighted combination of the node feature vectors according to the weight coefficient to obtain a node-level dynamic embedding vector; S25. Based on the node-level dynamic embedding vector, a density clustering method is used to identify abnormal traffic characteristics of the node, and the abnormality judgment threshold is set to an inter-cluster distance greater than 0.8; S26. Based on the identified abnormal traffic features, reversely trace the embedding vector similarity of adjacent nodes starting from the abnormal traffic features. When the cosine similarity value of the embedding vectors between nodes is greater than 0.9, it is determined that there is a network attack topology path between the nodes, and the topology path of the network attack is determined.
4. The method for blocking network attacks near the source of a layered and domain-specific security control protocol according to claim 1, characterized in that: The S25 specifically includes: S251, taking the node-level dynamic embedding vector as input, calculate the Euclidean distance between any two nodes and generate a node distance matrix D, where the matrix element d ij Represents the Euclidean distance between node i and node j; S252, according to the node distance matrix D, the distance values between nodes are arranged in ascending order, and the distance values of the nodes in the top 15% of the distance ranking are selected as the distance determination basis of the local high-density area; S253, calculate the local density value ρ for each node i : Where N is the total number of nodes, d c is the basis for determining the distance to the local high-density area, χ(x) is the cutoff function, when x<0, χ(x)=1, when x≥0, χ(x)=0; S254. Calculate the node density difference δ for each node i , For the node with the highest density, set δ i =max(d ij ); S255, according to the local density value ρ i and density difference δ i , calculate the cluster center metric value γ of the node i ,γ i =ρ i ×δ i , select the cluster center metric value γ i Nodes with a score greater than 0.5 are used as cluster centers; S256. Cluster nodes with the cluster center as the core node and the distance between clusters greater than 0.8 as the abnormality judgment threshold. Nodes that are not included in any cluster or have a distance from the cluster center greater than 0.8 are determined to be nodes with abnormal traffic characteristics.
5. The method for blocking network attacks near the source of a layered and domain-specific security control protocol according to claim 1, characterized in that: The S3 specifically includes: S31. Based on the network attack topology path, select the node on the topology path that directly corresponds to the abnormal traffic characteristics as the node to be located; S32, obtaining the data packet size, protocol type, and transmission timestamp data of the node to be located within three consecutive acquisition cycles; S33, calculate the variance value of the data packet size of the node to be located in three consecutive acquisition cycles And set A packet size greater than 1500 bytes is considered abnormal; S34, for the protocol type data of the node to be located, counting the ratio of packets of the Transmission Control Protocol (TCP) and the User Datagram Protocol (UDP), and determining that the protocol type is abnormal when the ratio of UDP packets exceeds 60%; S35, based on the timestamp data transmitted by the node to be located, calculate the average value t of the data transmission interval of the node in three consecutive acquisition cycles avg , when t avg When it is less than 50 milliseconds, the node transmission frequency is determined to be abnormal; S36. When the node to be located meets the three conditions of abnormal data packet size, abnormal protocol type, and abnormal node transmission frequency at the same time, the node to be located is determined to be the near source location of the network attack behavior.
6. The method for blocking network attacks near the source of a layered and domain-specific security control protocol according to claim 1, characterized in that: The S4 specifically includes: S41. After determining the near-source location of the network attack behavior, activate the defense agent deployed in each network domain to obtain the source IP address, destination IP address, and number of data packets of the near-source location node in the network domain; S42. Based on the number of data packets in the last three collection cycles, count the number of data packets sent by the nodes near the source location in the network domain to different destination IP addresses, and mark the destination IP addresses for which the number of data packets from a single destination IP address accounts for more than 30% as high-frequency destination IP addresses; S43. Based on the high-frequency target IP address, determine the link load value of the node connected to the high-frequency target IP address in the current network domain; S44. Identify links within the network domain with a link load value greater than 5000 bytes / second, and mark nodes associated with the links as key defense nodes; S45. Determine a node blocking priority based on a link load value of the key defense node. When the link load value is greater than 10,000 bytes / second, the blocking priority is set to high. When the link load value is between 5,000 bytes / second and 10,000 bytes / second, the blocking priority is set to medium. S46. Each defense agent generates an initial defense blocking strategy based on the blocking priority and link load value of the key defense nodes in the network domain. The initial defense blocking strategy includes the node blocking time interval and link blocking direction. The high-priority node blocking time interval is set to 5 seconds, and the medium-priority node blocking time interval is set to 10 seconds.
7. The method for blocking network attacks near the source of a layered and domain-specific security control protocol according to claim 1, characterized in that: The S5 specifically includes: S51. Take the initial defense blocking strategy of each defense agent as the initial strategy set, and define the strategy space of each defense agent as A. Each strategy in the strategy space A is denoted as a i ; S52: Execute strategy interaction based on the initial strategy set. Each defense agent calculates the reward value of its own strategy in the current state according to the strategy state of other network domain defense agents. S53. Each defensive agent updates its own strategy value function Q based on the reward value. i (s t ,a i ): Where η represents the learning rate, ranging from 0.01 to 0.1, and γ represents the policy value discount factor, ranging from 0.8 to 0.95; S54, adjusting the strategy selection probability in the strategy space A of each defense agent based on the updated strategy value function; S55, executing the strategy selection probability update of the defense agent, with the number of iterations set to 10 to 20 times; S56. After iterative optimization, the strategy with the highest probability of selection of the defense agent strategy of each network domain is selected to form a cross-domain collaborative blocking strategy.
8. The method for blocking network attacks near the source of a layered and domain-specific security control protocol according to claim 1, characterized in that: The S52 specifically includes: S521, record the current state of the defense agent as s t , current state s t It is defined as the set of link load values and blocking priorities of key defense nodes within the network domain; S522, obtain the current state s of other network domain defense agents t Strategy a j , strategy a j Including node blocking time interval and link blocking direction; S523, calculate other network domain defense agent execution strategy a j The number of data packets blocked after pj , the statistical time is set to the latest three collection cycles; S524, calculate other network domain defense agent execution strategy a j The resulting link blocking cost C j ; S525, based on the number of data packets blocked N pj and link blocking cost C j , calculate the network domain defense agent to execute its own strategy a i The return value r i (s t ,a i ): r i (s t ,a i )=α×N pi -β×C i ; Among them, α is the reward weight coefficient of the number of packet blocking, ranging from 0.1 to 1.0, β is the reward penalty coefficient of the link blocking cost, ranging from 0.1 to 1.0, N pi and C i Respectively represent the network domain defense agent strategy a i The number of blocked packets and the link blocking cost.
9. The method for blocking network attacks near the source of a layered and domain-specific security control protocol according to claim 1, characterized in that: The S6 specifically includes: S61. Based on the cross-domain collaborative blocking strategy, determine the blocking time interval and link blocking direction of key defense nodes in the near-source domain of the network attack; S62. Generate a blocking instruction for the key defense node based on the blocking time interval; S63. Send the blocking instruction to the network edge device in the domain close to the network attack source in an encrypted and authenticated manner. The authentication method adopts 128-bit symmetric key encryption, and the key update period is set to 12 hours. S64, after the network edge device in the network attack near source domain receives the blocking instruction, at the starting blocking time T s Modify the packet forwarding rules for the specified port to discard incoming or outgoing packets of the specified port; S65. The network edge device records the number of data packets blocked by the designated port in real time during the execution of the blocking instruction, with the time resolution of the data packet number statistics set to 1 second. S66, blocking duration T d After the end, the network edge device automatically restores the designated port packet forwarding rules to the state before blocking, and feeds back the number of packets during the blocking period to the defense agent of the corresponding network domain.
10. The method for blocking network attacks near the source of a layered and domain-specific security control protocol according to claim 1, characterized in that: The S62 specifically includes: S621, based on the blocking time interval of the key defense node, calculate the starting blocking time T of the blocking instruction s , starting blocking time T s Defined as the next integer second after the end of the current acquisition cycle; S622: Determine the duration of blocking of key defense nodes based on the cross-domain collaborative blocking strategy. d When the blocking priority is high, the blocking duration is T d Set to 5 seconds. When the blocking priority is medium, the blocking duration is T d Set to 10 seconds; S623. Obtain the port identifier of the link corresponding to the key defense node. The port identifier consists of the node IP address and the port number. The node IP address uses the IPv4 format, and the port number ranges from 1024 to 65535. S624, based on the starting blocking time T s , blocking duration T d and the port identifier to generate a port-level blocking instruction, which includes a start timestamp, an end timestamp, and a specific port number; S625: Verify the compliance of the port-level blocking instruction, including the duration of the blocking, T d Whether the time is within the range of 5 to 10 seconds and whether the port number is within the range of 1024 to 65535. After verification, the blocking instruction is confirmed to be valid; S626: Store the verified port-level blocking instruction into the blocking instruction queue to be sent, and s They will be issued in the order of time.
Citation Information
Cited By
Network security defense method and device for automatic fire alarm system
CN120675820A
Network security defense method and device of automatic fire alarm system
CN120675820B
High-speed rail system network risk dynamic intelligent positioning method based on cloud edge collaboration
CN120750637A
Adaptive network defense and topology reconstruction system based on attack feature learning
CN121396603A