A network attack defense method, device, equipment and storage medium

By deploying honeypot nodes in the network, extracting graph structure features and using graph neural network for malicious detection, the problem of inaccurate detection of honeypot technology in network defense is solved, and accurate detection and rapid response to network attacks are achieved.

CN120151116BActive Publication Date: 2025-08-08CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510623992.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-08
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

The existing honeypot technology cannot accurately detect complex, changeable and hidden cyber attacks, resulting in poor network defense effects.

Method used

By deploying multiple honeypot nodes in the target network, collecting data to be detected and extracting graph structure features, including node features, edge features and adjacency relationship features, using graph neural network for malicious detection, identifying the attacking party and activate the defense mechanism.

Benefits of technology

It realizes accurate detection of network attacks, improves attack response speed and defense efficiency, and realizes real-time response and automated defense.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151116B_ABST
    Figure CN120151116B_ABST
Patent Text Reader

Abstract

The present application relates to the field of network security technology, and in particular to a network attack defense method, apparatus, device and storage medium, which are used to solve the problem of poor network attack response and defense in related technologies. The method is: collecting data to be detected through multiple honeypot nodes deployed in a target network, a single honeypot node is a protection device deployed for a business server or business equipment in the target network, and a single data to be detected includes behavior data of an interacting party interacting with the corresponding honeypot node; based on the graph structure features of multiple honeypot nodes obtained by feature extraction of each data to be detected, malicious detection is performed on the behavior of the interacting party interacting with the multiple honeypot nodes, thereby effectively improving the accuracy of network attack detection. When maliciousness is detected, the first interacting party interacting with the honeypot node is determined to be the attacker, and a defense mechanism is activated to block the attacker's network attack, effectively improving the attack response speed and defense efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and in particular to a method, apparatus, device, and storage medium for defending against network attacks. Background Art

[0002] With the continuous development of network technology, the Internet plays a vital role in various fields, but this also brings with it a variety of new cyber attacks. Traditional static, passive network security defense technologies, such as firewalls and intrusion detection systems, are increasingly unable to meet today's network security needs, and defenders are often in a passive position in network attacks and defenses.

[0003] Honeypot technology is a technique used to protect networks by deploying honeypots, or decoys, within real business networks to attract attackers. However, honeypots themselves lack the ability to deeply analyze attack data. This results in networks using honeypot technology for network defense being unable to accurately detect complex, diverse, and covert cyberattacks, hindering the ability to accurately defend against cyberattacks within business networks. Summary of the Invention

[0004] The embodiments of the present application provide a network attack defense method, apparatus, device, and storage medium to improve attack detection accuracy and achieve real-time response and automated defense.

[0005] The specific technical solutions provided in the embodiments of this application are as follows:

[0006] In a first aspect, an embodiment of the present application provides a method for defending against network attacks, including:

[0007] Acquire data to be detected collected by multiple honeypot nodes deployed in a target network, wherein a single honeypot node is a protection device deployed for a service server or service device in the target network, and the single data to be detected includes behavioral data of an interacting party interacting with the corresponding honeypot node;

[0008] Extracting graph structure features of the plurality of honeypot nodes based on the collected data to be detected, wherein the graph structure features include node features of a single honeypot node, and edge features and adjacency relationship features between honeypot nodes;

[0009] Based on the graph structure features, malicious detection is performed on the behavior of the interacting parties interacting with the multiple honeypot nodes;

[0010] If it is detected that the behavior of the first interacting party interacting with any honeypot node is malicious, the first interacting party is determined to be an attacker, and a defense mechanism is activated to block the attacker from performing a network attack on the target network.

[0011] In a possible implementation, extracting graph structure features of the plurality of honeypot nodes based on the collected data to be detected includes:

[0012] Performing feature extraction on each of the data to be detected to obtain node features of the multiple honeypot nodes, wherein the node features are used to characterize the location and attack behavior information of each honeypot node;

[0013] Based on the node features and the data to be detected, feature analysis is performed on the honeypot nodes with interactive behaviors among the multiple honeypot nodes to obtain edge features of the multiple honeypot nodes, where the edge features are used to characterize the interactive behavior information between the honeypot nodes;

[0014] Based on the number of the multiple honeypot nodes and the edge features, the connection relationship between every two honeypot nodes in the multiple honeypot nodes is recorded to obtain adjacency relationship features of the multiple honeypot nodes.

[0015] In a possible implementation, the node features include some or all of the following features: location dimension features, time dimension features, attack type dimension features, attack frequency dimension features, scanning and detection behavior dimension features, and attempted attack method dimension features;

[0016] The edge features include some or all of the following features: time dimension features, access number dimension features, access method dimension features, attack type dimension features, scanning and detection behavior dimension features, and attempted attack method dimension features.

[0017] In a possible implementation, the detecting malicious behavior of the interacting parties interacting with the multiple honeypot nodes based on the graph structure features includes:

[0018] Normalizing the adjacency relationship features to obtain target adjacency relationship features;

[0019] According to the target adjacency relationship feature and the edge feature, feature aggregation is performed on the node features of the multiple honeypot nodes at least once, and target features corresponding to the multiple honeypot nodes after feature aggregation are classified and predicted to obtain predicted detection values corresponding to the multiple honeypot nodes;

[0020] If any predicted detection value is greater than the detection threshold, it is determined that the behavior of the interacting party interacting with the honeypot node corresponding to any predicted detection value is malicious;

[0021] If any predicted detection value is less than or equal to the detection threshold, it is determined that the behavior of the interacting party interacting with the honeypot node corresponding to the predicted detection value is not malicious.

[0022] In a possible implementation, the adjacency relationship feature, the edge feature, and the node feature are all represented in matrix form;

[0023] The step of performing at least one feature aggregation on the node features of the multiple honeypot nodes based on the target adjacency relationship features and the edge features, and performing classification prediction on the target features corresponding to the multiple honeypot nodes after feature aggregation to obtain predicted detection values corresponding to the multiple honeypot nodes includes:

[0024] Performing element-wise multiplication on the matrix corresponding to the target adjacency relationship feature and the matrix corresponding to the edge feature to obtain an intermediate feature matrix;

[0025] According to the intermediate feature matrix and the matrix corresponding to the node features of the multiple honeypot nodes, a graph convolution operation is performed on the node features of the multiple honeypot nodes using a weight matrix corresponding to a preset number of graph convolution layers to obtain a target feature matrix representing the target features corresponding to each of the multiple honeypot nodes after feature aggregation, wherein the preset number of layers is an integer greater than or equal to 1;

[0026] Based on the target feature matrix, target features corresponding to the multiple honeypot nodes are classified and predicted respectively to obtain predicted detection values corresponding to the multiple honeypot nodes.

[0027] In a possible implementation, after performing malicious detection on the behavior of the interacting parties interacting with the multiple honeypot nodes based on the graph structure features, the method further includes:

[0028] If it is detected that the behavior of the first interacting party interacting with any honeypot node is malicious, an attack report of the first interacting party is generated, and the attack report and the corresponding data to be detected are stored as new samples;

[0029] If it is detected that the behavior of the second interacting party interacting with another honeypot node is not malicious, an interaction report of the second interacting party is generated, and the interaction report after malicious identification by relevant personnel and the corresponding data to be detected are stored as new samples;

[0030] After receiving an optimization instruction for the weight matrix of the preset number of graph convolution layers, the weight matrix of the preset number of graph convolution layers is optimized based on the stored new samples, and the weight matrices of the original graph convolution layers are replaced by the optimized weight matrices.

[0031] In a possible implementation, the interaction types of the multiple honeypot nodes include part or all of a low interaction type, a medium interaction type, and a high interaction type;

[0032] The behavioral data collected by a single honeypot node includes part or all of the network traffic data, attack behavior data, system log data and interaction data; wherein, the attack behavior data is obtained by the honeypot node by identifying the behavior of the interacting party based on the network traffic data, the system log data and the interaction data, including the attack type, scanning and detection behavior data, attempted attack method data, and part or all of the time and frequency.

[0033] In a second aspect, an embodiment of the present application provides a network attack defense device, comprising:

[0034] a data acquisition unit, configured to acquire data to be detected collected by multiple honeypot nodes deployed in a target network, wherein a single honeypot node is a protection device deployed for a service server or service device in the target network, and a single data to be detected includes behavioral data of an interacting party interacting with the corresponding honeypot node;

[0035] A feature extraction unit is used to extract graph structure features of the plurality of honeypot nodes based on the collected data to be detected, wherein the graph structure features include node features of a single honeypot node, and edge features and adjacency relationship features between honeypot nodes;

[0036] a malicious detection unit, configured to perform malicious detection on behaviors of interacting parties interacting with the plurality of honeypot nodes based on the graph structure features;

[0037] The response and defense unit is used to determine that the first interacting party interacting with any honeypot node is an attacker if malicious behavior is detected, and to activate a defense mechanism to block the attacker from conducting a network attack on the target network.

[0038] In a possible implementation, the feature extraction unit is specifically configured to:

[0039] Performing feature extraction on each of the data to be detected to obtain node features of the multiple honeypot nodes, wherein the node features are used to characterize the location and attack behavior information of each honeypot node;

[0040] Based on the node features and the data to be detected, feature analysis is performed on the honeypot nodes with interactive behaviors among the multiple honeypot nodes to obtain edge features of the multiple honeypot nodes, where the edge features are used to characterize the interactive behavior information between the honeypot nodes;

[0041] Based on the number of the multiple honeypot nodes and the edge features, the connection relationship between every two honeypot nodes in the multiple honeypot nodes is recorded to obtain adjacency relationship features of the multiple honeypot nodes.

[0042] In a possible implementation, the node features include some or all of the following features: location dimension features, time dimension features, attack type dimension features, attack frequency dimension features, scanning and detection behavior dimension features, and attempted attack method dimension features;

[0043] The edge features include some or all of the following features: time dimension features, access number dimension features, access method dimension features, attack type dimension features, scanning and detection behavior dimension features, and attempted attack method dimension features.

[0044] In a possible implementation, the malicious detection unit is specifically configured to:

[0045] Normalizing the adjacency relationship features to obtain target adjacency relationship features;

[0046] According to the target adjacency relationship feature and the edge feature, feature aggregation is performed on the node features of the multiple honeypot nodes at least once, and target features corresponding to the multiple honeypot nodes after feature aggregation are classified and predicted to obtain predicted detection values corresponding to the multiple honeypot nodes;

[0047] If any predicted detection value is greater than the detection threshold, it is determined that the behavior of the interacting party interacting with the honeypot node corresponding to any predicted detection value is malicious;

[0048] If any predicted detection value is less than or equal to the detection threshold, it is determined that the behavior of the interacting party interacting with the honeypot node corresponding to the predicted detection value is not malicious.

[0049] In a possible implementation, the adjacency relationship feature, the edge feature, and the node feature are all represented in matrix form;

[0050] The malicious detection unit is specifically used to:

[0051] Performing element-wise multiplication on the matrix corresponding to the target adjacency relationship feature and the matrix corresponding to the edge feature to obtain an intermediate feature matrix;

[0052] According to the intermediate feature matrix and the matrix corresponding to the node features of the multiple honeypot nodes, a graph convolution operation is performed on the node features of the multiple honeypot nodes using a weight matrix corresponding to a preset number of graph convolution layers to obtain a target feature matrix representing the target features corresponding to each of the multiple honeypot nodes after feature aggregation, wherein the preset number of layers is an integer greater than or equal to 1;

[0053] Based on the target feature matrix, target features corresponding to the multiple honeypot nodes are classified and predicted respectively to obtain predicted detection values corresponding to the multiple honeypot nodes.

[0054] In a possible implementation, after performing malicious detection on the behavior of the interacting parties interacting with the multiple honeypot nodes based on the graph structure features, the response and defense unit is further configured to:

[0055] If it is detected that the behavior of the first interacting party interacting with any honeypot node is malicious, an attack report of the first interacting party is generated, and the attack report and the corresponding data to be detected are stored as new samples;

[0056] If it is detected that the behavior of the second interacting party interacting with another honeypot node is not malicious, an interaction report of the second interacting party is generated, and the interaction report after malicious identification by relevant personnel and the corresponding data to be detected are stored as new samples;

[0057] After receiving an optimization instruction for the weight matrix of the preset number of graph convolution layers, the weight matrix of the preset number of graph convolution layers is optimized based on the stored new samples, and the weight matrices of the original graph convolution layers are replaced by the optimized weight matrices.

[0058] In a possible implementation, the interaction types of the multiple honeypot nodes include part or all of a low interaction type, a medium interaction type, and a high interaction type;

[0059] The behavioral data collected by a single honeypot node includes part or all of the network traffic data, attack behavior data, system log data and interaction data; wherein, the attack behavior data is obtained by the honeypot node by identifying the behavior of the interacting party based on the network traffic data, the system log data and the interaction data, including the attack type, scanning and detection behavior data, attempted attack method data, and part or all of the time and frequency.

[0060] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0061] Memory, used to store computer programs or instructions;

[0062] A processor is used to execute the computer program or instructions in the memory so that any method described in the first aspect above is executed.

[0063] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which, when instructions in the storage medium are executed by a processor, enables the processor to execute any method described in the first aspect above.

[0064] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising: a computer program code, which, when executed on a computer, enables the computer to execute any of the methods described in the first aspect above.

[0065] In an embodiment of the present application, based on the graph structure features of multiple honeypot nodes extracted from the data to be detected collected by multiple honeypot nodes deployed in the target network, malicious detection is performed on the behavior of the interacting party interacting with the multiple honeypot nodes; if it is detected that the behavior of the first interacting party interacting with any honeypot node is malicious, the first interacting party is determined to be the attacker, and a defense mechanism is activated to block the attacker from conducting a network attack on the target network; wherein, a single honeypot node is a protection device deployed for a business server or business equipment in the target network, and a single data to be detected includes the behavior data of the interacting party interacting with the corresponding honeypot node, and the graph structure The structural features include the node features of a single honeypot node, as well as the edge features and adjacency relationship features between honeypot nodes. In this way, by deploying honeypots to collect the data to be detected in the target network in real time and extracting features from the data to be detected, the graph structural features can be obtained, and then malicious detection is performed based on the graph structural features. That is, by making a comprehensive judgment based on the connection relationship between nodes and the additional features of the edges, abnormal nodes, abnormal behaviors and potential attack modes in the network can be accurately detected, thereby improving the accuracy of network attack detection, thereby being able to trigger the defense mechanism in real time, realize real-time response and automated defense, and improve attack response speed and defense efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 This is a schematic diagram of the architecture of an optional network attack detection and defense system in an embodiment of the present application;

[0067] Figure 2 A flowchart of a method for defending against network attacks according to an embodiment of the present application is shown;

[0068] Figure 3 A schematic diagram of the composition of a graph structure feature of multiple honeypot nodes in an embodiment of the present application;

[0069] Figure 4 This is a schematic diagram of a process for extracting graph structure features in an embodiment of the present application;

[0070] Figure 5 This is a schematic diagram of a malicious detection process in an embodiment of the present application;

[0071] Figure 6 This is a schematic diagram of a specific malicious detection process in an embodiment of the present application;

[0072] Figure 7Schematic diagram of a process for further identifying maliciousness in non-malicious data to be detected according to an embodiment of the present application;

[0073] Figure 8 A flowchart of a real-time data collection and model optimization method in an embodiment of the present application;

[0074] Figure 9 This is a schematic diagram of the logical architecture of a network attack defense device according to an embodiment of the present application;

[0075] Figure 10 Schematic diagram of the physical structure of the electronic device in the embodiment of the present application. DETAILED DESCRIPTION

[0076] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0077] It should be noted that the term "and / or" in the specification, claims, and drawings of this application describes the association relationship between associated objects, indicating that three possible relationships exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0078] The terms "first," "second," "third," and the like in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequential sequence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the application described herein can be practiced in sequences other than those illustrated or described herein.

[0079] A honeypot is a security resource whose value lies in being scanned, attacked, and compromised. A honeypot does not provide any services to external users. All network traffic entering and leaving the honeypot is illegal and could indicate a scan or attack. The core value of a honeypot lies in monitoring, detecting, and analyzing these illegal activities.

[0080] With the continuous development of network technology, new network attacks, complex, changeable and covert network attacks continue to emerge. However, current honeypots can only passively collect attack data and cannot accurately detect the above-mentioned network attacks, and thus cannot achieve precise defense against network attacks in business networks.

[0081] In view of this, in order to improve the accuracy of network attack detection and achieve real-time response to automated defense, an embodiment of the present application provides a method for defending against network attacks. In the embodiment of the present application, data to be detected collected by multiple honeypot nodes deployed in the target network are obtained, wherein a single honeypot node is a protection device deployed for a business server or business equipment in the target network, and a single data to be detected includes behavior data of an interacting party interacting with the corresponding honeypot node; then, based on the collected data to be detected, graph structure features of multiple honeypot nodes are extracted, wherein the graph structure features include node features of a single honeypot node, as well as edge features and adjacency relationship features between honeypot nodes; based on the graph structure features, malicious detection is performed on the behavior of the interacting party interacting with multiple honeypot nodes; in this way, if it is detected that the behavior of the first interacting party interacting with any honeypot node is malicious, the first interacting party is determined to be an attacker, and a defense mechanism is activated to block the attacker from conducting a network attack on the target network.

[0082] In this application, multiple honeypot nodes deployed in the target network collect the data to be detected in the target network in real time. Then, by performing feature extraction on the data to be detected, the graph structure features of these honeypot nodes can be extracted, and malicious detection can be performed based on the graph structure features. The malicious interacting parties can be accurately detected, and the defense mechanism can be activated to make timely responses and active defenses to the interacting parties, effectively improving the attack response speed and defense efficiency.

[0083] The present application provides a method for defending against network attacks, which is applicable to electronic devices. The electronic device may be a server, such as a web server or database server, or a network device, such as an Internet of Things (IoT) device, without further limitation.

[0084] In some embodiments, the server can be an independent physical server or a server cluster or distributed system composed of multiple physical servers, which is a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0085] The preferred implementation methods of the present application are further described in detail below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. In addition, the embodiments of the present application and the features in the embodiments can be combined with each other if there is no conflict.

[0086] Figure 1This is a schematic diagram of the architecture of an optional network attack detection and defense system in an embodiment of the present application. The system is deployed in the aforementioned electronic device to achieve attack defense against the target network.

[0087] like Figure 1 As shown, the system includes a honeypot module, a data preprocessing module, a graph neural network (GNN) model module, a response and defense module, and a database module, among which,

[0088] The honeypot module is used to deploy honeypot nodes and collect data to be detected. It is also used to transmit the collected data to be detected to the data preprocessing module. Since the data captured by honeypots is usually offensive, the data to be detected can also be called attack data.

[0089] The data preprocessing module is used to preprocess the attack data collected by the honeypot node, including but not limited to data cleaning, format conversion, and feature extraction, so as to convert the attack data into a format suitable for CNN model input; it is also used to transmit the preprocessed data to the GNN model module;

[0090] The GNN model module is used to detect malicious behaviors of the parties interacting with the honeypot node based on the data transmitted by the data preprocessing module, and obtain detection results (or classification results); it is also used to transmit the detection results to the response and defense module;

[0091] The response and defense module is used to trigger the defense mechanism based on the detection results to block the detected malicious interacting party (attacker or attacker) from conducting network attacks on the target network;

[0092] The database module is used to store the above attack data, graph structure data, model parameters of the GNN model and detection results.

[0093] Based on the above system architecture, see Figure 2 As shown, the embodiment of the present application provides a method for defending against network attacks, and the specific process of the method may include but is not limited to the following steps:

[0094] Step 200: Acquire the data to be detected collected by multiple honeypot nodes deployed in the target network, wherein a single honeypot node is a protection device deployed for a business server or business device in the target network, and the single data to be detected includes the behavior data of the interacting party interacting with the corresponding honeypot node.

[0095] In the embodiment of the present application, the target network is a business network authorized to perform network attack defense, and the business network can provide network services to visitors.

[0096] A honeypot is a system, service, or resource intentionally exposed to a network to attract attackers and mitigate the damage to a business network. In other words, a honeypot can appear to be a server or computer host with one or more vulnerabilities that can be exploited by attackers. They are simply like a default operating system, full of vulnerabilities and the potential for compromise.

[0097] In an embodiment of the present application, the deployed honeypot nodes may include low-interaction type honeypot nodes, medium-interaction type honeypot nodes and high-interaction type honeypot nodes, among which low-interaction type honeypot nodes are usually used to simulate limited services or system functions, and attackers can only perform simple interactions; high-interaction type honeypot nodes are usually used to simulate complete systems or services, and attackers can perform deep interactions; and medium-interaction type honeypot nodes are usually between low interaction and high interaction, simulating partial system functions.

[0098] Honeypots used to protect real systems and usually deployed in real production environments are called production honeypots. In the embodiments of the present application, multiple production honeypot nodes can be pre-deployed in the target network to simulate real services and system environments, and the multiple production honeypot nodes deployed can cover a variety of business scenarios and attract attackers to attack. It should be noted that in the embodiments of the present application, the honeypot nodes can be deployed on nodes or ports in the target network that provide external services, or they can be deployed inside the target network. This application does not make specific restrictions.

[0099] In the embodiment of the present application, each honeypot node deployed in the target network is used to collect data to be detected, and the data to be detected includes the behavioral data of the interacting party interacting with the corresponding honeypot node, and the behavioral data collected by a single honeypot node includes part or all of the network traffic data, attack behavior data, system log data and interaction data.

[0100] Network traffic data may include the Internet Protocol Address (IP address), port information, protocol type used, and part or all of the data packet content of the interacting party (e.g., attacker or accessing party);

[0101] System log data may include some or all of file operation data, process creation data, and user login data;

[0102] Interaction data may include interaction commands between the interacting party (such as the attacker or access party) and the honeypot node;

[0103] Attack behavior data is obtained by the honeypot node based on the network traffic data, system log data, and interaction data it obtains, identifying the behavior of the interacting party. In some embodiments, the attack behavior data may include attack type, scanning and detection behavior data, attempted attack method data, and some or all of the time and frequency.

[0104] As a specific implementation method, the attack type can be obtained by analyzing network traffic data, system log data, and interaction data, such as extracting quintuple information and / or specific fields, and then comparing it with the characteristics of known attack types, such as quintuple feature information and / or specific fields, to determine the specific attack type, such as scanning, vulnerability exploitation, malware propagation, cross-site scripting attack, cross-site request forgery, file inclusion attack, etc. It should be noted that this application does not limit the specific determination method.

[0105] In the embodiment of the present application, the scanning and probing behavior may include the interacting party scanning and probing the honeypot node, such as port scanning. Usually, attackers will use various technical means to collect information and discover vulnerabilities in the target network or system, which is the process of scanning and probing. Among them, network scanning refers to the behavior of attackers actively probing the target network or system to collect key information (such as open ports, running services, operating system versions, etc.), with the purpose of discovering the exposed assets of the scanned target, finding exploitable attack entry points, and preparing for subsequent attacks. The probing stage usually includes the following key steps: reconnaissance and information collection, scanning and vulnerability discovery, attack and permission acquisition, maintenance and backdoor placement, trace removal and concealment, etc.

[0106] As a specific implementation method, the attempted attack method, such as a web attack method, a database attack method, an operating system attack method, etc., can be obtained by analyzing network traffic data and system log data and extracting relevant information therefrom. It should be noted that this application does not limit the specific method for determining the attempted attack method.

[0107] In the embodiment of the present application, when executing step 200, the data to be detected can be collected by each honeypot node deployed in the target network, thereby obtaining the data to be detected collected by multiple honeypot nodes deployed in the target network. As a specific implementation method, after the honeypot node collects the above-mentioned data to be detected, it can be stored in the database module and the data can be transmitted in real time to the data preprocessing module of the system to trigger subsequent processes.

[0108] Step 210: Based on the collected data to be detected, graph structure features of multiple honeypot nodes are extracted, wherein the graph structure features include node features of a single honeypot node, and edge features and adjacency relationship features between honeypot nodes.

[0109] In the embodiment of the present application, the data preprocessing module receives the data to be detected transmitted by each honeypot node; then, step 210 is executed to obtain the graph structure features of the above-mentioned multiple honeypot nodes.

[0110] In some embodiments, see Figure 3 As shown, the graph structure features include the node features of a single honeypot node, as well as the edge features and adjacency features between honeypot nodes; then, refer to Figure 4 As shown, when executing step 210, the data preprocessing module can specifically execute the following process:

[0111] Step 2101: extract features from each piece of data to be detected to obtain node features of multiple honeypot nodes, wherein the node features are used to characterize the location and attack behavior information of each honeypot node.

[0112] In the embodiment of the present application, the node feature characterizes the location of each honeypot node and the collected attack behavior information of the attacker, which may be the access behavior of an external user accessing the target network.

[0113] In some embodiments, the above-mentioned node features may include some or all of the following features: location dimension features, time dimension features, attack type dimension features, attack frequency dimension features, scanning and detection behavior dimension features, and attempted attack method dimension features. In this way, multi-dimensional information is used to describe the data to be detected collected by the honeypot node, that is, the behavioral data of the interacting party, which can more comprehensively describe the attack behavior of the interacting party interacting with the honeypot node, that is, the attacker, so as to facilitate the subsequent extraction of multi-dimensional features for detection and identification, improve detection accuracy, and thus improve real-time response and defense efficiency.

[0114] In a specific implementation, when executing step 2101, the node features can be obtained as follows; it should be noted that the node features may only include one or more of the above features, and the following description will be made using the example of including all the above features:

[0115] 1. Location dimension features, which can include Internet Protocol address features and geographic location features. The Internet Protocol address feature can be obtained by label encoding the IP address of the device corresponding to the honeypot node. It is a numerical feature. Similarly, the geographic location feature can also be obtained by encoding the geographic location of the device through label encoding. It is also a numerical feature.

[0116] 2. Time dimension features, that is, characterizing the attack time. The time in the above attack behavior data can be converted into the time difference relative to a certain reference time point through the timestamp, in minutes;

[0117] 3. Attack type dimension feature, which can be obtained by encoding the attack type in the aforementioned attack behavior data through label encoding, is a numerical feature;

[0118] 4. The attack frequency dimension feature, i.e., the attack frequency, can retain its original value, i.e., the frequency value in the above attack behavior data;

[0119] 5. Scanning and detection behavior dimension features, which can be obtained based on the scanning and detection behavior data in the above-mentioned attack behavior data. In the embodiment of the present application, the scanning and detection behavior dimension features can be Boolean value features, which can be converted to binary (True is 1, False is 0);

[0120] 6. Attack attempt method dimension feature can be obtained based on the attack attempt method data in the above attack behavior data. In the embodiment of the present application, the attack attempt method data can be converted into a numerical feature by label encoding.

[0121] Through the above processing method, the behavioral data of the interacting parties included in each detection data can be converted into numerical features, which is convenient for subsequent detection and classification using the GNN model.

[0122] Step 2102: Based on the node features and each data to be detected, feature analysis is performed on the honeypot nodes with interactive behaviors among the multiple honeypot nodes to obtain edge features of the multiple honeypot nodes, wherein the edge features are used to characterize the interactive behavior information between the honeypot nodes.

[0123] In the embodiment of the present application, the edge feature characterizes the access relationship between two honeypot nodes, which may belong to the access situation within the target network. This is similar to the aforementioned method of obtaining node features, and is achieved by executing step 2102.

[0124] In some embodiments, the above-mentioned edge features include some or all of the following features: time dimension features, access number dimension features, access method dimension features, attack type dimension features, scanning and detection behavior dimension features, and attempted attack method dimension features. In this way, multi-dimensional information is used to describe the access relationship between honeypot nodes, so that when the GNN model is subsequently used for feature aggregation, the model can capture the complex relationship between nodes, enhance the accuracy and efficiency of information dissemination, thereby accurately identifying potential attack behaviors, improving the accuracy of malicious detection, and thereby improving real-time response and defense efficiency.

[0125] In a specific implementation, when executing step 2102, the edge features can be obtained as follows; it should be noted that the edge features may only include one or more of the above features, and the following description will be given using the edge features including the above features as an example:

[0126] 1. Time dimension feature, that is, the time of initiating access. The time in the above attack behavior data can be converted into the time difference relative to the reference time point through the timestamp, in minutes. In the specific implementation, each honeypot node records the behavior data of the visitor, i.e., the interacting party, from the first-person perspective of "I". Then, if there is an access relationship between honeypot node a and honeypot node b, and the data collected by honeypot node b contains the data of honeypot node a accessing honeypot node b, then the time of initiating access between honeypot node a and honeypot node b is the time when honeypot node a initiates access to honeypot node b.

[0127] 2. The access count dimension feature, i.e., the number of visits, can retain its original value, i.e., the frequency value in the above-mentioned attack behavior data;

[0128] 3. Access method dimension features, namely, access methods, can be obtained by analyzing the packet header information of the data packets in the above network traffic data collected by the honeypot node, such as Hypertext Transfer Protocol (HTTP) / Hypertext Transfer Protocol Secure (HTTPS) access, File Transfer Protocol (FTP) access, Secure Shell (SSH) access, etc. In the embodiment of the present application, the access method data can be converted into a numerical feature by label encoding;

[0129] 4. Attack type dimension feature, which can be obtained by encoding the attack type in the aforementioned attack behavior data through label encoding, is a numerical feature;

[0130] 5. Scanning and detection behavior dimension features, which can be obtained based on the scanning and detection behavior data in the above-mentioned attack behavior data. In the embodiment of the present application, the scanning and detection behavior dimension features can be Boolean value features, which can be converted to binary (True is 1, False is 0);

[0131] 6. Attack attempt method dimension feature can be obtained based on the attack attempt method data in the above attack behavior data. In the embodiment of the present application, the attack attempt method data can be converted into a numerical feature by label encoding.

[0132] Through the above processing method, the behavioral data between honeypot nodes included in each detection data can be converted into numerical features.

[0133] Step 2103: Based on the number of the honeypot nodes and the edge features, the connection relationship between every two honeypot nodes in the honeypot nodes is recorded to obtain the adjacency relationship features of the honeypot nodes.

[0134] In the embodiment of the present application, after obtaining the node features and edge features of the plurality of honeypot nodes, step 2103 is executed to record the connection relationship between the plurality of honeypot nodes to obtain the above-mentioned adjacency relationship features.

[0135] In an embodiment of the present application, the above-mentioned node features, edge features and adjacency relationship features can all be represented in matrix form. Then, by executing step 2101, the node features represented in matrix form can be obtained, which is recorded as matrix X, wherein each column element in the matrix X represents a feature of the same dimension, and each row element represents a honeypot node; by executing step 2102, the edge features represented in matrix form can be obtained, which is recorded as matrix E, wherein each column element in the matrix E represents a feature of the same dimension, and each row element represents an edge.

[0136] The adjacency relationship features represented in matrix form, also known as an adjacency matrix and denoted as adjacency matrix A, are used to represent the connection relationships between nodes in a graph structure. The adjacency matrix is a symmetric binary matrix, where indicates the existence of an edge connection between honeypot node i and honeypot node j, and vice versa, that is, the absence of an edge connection between honeypot node i and honeypot node j. Therefore, when executing step 2103, an adjacency matrix can be constructed based on the number of honeypot nodes and the edge feature data obtained above, and filled with corresponding values, i.e., 1 or 0.

[0137] Then, the obtained matrix X, matrix E and adjacency matrix A can be stored.

[0138] Since the numerical value ranges of the elements in the above-mentioned matrices are not uniform, there are differences between different feature dimensions. Therefore, in order to eliminate the differences between different feature dimensions and improve the stability and convergence speed of GNN model training, in an embodiment of the present application, after obtaining the above-mentioned matrix X (node features), matrix E (edge features) and adjacency matrix A, normalization processing is performed.

[0139] As a specific implementation, a normalization process may be used to classify the value of each element in the matrix into the range (0, 1).

[0140] For the convenience of subsequent description, the matrix X, matrix E, and adjacency matrix A after normalization are still represented by X, E, and A. In this way, the above-mentioned graph structure features, or graph structure data, are obtained. In the embodiment of the present application, the obtained matrix A, matrix E, and adjacency matrix A can also be stored in the database module for subsequent GNN model calls.

[0141] Step 220: Based on the graph structure characteristics, malicious detection is performed on the behavior of the interacting party interacting with the multiple honeypot nodes.

[0142] In an embodiment of the present application, when executing step 220, the GNN model module can input the graph structure feature into a pre-trained GNN model, and the GNN model performs malicious detection on the behavior of the interacting parties interacting with multiple honeypot nodes.

[0143] In the embodiments of this application, the GNN model is a type of deep learning model specifically designed for processing graph-structured data. Graph-structured data is typically composed of nodes and edges, where nodes represent entities (i.e., honeypot nodes) and edges represent relationships between entities (i.e., relationships between honeypot nodes). By learning and reasoning about graph-structured data, the GNN model can capture complex relationships between nodes and the global information of the graph.

[0144] In the embodiment of the present application, the GNN model to be trained can be iteratively trained in advance. After the training is completed, a trained GNN model is obtained. Then, the GNN model is deployed on Figure 1 In the network attack detection and defense system shown, a network attack defense method provided by an embodiment of the present application is implemented to achieve network attack defense against the target network.

[0145] The graph structure data used for training the GNN model can be pre-collected by the above-mentioned method of obtaining graph structure features. The model training process enables the GNN model to capture the complex relationship between honeypot nodes by learning and reasoning about the graph structure features, thereby enabling efficient malicious detection and classification, and accurately identifying abnormal nodes, abnormal behaviors and potential attack patterns.

[0146] The following describes the training process of the GNN model in detail, taking the matrix X (node features), matrix E (edge features), and adjacency matrix A (adjacency relationship features) as examples.

[0147] During the model training process, the graph structure features obtained above (collectively referred to as graph structure data in the subsequent training process description) are input into the GNN model to be trained, where the graph structure data includes matrix X, matrix E and adjacency matrix A: the feature vector of each honeypot node in matrix X has a shape of N×F, where N is the number of honeypot nodes (number), and F is the number of feature dimensions of each honeypot node; the features of each edge in matrix E usually represent the interaction behavior information between honeypot nodes; the adjacency matrix A is a matrix representing the connection relationship between honeypot nodes, with a size of N×N.

[0148] Graph Convolution (GCN) is one of the core operations in the GNN model, used to perform feature aggregation and propagation on graph-structured data. The goal of graph convolution is to update the features of the current node by using the features of neighboring nodes, thereby capturing the local structural information of the graph.

[0149] In an embodiment of the present application, the GNN model includes at least one graph convolution layer. The goal of the graph convolution layer is to aggregate the features of each node with the traffic and attack data features of the neighboring nodes through the adjacency matrix, perform linear transformation through the weight matrix, and finally complete the nonlinear transformation through the activation function.

[0150] As a specific implementation method, first, the matrix X, matrix E, and adjacency matrix A are used as input to initialize the feature vector of each honeypot node; then, at least one graph convolution process is performed, where the target feature matrix of the convolution layer of each node obtained after the convolution layer processing can be specifically expressed by the following formula:

[0151]

[0152] in, represents the target feature matrix; Represents the ReLU activation function; is the normalized adjacency matrix, which is obtained using the general normalization method. See the subsequent content for details; Indicates the l The intermediate feature matrix of each node after the convolutional layer processing, initially (i.e. node features); It is l The weight matrix of the graph convolution layer is used to transform the dimension of node features; Represents element-wise multiplication processing.

[0153] As shown in the above formula, in the embodiment of the present application, the adjacency matrix A (adjacency relationship features) is just a binary matrix, which indicates whether there is an edge connection between the honeypot nodes, and the value is 0 or 1. By combining the matrix E (edge features), the representation of each edge and its influence on information dissemination can be enhanced. The matrix E (edge features) acts as the weight of the edge to affect the adjacency matrix to adjust the information transmission intensity.

[0154] The normalized adjacency matrix above is The determination method of is briefly described. In the embodiment of the present application, It can be obtained by the following formula:

[0155]

[0156] in, is the normalized adjacency matrix;

[0157] represents symmetric normalization;

[0158] The adjacency matrix representing a self-connected network can be obtained by the following formula: , indicating that each node also propagates information to itself, A is the original adjacency matrix, I is the identity matrix, which is a square matrix with all 1s on the diagonal and all 0s elsewhere;

[0159] Represents the degree matrix of the adjacency matrix with self-connection, that is The degree matrix of the self-connected degree matrix is a diagonal matrix with only diagonal elements having values. The diagonal elements are the degrees of each node. The self-connected degree matrix can be obtained by the following formula: , i, j are serial numbers.

[0160] above , which adds "self-connection" to each node, that is, each node can not only obtain information from its neighbors, but also retain its own feature information.

[0161] In the embodiment of the present application, the adjacency matrix A is normalized because if the adjacency matrix is used directly to propagate node features, the node features will be weighted and summed in each layer, and the influence of nodes with high degrees will be sharply amplified, resulting in numerical instability or gradient explosion / vanishment; and, nodes with high degrees themselves have many neighbors. If they are not normalized, their information propagation will be stronger, causing unfairness and affecting the generalization ability of the model; therefore, through the above-mentioned normalization processing, feature explosion can be prevented, the deviation caused by the difference in node degrees can be eliminated, and stable training can be achieved. At the same time, the normalized adjacency matrix can be regarded as the "weight distribution" when each node receives information from the neighboring node, so that the aggregation of each node is "average" rather than "sum", achieving average information propagation, and ultimately improving the performance of the model.

[0162] After multiple graph convolution operations, the features of each honeypot node are finally classified through a fully connected layer to obtain the classification results. Usually, there are two types of classification results: normal or abnormal. Among them, normal indicates that the behavior of the interacting party interacting with the honeypot node is not malicious, and abnormal indicates that the behavior of interacting with the honeypot node is malicious.

[0163] Then, the classification results are compared with the label information pre-annotated for each graph structure data used for training to determine the loss value, and the model parameters of the GNN model are adjusted based on the loss value, such as adjusting the weight matrix of the aforementioned graph convolution layers.

[0164] Usually the classification results output by the fully connected layer (i.e., classification layer) in the GNN model are mostly expressed in the form of probability distribution. Then, the predicted detection value output by the GNN model is It can be expressed as follows:

[0165]

[0166] in, is the predicted detection value of each honeypot node; is the weight matrix of the output layer; Indicates the L Honeypot nodes at the output of the graph convolution layer i The target characteristics, L is the total number of graph convolutional layers; i is the serial number of the honeypot node; Represents a normalization function, used to generate a probability distribution.

[0167] In this way, the model is trained through the data collected by the honeypot nodes, the prediction value is gradually adjusted, the accuracy of anomaly detection is continuously improved, and when it is determined that the model convergence is in line with expectations, the model parameters are output to obtain a trained GNN model.

[0168] In the embodiments of this application, the selection of the matrix E (edge features) plays a crucial role. It provides additional contextual information for each edge in the graph. In this way, in the graph neural network (GNN), the graph convolutional layer (GCN) aggregates the features of neighboring nodes based on the adjacency matrix. The advantages of incorporating edge features are explained below from the following aspects:

[0169] 1. The visit count dimension, i.e., the number of visits between nodes, can indicate the strength of the connection. A higher number of visits means frequent interactions between two nodes, which may be a sign of attack. The GNN model adjusts the strength of information aggregation by weighting this edge feature, making the relationship between frequently interacting nodes more prominent.

[0170] 2. Time dimension features, namely the access time dimension features - the time when the access is initiated. The interaction time information between nodes is important for capturing the temporal nature of attacks. Certain attack patterns may occur within specific time periods or time intervals. The access time features of edges help the model learn these temporal patterns.

[0171] 3. Access method dimension features, i.e., access method (e.g., "SSH Brute Force") and attack type dimension features, i.e., attack type (e.g., "DDoS"), as edge features, can help GNN distinguish normal network traffic from malicious attack traffic;

[0172] 4. Attack attempt dimension features and scanning behavior dimension features, that is, information such as attack attempt behavior and scanning behavior, help the model capture the propagation path and expansion pattern of specific attack behavior, thereby identifying the source node and target node of the attack.

[0173] Compared with the scheme in the related art that does not consider edge features, the scheme that does not consider edge features will consider the information of all neighbor nodes equivalent during aggregation, and will ignore the node relationship characteristics, thus failing to accurately detect potential attack behaviors; in the embodiment of the present application, due to the combination of edge features, the edge features are used to provide additional weights or contextual information for each edge, so that when performing feature aggregation on nodes, not only the nodes themselves can be considered, but also the relationship characteristics between nodes, so that the GNN model can accurately identify abnormal nodes, abnormal behaviors and potential attack patterns, thereby improving detection accuracy, and then achieving real-time response and automated defense.

[0174] After briefly describing the training process of the GNN model, the following describes in detail the processing flow of the GNN model, i.e., the GNN model module, during the implementation process. Figure 5 As shown, malicious detection can be achieved by executing the following process:

[0175] Step 2201: normalize the adjacency relationship features to obtain target adjacency relationship features.

[0176] In the embodiment of the present application, when executing step 2201, the GNN model module uses the aforementioned normalization processing method to normalize the input adjacency relationship features, thereby obtaining the target adjacency relationship features. If the aforementioned adjacency relationship features, edge features, and node features are all represented in matrix form, then the target adjacency relationship features are the aforementioned matrix.

[0177] Step 2202: Based on the target adjacency relationship features and edge features, perform at least one feature aggregation on the node features of multiple honeypot nodes, and perform classification prediction on the target features corresponding to the multiple honeypot nodes after feature aggregation to obtain the predicted detection values corresponding to the multiple honeypot nodes.

[0178] In the embodiment of the present application, when executing step 2202, refer to Figure 6 As shown, the GNN model module can specifically execute the following process:

[0179] Step 600: Perform element-wise multiplication on the matrix corresponding to the target adjacency relationship feature and the matrix corresponding to the edge feature to obtain an intermediate feature matrix.

[0180] Step 610: Based on the intermediate feature matrix and the matrix corresponding to the node features of the multiple honeypot nodes, a graph convolution operation is performed on the node features of the multiple honeypot nodes using the weight matrix corresponding to the preset number of graph convolution layers to obtain a target feature matrix that characterizes the target features corresponding to the multiple honeypot nodes after feature aggregation, wherein the preset number of layers is an integer greater than or equal to 1.

[0181] Still referring to the above formula, when the GNN model module executes steps 600 to 610, it can be specifically obtained by the above The formula is calculated.

[0182] Step 620: Based on the target feature matrix, classify and predict the target features corresponding to the multiple honeypot nodes respectively to obtain the predicted detection values corresponding to the multiple honeypot nodes.

[0183] In the embodiments of the present application, the node probability (probability distribution value) output by the classification layer is used as the confidence score for whether the node is abnormal. In some embodiments, the probability of the abnormal class is selected as the node confidence, while in other embodiments, the probability of the normal class is selected as the node confidence.

[0184] Whether using the probability of the normal or abnormal class as the node confidence, the corresponding threshold must be set according to the actual application scenario. Taking the probability of the abnormal class as the node confidence as an example, in actual implementation, asymmetric threshold settings can be used according to the actual application scenario: if you are sensitive to false positives in anomaly detection (misclassifying normal nodes as abnormal), you can set a higher threshold, such as 0.7, to reduce false positives; conversely, if you are more concerned about false negatives (failure to identify truly abnormal nodes), you can set a lower threshold.

[0185] In the embodiment of the present application, when executing step 620, the GNN model module can output the predicted detection value of each honeypot node based on the processing flow of the aforementioned fully connected layer, which is still recorded as . Then, the predicted detection value of each honeypot node is transmitted to the response and defense module.

[0186] After receiving each predicted detection value, the response and defense module compares each predicted detection value with the preset detection threshold, and automatically triggers the defense mechanism based on the comparison result to achieve attack defense against the target network.

[0187] In some embodiments, if the response and defense module determines that any predicted detection value is greater than the detection threshold, it means that there is an abnormality in the honeypot node corresponding to the predicted detection value, then step 2203 is executed; in other embodiments, if it is determined that any predicted detection value is less than or equal to the detection threshold, it means that there is no abnormality in the honeypot node corresponding to the predicted detection value, then step 2204 is executed.

[0188] Step 2203: If any predicted detection value is greater than the detection threshold, it is determined that the behavior of the interacting party interacting with the honeypot node corresponding to the predicted detection value is malicious.

[0189] Step 2204: If any predicted detection value is less than or equal to the detection threshold, it is determined that the behavior of the interacting party interacting with the honeypot node corresponding to the predicted detection value is not malicious.

[0190] Step 230: If it is detected that the behavior of the first interacting party interacting with any honeypot node is malicious, the first interacting party is determined to be an attacker, and a defense mechanism is activated to block the attacker from conducting a network attack on the target network.

[0191] In some embodiments of the present application, if the response and defense module detects that the behavior of the first interacting party interacting with any honeypot node is malicious, step 230 is executed to determine that the first interacting party is an attacker, and a defense mechanism is activated to block the attacker from conducting a network attack on the target network. For example, defense methods such as blocking the IP of the first interacting party and isolating the infected host (i.e., the host where the above-mentioned honeypot node is located) can be used to respond to and defend the target network in real time.

[0192] In some other embodiments of the present application, if the response and defense module detects that the behavior of the second interacting party interacting with another honeypot node is not malicious, it can also send an alarm message to relevant personnel (such as administrators) for further manual malicious identification, analysis and monitoring to improve attack response speed and defense efficiency.

[0193] like Figure 7 As shown, assuming that the response and defense module detects that the behavior of user F interacting with honeypot node a is not malicious, it pushes an alarm message to the administrator, and also provides a detailed interaction report to the administrator, and stores the attack details between user F and honeypot node a (including but not limited to interaction reports and data to be detected) in the database module.

[0194] After receiving the alarm information, the administrator can obtain the relevant data between user F and honeypot node a from the database module, and based on the obtained relevant data, further identify the maliciousness of the interaction information between user F and honeypot node a, thereby marking whether the interaction information between user F and honeypot node a (such as the data to be detected and the interaction report) is attack data, and then mark whether user F is the attacker; and store the interaction report after malicious identification by the administrator in the database module to update the original interaction report.

[0195] In the embodiment of the present application, the response and defense module can also provide detailed reports, such as subsequent attack reports and the aforementioned interactive reports, and store the attack details obtained by analyzing the data to be detected in the database module. Then, in some embodiments of the present application, after executing step 220, refer to Figure 8 As shown in the figure, the response and defense module can also perform the following process to continuously collect real attack data of the target network to prepare for subsequent updates to the GNN model or further research on attack defense:

[0196] Step 800: If it is detected that the behavior of the first interacting party interacting with any honeypot node is malicious, an attack report of the first interacting party is generated, and the attack report and the corresponding data to be detected are stored as new samples;

[0197] Step 810: If it is detected that the behavior of the second interacting party interacting with another honeypot node is not malicious, an interaction report of the second interacting party is generated, and the interaction report after malicious identification by relevant personnel and the corresponding data to be detected are stored as new samples;

[0198] Step 820: After receiving the optimization instruction for the weight matrix of the above-mentioned preset number of graph convolution layers, the weight matrix of the above-mentioned preset number of graph convolution layers is optimized based on the stored new samples, and the weight matrices of the original graph convolution layers are replaced with the optimized weight matrices.

[0199] In an embodiment of the present application, the response and defense module can also store the detection results, that is, whether malicious or not, and the corresponding data to be detected in the database module, that is, execute step 800 or step 810, and then, when the relevant personnel determine that the GNN model needs to be updated, the optimization instruction is triggered. The network attack detection and defense system selects a certain number of new samples from the stored new samples based on the optimization instruction, and optimizes and updates the GNN model, that is, optimizes and updates the weight matrix of each graph convolution layer in the GNN model, and replaces the original weight matrix in the GNN model with the optimized weight matrix to achieve the update of the model.

[0200] It should be noted that the above model optimization process can be performed online or offline, and this application does not make any specific restrictions.

[0201] In the embodiment of the present application, by deploying a honeypot, it is possible to collect the data to be detected in the target network in real time, and perform feature extraction on the data to be detected to obtain graph structure features, wherein the graph structure features include node features, edge features and adjacency relationship features. Then, malicious detection is performed based on the graph structure features, and abnormal nodes, abnormal behaviors and potential attack modes in the network can be accurately detected. By making a comprehensive judgment based on the connection relationship between nodes and the additional features of the edges, the accuracy of anomaly detection is improved, and the occurrence of false positives and missed reports is reduced. Ultimately, based on the detection results, the defense mechanism can be triggered in real time, such as automatically blocking the attack source IP, isolating the infected host, etc. For low-confidence abnormal behaviors, the system will also conduct further analysis and monitoring, thereby improving the attack response speed and defense efficiency.

[0202] Furthermore, in an embodiment of the present application, a method of cleaning, format conversion, and standardization of the data to be detected is adopted to generate a feature matrix (node feature matrix X, edge feature matrix E, adjacency matrix A) suitable for input into a graph neural network (GNN) model. Through feature extraction, the differences between different features are eliminated, stable training data is provided for the GNN model, and the training efficiency and accuracy of the model are improved. Then, GNN is used to conduct deep learning and analysis of network attack behaviors. By learning the graph structure features (node features, edge features, and adjacency relationship features) of the GNN model, the complex relationships between nodes can be captured, thereby accurately identifying potential attack behaviors. At the same time, the edge features (matrix E) play a vital role in the GNN model as auxiliary information, enhancing the accuracy and efficiency of information dissemination.

[0203] Based on the same inventive concept, see Figure 9 As shown, an embodiment of the present application provides a network attack defense device, including:

[0204] A data acquisition unit 910 is configured to acquire data to be detected collected by multiple honeypot nodes deployed in a target network, wherein a single honeypot node is a protection device deployed for a service server or service device in the target network, and a single data to be detected includes behavioral data of an interacting party interacting with the corresponding honeypot node;

[0205] A feature extraction unit 920 is configured to extract graph structure features of the plurality of honeypot nodes based on the collected data to be detected, wherein the graph structure features include node features of a single honeypot node, and edge features and adjacency relationship features between honeypot nodes;

[0206] a malicious detection unit 930, configured to perform malicious detection on behaviors of interacting parties interacting with the plurality of honeypot nodes based on the graph structure features;

[0207] The response and defense unit 940 is used to determine that the first interacting party interacting with any honeypot node is an attacker if malicious behavior is detected, and to activate a defense mechanism to block the attacker from launching a network attack on the target network.

[0208] In a possible implementation, the feature extraction unit 920 is specifically configured to:

[0209] Performing feature extraction on each of the data to be detected to obtain node features of the multiple honeypot nodes, wherein the node features are used to characterize the location and attack behavior information of each honeypot node;

[0210] Based on the node features and the data to be detected, feature analysis is performed on the honeypot nodes with interactive behaviors among the multiple honeypot nodes to obtain edge features of the multiple honeypot nodes, where the edge features are used to characterize the interactive behavior information between the honeypot nodes;

[0211] Based on the number of the multiple honeypot nodes and the edge features, the connection relationship between every two honeypot nodes in the multiple honeypot nodes is recorded to obtain adjacency relationship features of the multiple honeypot nodes.

[0212] In a possible implementation, the node features include some or all of the following features: location dimension features, time dimension features, attack type dimension features, attack frequency dimension features, scanning and detection behavior dimension features, and attempted attack method dimension features;

[0213] The edge features include some or all of the following features: time dimension features, access number dimension features, access method dimension features, attack type dimension features, scanning and detection behavior dimension features, and attempted attack method dimension features.

[0214] In a possible implementation, the malicious detection unit 930 is specifically configured to:

[0215] Normalizing the adjacency relationship features to obtain target adjacency relationship features;

[0216] According to the target adjacency relationship feature and the edge feature, feature aggregation is performed on the node features of the multiple honeypot nodes at least once, and target features corresponding to the multiple honeypot nodes after feature aggregation are classified and predicted to obtain predicted detection values corresponding to the multiple honeypot nodes;

[0217] If any predicted detection value is greater than the detection threshold, it is determined that the behavior of the interacting party interacting with the honeypot node corresponding to any predicted detection value is malicious;

[0218] If any predicted detection value is less than or equal to the detection threshold, it is determined that the behavior of the interacting party interacting with the honeypot node corresponding to the predicted detection value is not malicious.

[0219] In a possible implementation, the adjacency relationship feature, the edge feature, and the node feature are all represented in matrix form;

[0220] The malicious detection unit 930 is specifically configured to:

[0221] Performing element-wise multiplication on the matrix corresponding to the target adjacency relationship feature and the matrix corresponding to the edge feature to obtain an intermediate feature matrix;

[0222] According to the intermediate feature matrix and the matrix corresponding to the node features of the multiple honeypot nodes, a graph convolution operation is performed on the node features of the multiple honeypot nodes using a weight matrix corresponding to a preset number of graph convolution layers to obtain a target feature matrix representing the target features corresponding to each of the multiple honeypot nodes after feature aggregation, wherein the preset number of layers is an integer greater than or equal to 1;

[0223] Based on the target feature matrix, target features corresponding to the multiple honeypot nodes are classified and predicted respectively to obtain predicted detection values corresponding to the multiple honeypot nodes.

[0224] In a possible implementation, after performing malicious detection on the behavior of the interacting parties interacting with the multiple honeypot nodes based on the graph structure features, the response and defense unit 940 is further configured to:

[0225] If it is detected that the behavior of the first interacting party interacting with any honeypot node is malicious, an attack report of the first interacting party is generated, and the attack report and the corresponding data to be detected are stored as new samples;

[0226] If it is detected that the behavior of the second interacting party interacting with another honeypot node is not malicious, an interaction report of the second interacting party is generated, and the interaction report after malicious identification by relevant personnel and the corresponding data to be detected are stored as new samples;

[0227] After receiving an optimization instruction for the weight matrix of the preset number of graph convolution layers, the weight matrix of the preset number of graph convolution layers is optimized based on the stored new samples, and the weight matrices of the original graph convolution layers are replaced by the optimized weight matrices.

[0228] In a possible implementation, the interaction types of the multiple honeypot nodes include part or all of a low interaction type, a medium interaction type, and a high interaction type;

[0229] The behavioral data collected by a single honeypot node includes part or all of the network traffic data, attack behavior data, system log data and interaction data; wherein, the attack behavior data is obtained by the honeypot node by identifying the behavior of the interacting party based on the network traffic data, the system log data and the interaction data, including the attack type, scanning and detection behavior data, attempted attack method data, and part or all of the time and frequency.

[0230] See Figure 10 As shown, an electronic device is provided in an embodiment of the present application, which can implement the functions of the above method, referring to Figure 10 , the electronic device comprises:

[0231] At least one processor 101, and a memory 102 connected to the at least one processor 101. The specific connection medium between the processor 101 and the memory 102 is not limited in the embodiment of the present application. Figure 10 In the example, the processor 101 and the memory 102 are connected via the bus 100. Figure 10 The bus 100 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 10 The diagram is represented by only one thick line, but this does not mean that there is only one bus or one type of bus. Alternatively, the processor 101 may also be referred to as a controller, without limitation to the name.

[0232] In the embodiment of the present application, the memory 102 stores instructions that can be executed by at least one processor 101. The at least one processor 101 can perform the method discussed above by executing the instructions stored in the memory 102. The processor 101 can implement the functions of the aforementioned device.

[0233] In one possible design, processor 101 may include one or more processing units. Processor 101 may integrate an application processor and a modem processor. The application processor primarily processes the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 101. In some embodiments, processor 101 and memory 102 may be implemented on the same chip. In some embodiments, they may also be implemented on separate chips.

[0234] The processor 101 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application may be directly implemented and executed by a hardware processor, or by a combination of hardware and software modules in the processor.

[0235] The memory 102 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 102 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 102 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 102 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.

[0236] By designing and programming the processor 101, the code corresponding to the method described in the aforementioned embodiment can be embedded in the chip, so that the chip can execute the steps of the method described in the aforementioned embodiment when it is running. How to design and program the processor 101 is a technique well known to those skilled in the art and will not be described in detail here.

[0237] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium. When the instructions in the storage medium are executed by a processor, the processor is enabled to execute any one of the methods in the above embodiments.

[0238] In some possible implementations, various aspects of the method provided in the present application may also be implemented in the form of a program product, which includes program code. When the program product is run on a device, the program code is used to enable the device to execute the steps of the method according to various exemplary implementations of the present application described above in this specification.

[0239] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0240] This application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0241] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0242] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0243] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A method for defending against network attacks, characterized in that: include: Acquire data to be detected collected by multiple honeypot nodes deployed in a target network, wherein a single honeypot node is a protection device deployed for a service server or service device in the target network, and the single data to be detected includes behavioral data of an interacting party interacting with the corresponding honeypot node; Extracting graph structure features of the plurality of honeypot nodes based on the collected data to be detected, wherein the graph structure features include node features of a single honeypot node, and edge features and adjacency relationship features between honeypot nodes; Normalizing the adjacency relationship features to obtain target adjacency relationship features; performing at least one graph convolution process on the node features of the multiple honeypot nodes based on the target adjacency relationship features and the edge features, and performing classification prediction on the target features corresponding to the multiple honeypot nodes obtained after the graph convolution process to obtain predicted detection values corresponding to the multiple honeypot nodes; Determining whether the behavior of an interacting party interacting with the multiple honeypot nodes is malicious based on the predicted detection values corresponding to the multiple honeypot nodes; If it is determined that the behavior of the first interacting party interacting with any honeypot node is malicious, the first interacting party is determined to be an attacker, and a defense mechanism is activated to block the attacker from performing a network attack on the target network; The target feature matrix representing the target features corresponding to each of the multiple honeypot nodes is expressed as follows: in, represents the target feature matrix; Represents the ReLU activation function; represents the matrix corresponding to the target adjacency feature; E represents the matrix corresponding to the edge feature; Represents element-level multiplication processing; Indicates the l The feature matrix of each node after layer graph convolution processing, initially , X represents the matrix corresponding to the node features; It is l Layer Weight matrix of the graph convolution layer.

2. The method according to claim 1, wherein The step of extracting the graph structure features of the plurality of honeypot nodes based on the collected data to be detected includes: Performing feature extraction on each of the data to be detected to obtain node features of the multiple honeypot nodes, wherein the node features are used to characterize the location and attack behavior information of each honeypot node; Based on the node features and the data to be detected, feature analysis is performed on the honeypot nodes with interactive behaviors among the multiple honeypot nodes to obtain edge features of the multiple honeypot nodes, where the edge features are used to characterize the interactive behavior information between the honeypot nodes; Based on the number of the multiple honeypot nodes and the edge features, the connection relationship between every two honeypot nodes in the multiple honeypot nodes is recorded to obtain adjacency relationship features of the multiple honeypot nodes.

3. The method according to claim 2, wherein The node features include some or all of the following features: location dimension features, time dimension features, attack type dimension features, attack frequency dimension features, scanning and detection behavior dimension features, and attempted attack method dimension features; The edge features include some or all of the following features: time dimension features, access number dimension features, access method dimension features, attack type dimension features, scanning and detection behavior dimension features, and attempted attack method dimension features.

4. The method according to claim 2, wherein The determining, based on the predicted detection values corresponding to the multiple honeypot nodes, whether the behavior of the interacting party interacting with the multiple honeypot nodes is malicious includes: If any predicted detection value is greater than the detection threshold, it is determined that the behavior of the interacting party interacting with the honeypot node corresponding to any predicted detection value is malicious; If any predicted detection value is less than or equal to the detection threshold, it is determined that the behavior of the interacting party interacting with the honeypot node corresponding to the predicted detection value is not malicious.

5. The method according to claim 4, wherein The adjacency relationship features, the edge features and the node features are all represented in matrix form; The step of performing at least one graph convolution process on the node features of the multiple honeypot nodes according to the target adjacency relationship features and the edge features, and performing classification prediction on the target features corresponding to the multiple honeypot nodes obtained after the graph convolution process to obtain predicted detection values corresponding to the multiple honeypot nodes includes: Performing element-wise multiplication on the matrix corresponding to the target adjacency relationship feature and the matrix corresponding to the edge feature to obtain an intermediate feature matrix; According to the intermediate feature matrix and the matrix corresponding to the node features of the multiple honeypot nodes, a graph convolution operation is performed on the node features of the multiple honeypot nodes using a weight matrix corresponding to a preset number of graph convolution layers to obtain a target feature matrix representing the target features corresponding to each of the multiple honeypot nodes after feature aggregation, wherein the preset number of layers is an integer greater than or equal to 1; Based on the target feature matrix, target features corresponding to the multiple honeypot nodes are classified and predicted respectively to obtain predicted detection values corresponding to the multiple honeypot nodes.

6. The method according to claim 5, wherein After determining whether the behavior of the interacting party interacting with the multiple honeypot nodes is malicious based on the predicted detection values corresponding to the multiple honeypot nodes, the method further includes: If it is determined that the behavior of the first interacting party interacting with any honeypot node is malicious, an attack report of the first interacting party is generated, and the attack report and the corresponding data to be detected are stored as new samples; If it is determined that the behavior of the second interacting party interacting with another honeypot node is not malicious, an interaction report of the second interacting party is generated, and the interaction report after malicious identification by relevant personnel and the corresponding data to be detected are stored as new samples; After receiving an optimization instruction for the weight matrix of the preset number of graph convolution layers, the weight matrix of the preset number of graph convolution layers is optimized based on the stored new samples, and the weight matrices of the original graph convolution layers are replaced by the optimized weight matrices.

7. The method according to any one of claims 1 to 6, wherein: The interaction types of the plurality of honeypot nodes include part or all of a low interaction type, a medium interaction type, and a high interaction type; The behavioral data collected by a single honeypot node includes part or all of the network traffic data, attack behavior data, system log data and interaction data; wherein, the attack behavior data is obtained by the honeypot node by identifying the behavior of the interacting party based on the network traffic data, the system log data and the interaction data, including the attack type, scanning and detection behavior data, attempted attack method data, and part or all of the time and frequency.

8. A network attack defense device, characterized in that: include: a data acquisition unit, configured to acquire data to be detected collected by multiple honeypot nodes deployed in a target network, wherein a single honeypot node is a protection device deployed for a service server or service device in the target network, and a single data to be detected includes behavioral data of an interacting party interacting with the corresponding honeypot node; A feature extraction unit is used to extract graph structure features of the plurality of honeypot nodes based on the collected data to be detected, wherein the graph structure features include node features of a single honeypot node, and edge features and adjacency relationship features between honeypot nodes; a malicious detection unit, configured to normalize the adjacency relationship features to obtain target adjacency relationship features; perform at least one graph convolution process on the node features of the multiple honeypot nodes based on the target adjacency relationship features and the edge features, and perform classification prediction on the target features corresponding to each of the multiple honeypot nodes obtained after the graph convolution process to obtain predicted detection values corresponding to the multiple honeypot nodes; and determine whether the behavior of an interacting party interacting with the multiple honeypot nodes is malicious based on the predicted detection values corresponding to the multiple honeypot nodes; a response and defense unit, configured to, if it is determined that the behavior of a first interacting party interacting with any honeypot node is malicious, determine that the first interacting party is an attacker, and activate a defense mechanism to block the attacker from performing a network attack on the target network; The target feature matrix representing the target features corresponding to each of the multiple honeypot nodes is expressed as follows: in, represents the target feature matrix; Represents the ReLU activation function; represents the matrix corresponding to the target adjacency feature; E represents the matrix corresponding to the edge feature; Represents element-level multiplication processing; Indicates the l The feature matrix of each node after layer graph convolution processing, initially , X represents the matrix corresponding to the node features; It is l Layer Weight matrix of the graph convolution layer.

9. An electronic device, characterized in that: include: Memory, used to store computer programs or instructions; A processor is configured to execute the computer program or instructions in the memory so that the method according to any one of claims 1 to 7 is performed.

10. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor, the processor is enabled to perform the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • GCN-based honey situation map analysis method

    CN119316238A