Network defense method and device
By dividing the power grid system into security domains and global domains and using topology encoders and intelligent agents to optimize defense strategies, the problems of update lag and inefficient resource allocation in traditional network defense architectures are solved, achieving efficient network threat protection.
Patent Information
- Application Number
- CN202510970100.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-15
AI Technical Summary
When faced with complex and ever-changing network attacks, traditional network defense architectures have delayed defense strategy updates and inefficient resource allocation, making it difficult to cope with advanced persistent threats. They also fail to effectively consider the heterogeneity and differentiated security needs of different areas within the network.
By establishing a device topology map of the power grid system equipment, using the topology encoder to extract the characteristics of devices and edges, dividing the security domain and the global domain, and performing dynamic defense optimization based on the attribute characteristics of devices and edges, the defense strategy is adjusted by combining intelligent agents and historical experience rules.
It achieves efficient utilization of domain resources and threat protection capabilities, and improves the dynamic adaptability of network defense and the targeted nature of defense strategies.
Smart Images

Figure CN120498902B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of network security technology, and in particular to a network defense method and device. Background Art
[0002] With the rapid development of Internet technology and the in-depth advancement of digital transformation, network security issues have become increasingly prominent. Traditional network defense architectures mostly adopt static configuration and passive response strategies. When faced with the current complex, changeable and continuously evolving network attacks, they will reveal many shortcomings, such as lagging defense strategy updates, limited environmental adaptability, inefficient resource allocation, and difficulty in responding to advanced persistent threats.
[0003] Furthermore, dynamic defense strategies in related technologies often assume that the entire network is optimized as a single entity, without considering the heterogeneity and differentiated security requirements of different areas within the network. Consequently, the potential value of network topology partitioning for security domain division is overlooked. Summary of the Invention
[0004] In view of the above problems, the present disclosure provides a network defense method and device, which can ensure efficient utilization of domain resources and threat protection capabilities within the domain.
[0005] According to a first aspect of the present disclosure, a network defense method is provided, comprising: determining a device topology graph based on multiple devices in a power grid system, wherein the device topology graph includes devices and edges, and the edges represent the communication relationships between multiple devices; inputting the device topology graph into a topology encoder to obtain device topology features; dividing the multiple devices in the power grid system based on the device topology features, device attribute features of multiple devices in the device topology graph, and edge attribute features of multiple edges to obtain multiple security domains and global domains, wherein the security domain includes at least one device with an intra-domain communication relationship, and the global domain includes multiple security domains with inter-domain communication relationships; and obtaining network defense action information for each device based on the state vectors of each of the multiple security domains and the state vector of the global domain.
[0006] According to an embodiment of the present disclosure, the network defense method also includes: obtaining multiple attribute information sets of different types for each device; using a coding strategy that matches the type of the attribute information set to encode different types of attribute information sets to obtain sub-device attribute characteristics; and obtaining device attribute characteristics based on multiple sub-device attribute characteristics.
[0007] According to an embodiment of the present disclosure, a coding strategy that matches the type of attribute information set is adopted to encode different types of attribute information sets to obtain sub-device attribute characteristics, including: in a case where the type of attribute information set is an equipment indicator type, based on the correlation between the multiple equipment indicators in the attribute information set, determining the weights of the multiple equipment indicators; based on the multiple weights, taking a weighted sum of the indicator values of the multiple equipment indicators to obtain sub-device attribute characteristics of the equipment indicator type; in a case where the type of attribute information set is an equipment location type, based on the pre-divided network hierarchy, determining the network location information of the equipment in the power grid system; encoding the network location information to obtain sub-device attribute characteristics of the equipment location type; in a case where the type of attribute information set is an equipment type, encoding the equipment type information in the attribute information set to obtain sub-device attribute characteristics of the equipment type.
[0008] According to an embodiment of the present disclosure, the network defense method also includes: obtaining a communication attribute information set between two connected devices in a device topology diagram, wherein the communication attribute information set includes communication type information, communication frequency information, and communication importance information; and encoding the communication attribute information set to obtain edge attribute features.
[0009] According to an embodiment of the present disclosure, device attribute features include sub-device attribute features of a device location type, and based on device topology features, device attribute features of multiple devices in a device topology graph, and edge attribute features of multiple edges, multiple devices in a power grid system are divided to obtain multiple security domains and a global domain, including: obtaining device embedding features of multiple devices based on device topology features, device attribute features of multiple devices, and sub-device attribute features of a device location type; performing attention feature extraction and feature update on each device embedding feature in turn to obtain multiple target device features; performing feature fusion on multiple target device features, edge attribute features of multiple edges, and device topology features to obtain a target fusion feature; based on the target fusion feature, determining domain identifiers of multiple devices for representing the security domains to which they belong; and dividing multiple devices in the power grid system based on the domain identifiers of multiple devices to obtain multiple security domains and a global domain.
[0010] According to an embodiment of the present disclosure, network defense action information of each device is obtained based on the state vectors of each of the multiple security domains and the state vector of the global domain, including: inputting the state vector of the security domain into the intra-domain intelligent agent to obtain the intra-domain network defense action information of the security domain; inputting the state vector of the global domain into the inter-domain intelligent agent to obtain the inter-domain network defense action information of the global domain; and obtaining network defense action information based on the intra-domain network defense action information and the inter-domain network defense action information.
[0011] According to an embodiment of the present disclosure, the network defense method further includes: obtaining a state vector of the security domain based on device attribute characteristics, edge attribute characteristics, network configuration status characteristics, and network abnormal event characteristics.
[0012] According to an embodiment of the present disclosure, the network defense method further includes: obtaining a state vector of the global domain based on inter-domain communication characteristics and state vectors of each of the multiple security domains.
[0013] According to an embodiment of the present disclosure, the network defense method further includes: updating the network defense action information based on historical experience rules to obtain target network defense action information.
[0014] A second aspect of the present disclosure provides a network defense device, including: a topology map determination module, used to determine a device topology map based on multiple devices in the power grid system, wherein the device topology map includes devices and edges, and the edges represent the communication relationship between multiple devices; a feature extraction module, used to input the device topology map into a topology encoder to obtain device topology features; a device division module, used to divide multiple devices in the power grid system based on the device topology features, device attribute features of multiple devices in the device topology map, and edge attribute features of multiple edges, to obtain multiple security domains and global domains, wherein the security domain includes at least one device with an intra-domain communication relationship, and the global domain includes multiple security domains with inter-domain communication relationships; and an action determination module, used to obtain network defense action information for each device based on the state vectors of each of the multiple security domains and the state vector of the global domain.
[0015] According to an embodiment of the present disclosure, the network defense device further includes a first information acquisition module, a first information encoding module, and a feature determination module.
[0016] The first information acquisition module is used to acquire multiple attribute information sets of different types for each device.
[0017] The first information encoding module is configured to encode different types of attribute information sets using an encoding strategy that matches the type of the attribute information set to obtain sub-device attribute characteristics.
[0018] The feature determination module is used to obtain device attribute features based on multiple sub-device attribute features.
[0019] According to an embodiment of the present disclosure, the first information encoding module includes a weight determination submodule, a first feature determination submodule, a position determination submodule, a second feature determination submodule, and a third feature determination submodule.
[0020] The weight determination submodule is used to determine the weights of the multiple device indicators based on the correlation between the multiple device indicators in the attribute information set when the type of the attribute information set is a device indicator type.
[0021] The first feature determination submodule is configured to perform weighted summation of the respective index values of the plurality of device indexes based on a plurality of weights to obtain a sub-device attribute feature of the device index type.
[0022] The location determination submodule is used to determine the network location information of the device in the power grid system based on the pre-divided network hierarchy when the type of the attribute information set is the device location type.
[0023] The second feature determination submodule is used to encode the network location information to obtain the sub-device attribute features of the device location type.
[0024] The third feature determination submodule is configured to, when the type of the attribute information set is a device type, encode the device type information in the attribute information set to obtain a sub-device attribute feature of the device type.
[0025] According to an embodiment of the present disclosure, the network defense device further includes a second information acquisition module and a second information encoding module.
[0026] The second information acquisition module is used to obtain a communication attribute information set between two connected devices in the device topology diagram, wherein the communication attribute information set includes communication type information, communication frequency information and communication importance information.
[0027] The second information encoding module is used to encode the communication attribute information set to obtain edge attribute features.
[0028] According to an embodiment of the present disclosure, the device attribute characteristics include sub-device attribute characteristics of the device location type, and the device division module includes an embedded feature determination sub-module, a target feature determination sub-module, a fusion feature determination sub-module, an identification determination sub-module and a device division sub-module.
[0029] The embedding feature determination submodule is used to obtain the device embedding features of the multiple devices based on the device topology features, the device attribute features of the multiple devices and the sub-device attribute features of the device location types.
[0030] The target feature determination submodule is used to extract and update the attention features of each device embedding feature in turn to obtain multiple target device features.
[0031] The fusion feature determination submodule is used to fuse multiple target device features, edge attribute features of multiple edges and device topology features to obtain target fusion features.
[0032] The identification determination submodule is used to determine the domain identification of each of the multiple devices used to characterize the security domain to which they belong based on the target fusion feature.
[0033] The device division submodule is used to divide multiple devices in the power grid system based on their respective domain identifiers to obtain multiple security domains and a global domain.
[0034] According to an embodiment of the present disclosure, the action determination module includes a first action determination submodule, a second action determination submodule, and a third action determination submodule.
[0035] The first action determination submodule is used to input the state vector of the security domain into the intra-domain intelligent agent to obtain the intra-domain network defense action information of the security domain.
[0036] The second action determination submodule is used to input the state vector of the global domain into the inter-domain intelligent agent to obtain the inter-domain network defense action information of the global domain.
[0037] The third action determination submodule is used to obtain network defense action information based on the intra-domain network defense action information and the inter-domain network defense action information.
[0038] According to an embodiment of the present disclosure, the network defense device further includes a first vector determination module.
[0039] The first vector determination module is used to obtain the state vector of the security domain based on device attribute characteristics, edge attribute characteristics, network configuration state characteristics and network abnormal event characteristics.
[0040] According to an embodiment of the present disclosure, the network defense device further includes a second vector determination module.
[0041] The second vector determination module is configured to obtain a state vector of the global domain based on inter-domain communication characteristics and state vectors of each of the plurality of security domains.
[0042] According to an embodiment of the present disclosure, the network defense device further includes an action updating module.
[0043] The action update module is used to update the network defense action information based on historical experience rules to obtain the target network defense action information.
[0044] A third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0045] The fourth aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0046] The fifth aspect of the present disclosure further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.
[0047] According to the embodiments of the present disclosure, by establishing a device topology diagram for the devices in the power grid system, the actual power grid system is abstracted into a data structure for processing. By analyzing the device topology characteristics, device attribute characteristics, and edge attribute characteristics, multiple devices in the power grid system are divided into security domains and global domains based on the above characteristics, realizing a two-level division within the domain and between domains. Defense optimization is performed at both the within-domain and between-domain levels. Each intelligent agent in the domain is independently responsible for optimizing the defense strategy within its own security domain, and dynamically adjusts it according to the security status and defense configuration of the devices in the domain to ensure efficient utilization of resources within the domain and threat protection capabilities within the domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0049] Figure 1 The following schematically illustrates an application scenario of the network defense method and apparatus according to an embodiment of the present disclosure;
[0050] Figure 2 The flowchart of the network defense method according to the embodiment of the present disclosure is schematically shown;
[0051] Figure 3 Schematically shows a data flow diagram of the security domain division process of the network defense method according to an embodiment of the present disclosure;
[0052] Figure 4 A flowchart of a network defense method according to another embodiment of the present disclosure is schematically shown;
[0053] Figure 5 The following schematically shows a structural block diagram of a network defense device according to an embodiment of the present disclosure;
[0054] Figure 6 The block diagram schematically shows an electronic device suitable for implementing the network defense method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0055] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0056] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0057] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0058] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0059] In the technical solutions disclosed herein, the user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0060] In scenarios where personal information is used for automated decision-making, the methods, devices, and systems provided by the embodiments of the present disclosure all provide users with corresponding operation portals for them to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision-making" here refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests and hobbies, or economic, health, credit status, etc. through computer programs and making decisions. The expression "expert decision-making" here refers to the activity of making decisions by people who specialize in a certain field, have specialized experience, knowledge, and skills, and have reached a certain level of professionalism.
[0061] An embodiment of the present disclosure provides a network defense method, including: determining a device topology map based on multiple devices in a power grid system, wherein the device topology map includes devices and edges, and the edges represent the communication relationships between multiple devices; inputting the device topology map into a topology encoder to obtain device topology features; dividing the multiple devices in the power grid system based on the device topology features, device attribute features of multiple devices in the device topology map, and edge attribute features of multiple edges to obtain multiple security domains and global domains, wherein the security domain includes at least one device with an intra-domain communication relationship, and the global domain includes multiple security domains with inter-domain communication relationships; and obtaining network defense action information for each device based on the state vectors of each of the multiple security domains and the state vector of the global domain.
[0062] Figure 1 The following schematically illustrates an application scenario diagram of the network defense method and device according to an embodiment of the present disclosure.
[0063] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.
[0064] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0065] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0066] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0067] It should be noted that the network defense method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the network defense device provided in the embodiment of the present disclosure can generally be set in the server 105. The network defense method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the network defense device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0068] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0069] The following will be based on Figure 1 The scene described by Figures 2 to 4 The network defense method of the disclosed embodiment is described in detail.
[0070] Figure 2 The flowchart of the network defense method according to the embodiment of the present disclosure is schematically shown.
[0071] like Figure 2 As shown, the network defense of this embodiment includes operations S210 to S240.
[0072] In operation S210 , a device topology map is determined based on a plurality of devices in a power grid system.
[0073] In some examples, the devices in the power grid system may be firewall devices, intrusion detection devices, intrusion prevention devices, and common asset devices, etc.
[0074] According to an embodiment of the present disclosure, the device topology graph includes devices and edges. Devices represent power grid devices, and edges represent the communication relationships between multiple devices. By abstracting multiple devices in the power grid system and the communication relationships between multiple devices, a device set V and an edge set E are obtained. Based on the device set V and the edge set E, the device topology graph can be determined. .
[0075] In operation S220 , the device topology map is input into a topology encoder to obtain device topology features.
[0076] According to an embodiment of the present disclosure, the device topology feature may be used to represent a topological relationship between multiple devices in a device topology graph.
[0077] In operation S230 , multiple devices in the power grid system are divided based on device topology characteristics, device attribute characteristics of multiple devices in the device topology graph, and edge attribute characteristics of multiple edges to obtain multiple security domains and global domains.
[0078] According to an embodiment of the present disclosure, a security domain includes at least one device having an intra-domain communication relationship, and a global domain includes multiple security domains having an inter-domain communication relationship.
[0079] According to the embodiments of the present disclosure, based on the device topology characteristics, the device attribute characteristics of multiple devices in the device topology diagram, and the edge attribute characteristics of multiple edges, the communication relationship between the devices corresponding to the multiple devices, the respective attributes and usage status of the multiple devices can be determined, thereby determining the intra-domain communication relationship between the multiple devices. Further, based on the intra-domain communication relationship between the multiple devices, the multiple devices in the power grid system are divided to obtain multiple security domains and global domains.
[0080] In operation S240 , network defense action information of each device is obtained based on the state vectors of each of the plurality of security domains and the state vector of the global domain.
[0081] According to an embodiment of the present disclosure, the state vectors of each of the multiple security domains can be used to represent all state information in the intelligent agent, and to determine the operating status of the multiple devices included in each of the multiple security domains. The state vector of the global domain can be used to determine the status of inter-domain communication between the multiple security domains. Therefore, based on the state vectors of each of the multiple security domains and the state vector of the global domain, the operating status of each of the multiple devices in the current power grid system and the existing operating problems can be determined. Based on the above-mentioned operating status and operating problems, network defense action information for each device can be obtained, wherein the network defense action information can be used to solve problems in the operation of the device by performing operations on the device or changing the configuration of the device.
[0082] According to the embodiments of the present disclosure, by establishing a device topology diagram for the devices in the power grid system, the actual power grid system is abstracted into a data structure for processing. By analyzing the device topology characteristics, device attribute characteristics, and edge attribute characteristics, multiple devices in the power grid system are divided into security domains and global domains based on the above characteristics, realizing a two-level division within the domain and between domains, and performing defense optimization at the two levels within the domain and between domains respectively. Each intelligent agent within the domain is independently responsible for optimizing the defense strategy within its own security domain, and dynamically adjusts it according to the security status and defense configuration of the devices within the domain, ensuring the efficient utilization of resources within the domain and the threat protection capability within the domain.
[0083] According to an embodiment of the present disclosure, the network defense method also includes: obtaining multiple attribute information sets of different types for each device; using a coding strategy that matches the type of the attribute information set to encode different types of attribute information sets to obtain sub-device attribute characteristics; and obtaining device attribute characteristics based on multiple sub-device attribute characteristics.
[0084] According to embodiments of the present disclosure, device attribute information types may include device indicator type, device type, device location type, device value type, workload type, connectivity type, vulnerability type, and device configuration type. According to embodiments of the present disclosure, different attribute information types correspond to different encoding strategies. For each device, multiple attribute information sets of the device are encoded using encoding strategies that match the attribute information set type, thereby obtaining multiple sub-device attribute features of the device.
[0085] For example, for device v, the attribute information sets of its device indicator type, device type, device location type, device value type, workload type, connectivity type, vulnerability type, and device configuration type are obtained respectively. Multiple attribute information sets are encoded using encoding strategies that match the types of attribute information sets, and multiple sub-device attribute features corresponding to the above multiple attribute information sets are obtained. 、 、 、 、 、 、 and .
[0086] According to the embodiment of the present disclosure, after determining multiple sub-device attribute features, the multiple sub-device attribute features are spliced together to obtain the device Device attribute characteristics , as shown in formula (1):
[0087] (1)
[0088] in, , Represents a collection of multiple devices in a power grid system.
[0089] According to the embodiments of the present disclosure, multiple attribute information sets of different device types are encoded separately to obtain sub-device attribute features, which are then combined to form device attribute features. Therefore, device attribute features are derived by comprehensively considering multiple aspects of device attribute information, ensuring the comprehensiveness of device attribute features and further ensuring the effectiveness of subsequent network defenses.
[0090] According to an embodiment of the present disclosure, a coding strategy that matches the type of attribute information set is adopted to encode different types of attribute information sets to obtain sub-device attribute characteristics, including: in a case where the type of attribute information set is an equipment indicator type, based on the correlation between the multiple equipment indicators in the attribute information set, determining the weights of the multiple equipment indicators; based on the multiple weights, taking a weighted sum of the indicator values of the multiple equipment indicators to obtain sub-device attribute characteristics of the equipment indicator type; in a case where the type of attribute information set is an equipment location type, based on the pre-divided network hierarchy, determining the network location information of the equipment in the power grid system; encoding the network location information to obtain sub-device attribute characteristics of the equipment location type; in a case where the type of attribute information set is an equipment type, encoding the equipment type information in the attribute information set to obtain sub-device attribute characteristics of the equipment type.
[0091] According to an embodiment of the present disclosure, when the type of attribute information set is a device indicator type, the weights of the multiple device indicators can be determined based on the correlation between the multiple device indicators in the attribute information set, wherein the device indicators may include business continuity impact, data sensitivity, user dependence, system dependence, resource utilization and security risk level. Among the multiple device indicators, The device index and The correlation between the device indicators can be express.
[0092] Preferably, the value range of the correlation may be [1 / 9, 9]. The device index and The correlation between the device indicators can be used to express the Compared with the equipment indicators, The importance of each device indicator, therefore, the relevance The larger the value, the better the device index Relative to device indicators The higher the importance of . In particular, ,exist = In the case of .
[0093] According to an embodiment of the present disclosure, based on multiple correlations, a correlation matrix P can be determined, as shown in formula (2):
[0094] (2)
[0095] Where n is the number of device indicators.
[0096] According to an embodiment of the present disclosure, the correlation matrix is determined using formula (3): The eigenvalues of:
[0097] (3)
[0098] in, is the identity matrix, that is, a matrix with all 1s on the diagonal and 0s on the rest. Representation matrix The determinant of . From the correlation matrix Among at least one eigenvalue of .
[0099] According to the embodiment of the present disclosure, the maximum eigenvalue can be determined using formula (4). The corresponding eigenvector , as shown in formula (4):
[0100] (4)
[0101] According to the embodiments of the present disclosure, Normalizing multiple components in , we can get , The sum of the multiple components in is 1, Indicates business continuity impact, Indicates data sensitivity, Indicates user dependence, Indicates system dependency, Indicates resource utilization and Indicates the weight of each security risk level.
[0102] According to the embodiment of the present disclosure, data collection is performed on the device to obtain the index values of multiple device indicators. According to the weights of the multiple device indicators, the index values of the multiple device indicators are weighted and summed to obtain the sub-device attribute characteristics of the device indicator type. , as shown in formula (5):
[0103] (5)
[0104] in, 、 、 、 、 and They are the indicator values of business continuity impact, data sensitivity, user dependence, system dependence, resource utilization and security risk level respectively.
[0105] According to an embodiment of the present disclosure, when the type of the attribute information set is the device location type, the network location information of the device in the power grid system is determined based on the pre-divided network layers, wherein, for example, the network can be pre-divided into three network layers, including the access layer, the aggregation layer and the core layer, and the network location information of the device in the power grid system can be determined according to the actual network layer where the device is located.
[0106] According to the embodiment of the present disclosure, the network location information can be encoded to obtain the network location information encoding of the device. For example, the code of the access layer device is set to 1, the code of the convergence layer device is set to 2, and the code of the core layer device is set to 3. Since the relative positions of different devices in the same network layer can be different, the relative positions of the devices in the network layer can also be encoded to obtain and In the specific implementation, and The way to obtain is to manually define the two-dimensional coordinate mapping based on the actual distribution of devices in physical space or logical topology. For example, in the access layer network, if multiple devices are deployed in the same room, their Set it as the computer room logo (such as Indicates the computer room CR A ), The equipment is coded incrementally according to its specific location in the computer room (such as rack number or row number), for example, the computer room CR A Medium equipment CR A1 of , equipment CR A2 of If the equipment is located in different computer rooms, will change (e.g. Indicates the computer room CR B ), and z(v) starts numbering again from 1 in the new room. This coding method ensures that the equipment in the same room has the same , and the relative position relationship of equipment across computer rooms is reflected through coordinate differences. The larger the difference in values, the farther the computer room is from the physical location.
[0107] According to the above encoding of network location information 、 and , the sub-device attribute characteristics of the device location type can be determined using formula (6) :
[0108] (6)
[0109] According to an embodiment of the present disclosure, when the type of the attribute information set is the device type, the device type of the device is determined, wherein the device type may include two categories: security devices and ordinary devices. Security devices may specifically include firewalls, intrusion detection systems, intrusion prevention systems, virtual private networks, Web application firewalls, unified threat management, next-generation firewalls, etc. Ordinary devices may specifically include servers, workstations, network devices, storage devices, wireless devices, Internet of Things devices, industrial control devices, etc.
[0110] According to the embodiment of the present disclosure, the device type information in the attribute information set can be encoded in the form of a one-hot vector to obtain the sub-device attribute characteristics of the device type. ,in, The dimension is the same as the total number of categories of device types. In the case of the device type, No. The first component is set to 1, and the other components are set to 0.
[0111] According to an embodiment of the present disclosure, when the type of the attribute information set is the device value type, the initialization cost of the device is , maintenance costs and business value , the sub-device attribute characteristics of the device value type can be determined using formula (7) :
[0112] (7)
[0113] in, The initialization cost is , maintenance costs and business value The weight coefficient of Non-negative, satisfying Initialization cost Including equipment purchase cost and initial deployment cost, maintenance cost Including equipment including daily operation and maintenance costs, update and upgrade costs and technical support costs, business value The impact of the equipment on business continuity can be determined.
[0114] According to an embodiment of the present disclosure, when the type of the attribute information set is a workload type, the CPU usage of the device can be used to calculate the workload type. , memory usage and network bandwidth usage , according to formula (8), determine the sub-device attribute characteristics of the workload type :
[0115] (8)
[0116] in, CPU usage , memory usage and network bandwidth usage The weight coefficient of Non-negative, satisfying , CPU usage , memory usage and network bandwidth usage It can be determined by formulas (9) to (11):
[0117] (9)
[0118] (10)
[0119] (11)
[0120] in, is the moment within the sampling period T, is the number of moments in the sampling period T, for The CPU usage of the device at the moment, for The memory usage of the device at the moment, for The network bandwidth usage of the device at all times, The maximum bandwidth capacity of the device.
[0121] According to an embodiment of the present disclosure, when the type of attribute information set is a connectivity type, based on the connection status of the device with other devices in the power grid system, the attribute characteristics of the sub-device of the connectivity type can be determined using formula (12): :
[0122] (12)
[0123] in, Representation device The in-degree, Representation device The out-degree of the device is here Indicates the number of elements in the collection.
[0124] According to the embodiment of the present disclosure, when the type of the attribute information set is a vulnerability type, after checking the device and determining the security vulnerability, the device can be used to determine the device's security vulnerability using formula (13). At least one security vulnerability is assessed and the sub-device attribute characteristics of the vulnerability type are determined :
[0125] (13)
[0126] in, Indicates the The inherent characteristics of a security vulnerability, Indicates the The timeliness of security vulnerabilities, Indicates the The environmental score of a security vulnerability can be used to represent the device The impact of the security vulnerability in the environment, Representation device A collection of existing security vulnerabilities.
[0127] According to an embodiment of the present disclosure, when the type of the attribute information set is a device configuration type, the operating system type, the open port set, and the running service set in the device can be determined, and the sub-device attribute characteristics of the device configuration type can be determined using formula (14). :
[0128] (14)
[0129] in, Representation device The operating system type can be encoded by an enumeration type. Representation device The open port set is the set of network port numbers that are currently in listening state and allow external communication. For example, {22,80,443} means that the device has opened SSH (port 22), HTTP (port 80), and HTTPS (port 443) services; Representation device The running service collection in includes the service names. For example, {sshd,httpd,nginx} indicates that the device is running the SSH service (sshd), Apache HTTP service (httpd), and Nginx service. By sequentially concatenating the above three types of features to form a complete feature vector, The length is the sum of three characteristic lengths, where It is a single-value enumeration code with a length of 1. and They are a set of port numbers and a set of service names, respectively. Their length is determined by the number of ports actually opened on the device and the number of services running.
[0130] According to an embodiment of the present disclosure, for multiple attribute information types, attribute information corresponding to the attribute information type is first collected, and then the sub-device attribute characteristics are determined by calculation, so that the characteristics of each of the multiple sub-device attribute characteristics are clearer, ensuring the accuracy and comprehensiveness of the device attribute characteristics.
[0131] According to an embodiment of the present disclosure, the network defense method also includes: obtaining a communication attribute information set between two connected devices in a device topology diagram, wherein the communication attribute information set includes communication type information, communication frequency information, and communication importance information; and encoding the communication attribute information set to obtain edge attribute features.
[0132] According to embodiments of the present disclosure, a communication attribute information set can be used to represent the communication status between two connected devices. Communication type information can be determined based on the connection mode between the two connected devices, which can include local connection and remote connection. Communication frequency information can be determined based on the number of communications between the two connected devices per unit time. Communication importance information can be determined based on the sensitivity of the communication content and the importance of the service between the two connected devices.
[0133] According to the embodiment of the present disclosure, the information in the communication attribute information set is encoded to obtain the edge attribute feature for representing the edge connecting the two connected devices. , as shown in formula (15):
[0134] (15)
[0135] in, Indicates the communication type information. For example, the connection mode includes local connection and remote connection. In the case of local connection, you can The encoding result is assigned a value of 1. When the connection mode is remote connection, The encoding result is assigned a value of 0.5. Indicates communication frequency information, Indicates the importance of communication information. According to formulas (16) to (17), and Perform the calculation:
[0136] (16)
[0137] in, and For two connected devices, for and The number of communications per unit time, The preset communication number threshold.
[0138] (17)
[0139] in, for and The sensitivity of the content of the communication between for and The importance of the business handled by the communication, They are and The weight coefficient of Non-negative, satisfying .
[0140] According to an embodiment of the present disclosure, the communication attribute information set between two connected devices is encoded to obtain edge attribute features, which can ensure that the edge in the device topology map correctly represents the communication attributes between the two devices connected by the edge, improve the mapping ability of the device topology map to the power grid system, and ensure the effectiveness and accuracy of network defense based on the device topology map.
[0141] According to an embodiment of the present disclosure, device attribute features include sub-device attribute features of a device location type, and based on device topology features, device attribute features of multiple devices in a device topology graph, and edge attribute features of multiple edges, multiple devices in a power grid system are divided to obtain multiple security domains and a global domain, including: obtaining device embedding features of multiple devices based on device topology features, device attribute features of multiple devices, and sub-device attribute features of a device location type; performing attention feature extraction and feature update on each device embedding feature in turn to obtain multiple target device features; performing feature fusion on multiple target device features, edge attribute features of multiple edges, and device topology features to obtain a target fusion feature; based on the target fusion feature, determining domain identifiers of multiple devices for representing the security domains to which they belong; and dividing multiple devices in the power grid system based on the domain identifiers of multiple devices to obtain multiple security domains and a global domain.
[0142] According to an embodiment of the present disclosure, based on the device attribute characteristics of each of the multiple devices , the initial device embedding characteristics of multiple devices can be determined by formula (18) :
[0143] (18)
[0144] in, and is a learnable parameter, .
[0145] According to an embodiment of the present disclosure, the sub-device attribute characteristics of the device location type , the logical position of each device can be encoded to obtain the position embedding vector , as shown in formula (19):
[0146] (19)
[0147] in, and are learnable parameters.
[0148] According to an embodiment of the present disclosure, the initial device of device v can be embedded with a feature and position embedding vector Fusion, and update the initial device embedding features to obtain device embedding features , so that the message can be passed later, as shown in formula (20):
[0149] (20)
[0150] According to an embodiment of the present disclosure, for each pair of adjacent devices , the attention weight can be calculated using formula (21) :
[0151] (twenty one)
[0152] in, , , is the parameter matrix, is a trainable vector, is the vector concatenation operation, is the edge between devices u and v, For the edge The edge attribute characteristics of For the edge The edge attribute features of 、 and are devices v, u and Device embedding features, For devices The set of neighbor devices, is the first activation function, when Output when it is a positive number ,when Output when it is non-positive According to an embodiment of the present disclosure, attention feature extraction and feature update are performed on the device embedded features based on the attention weight, as shown in formula (22):
[0153] )(twenty two)
[0154] in, is the trainable matrix of the graph neural network layer, For device v Layer device embedding features, the formula represents the first Layer device embedding features through Neighbor device embedding features of the layer Calculated, when iteratively updating to the last layer L, the target device characteristics can be determined .
[0155] According to an embodiment of the present disclosure, determining The corresponding multiple neighboring devices, and through the weighted aggregation function of the neighboring devices, the device topology characteristics of device v can be obtained , as shown in formula (23):
[0156] (twenty three)
[0157] in, Represents an aggregate function, For devices Initial device embedding characteristics.
[0158] According to an embodiment of the present disclosure, the target device feature of device v , edge attribute characteristics of adjacent edges of device v and device topology characteristics By performing feature fusion, we can obtain the target fusion features of device v , as shown in formula (24):
[0159] (twenty four)
[0160] According to an embodiment of the present disclosure, a classifier can be used to determine the domain identifiers of multiple devices used to characterize their respective security domains. , as shown in formula (25):
[0161] (25)
[0162] in, and are the classifier parameters, It is the second activation function, which can normalize a numerical vector into a probability distribution vector, and the sum of each probability is 1.
[0163] According to an embodiment of the present disclosure, based on the domain identifiers of multiple devices, the security domains to which multiple devices in the power grid system belong can be determined, and the multiple devices in the power grid system can be divided into the security domains to which they belong, thereby obtaining multiple security domains and a global domain.
[0164] According to an embodiment of the present disclosure, in the process of partitioning multiple devices in a power grid system, an optimization objective function can be set so that the partitioning result meets the requirements. The optimization objective function may include maximizing the sum of intra-domain communication weights to reduce inter-domain communication overhead, as shown in formula (26):
[0165] (26)
[0166] in, For the The set of edges in a security domain, is the number of security domains.
[0167] According to an embodiment of the present disclosure, the optimization objective function may further include: preventing certain security domains from being overly concentrated by constraining the total weight of critical devices in each security domain, as shown in formula (27):
[0168] (27)
[0169] in, For the The device set in a security domain, c is For any device in the, K(c) represents the device indicator type characteristics of device c. The larger the K(c) value, the more critical the device c is. The preset upper limit of the critical equipment weight.
[0170] According to an embodiment of the present disclosure, the optimization objective function may further include: introducing a location-aware regularization term to penalize the partitioning results of devices in the security domain that are too far apart, as shown in formula (28):
[0171] (28)
[0172] in and Represents a collection of security domain devices Any two different devices and The device location type attribute characteristics.
[0173] According to the embodiments of the present disclosure, based on the above optimization objective function, the classifier parameters and attention weights can be iteratively adjusted to minimize the communication overhead and location dispersion and meet the key device constraints, thereby obtaining the optimal device security domain division result. The divided security domain device set includes .
[0174] According to the embodiments of the present disclosure, based on device topology features, device attribute features, and edge attribute features, the device embedding features of each of multiple devices can be determined. By performing attention feature extraction and feature update on the device embedding features, the target device features are obtained, thereby ensuring the authenticity and accuracy of the target device features. Feature fusion of the target device features, edge attribute features, and device topology features can ensure that the obtained target fusion features are richer and contain more comprehensive features. Based on the target fusion features, the domain identifiers of each of the multiple devices are determined, thereby achieving accurate domain division of the devices and obtaining security domains and global domains.
[0175] According to the embodiment of the present disclosure, after completing the security domain division, the inter-domain security rule configuration method based on communication requirements and security policies can be used to configure the security rules between any two security domains. and , and the device sets in its security domain are and ,in , define its communication relationship matrix ,in, and Represents security domains and The number of devices in The elements in the matrix are shown in formula (29):
[0176] , (29)
[0177] in, Security Domain Device collection Any device in Security Domain Any device in For devices and The number of communications between and Edge Communication frequency information and communication importance information, if there is no connection relationship between devices c and d, then and All are 0.
[0178] According to the embodiment of the present disclosure, based on the communication relationship matrix, the inter-domain communication strength index can be calculated using formula (30): :
[0179] (30)
[0180] Among them, p and q represent two security domains respectively. and are the device sets in security domains p and q respectively, for Any device in for Any device in Security Domain and The communication relationship matrix between Represents a function Take the maximum value among all possible values of the variable x.
[0181] According to the embodiment of the present disclosure, the inter-domain security risk index can be calculated using formula (31): :
[0182] (31)
[0183] in, Indicates the number of elements in the collection. and Equipment and the sub-device attribute characteristics of the vulnerability type of device d, It is the maximum value of the sub-device attribute characteristics of the vulnerability type in the network.
[0184] According to the embodiment of the present disclosure, for a security domain with communication requirements , according to its communication strength index and security risk indicators , you can design a set of inter-domain firewall rules .in, The following rules can be included, source address range: , destination address range: , service port collection: , communication protocol: ,action: {Allow, Deny}.
[0185] According to the embodiment of the present disclosure, the generation of rules needs to follow multiple principles, including the inter-domain communication strategy function , as shown in formula (32):
[0186] (32)
[0187] in, and is the security risk threshold. Based on the output of the above policy function, corresponding rules can be used to generate policies. If the output is strict (strict policy), rules can include denying all communication by default, allowing only necessary business communication, restricting service port ranges, and implementing stricter protocol restrictions. If the output is moderate (moderate policy), rules can include denying all communication by default, allowing general business communication, appropriately relaxing port restrictions, and allowing common protocols. If the output is loose (loose policy), rules can include allowing intra-domain communication by default, restricting only high-risk services, using a larger port range, and allowing multiple protocols.
[0188] According to an embodiment of the present disclosure, for each pair of security domains , the final rule set can be generated by the optimization objective shown in formula (33):
[0189] (33)
[0190] in, is the complexity cost of the rule.
[0191] According to the embodiment of the present disclosure, the generation of rules also needs to satisfy the constraints shown in formulas (34) to (35):
[0192] (34)
[0193] (35)
[0194] in, is the coverage of the necessary communication by the rule set, security risks introduced to the rule set, and is the adjustment parameter.
[0195] According to the embodiments of the present disclosure, by intelligently dividing the network structure and security features into security domains and establishing a corresponding inter-domain security policy configuration mechanism, a scientific partitioning management and control solution is provided for network security protection.
[0196] Figure 3 The data flow diagram of the security domain division process of the network defense method according to the embodiment of the present disclosure is schematically shown.
[0197] like Figure 3 As shown, the device topology diagram , sub-device attribute characteristics of device location type and device attribute characteristics The feature embedding layer is input to process the device topology map using the topology encoder, the sub-device attribute features of the device location type are processed using the location-aware encoder, and the device attribute features are processed using the node embedder. The processing results of the above three parts are fused to obtain the device embedding features. The device embedding features are processed using the message passing layer, and after attention feature extraction and feature update, the target device features can be obtained. The target device features are processed using the feature fusion layer, and the target device features, the edge attribute features of the edges in the device topology map, and the device topology features are fused to obtain the target fusion features. The target fusion features are input to the classifier to obtain the domain identifier of the device used to characterize the security domain to which it belongs. Based on the domain identifiers of multiple devices, the security domain can be divided to obtain the security domain division result.
[0198] According to an embodiment of the present disclosure, network defense action information of each device is obtained based on the state vectors of each of the multiple security domains and the state vector of the global domain, including: inputting the state vector of the security domain into the intra-domain intelligent agent to obtain the intra-domain network defense action information of the security domain; inputting the state vector of the global domain into the inter-domain intelligent agent to obtain the inter-domain network defense action information of the global domain; and obtaining network defense action information based on the intra-domain network defense action information and the inter-domain network defense action information.
[0199] According to an embodiment of the present disclosure, each security domain includes an intra-domain intelligent agent, which is used to independently be responsible for optimizing the defense strategy within the security domain, and dynamically adjust it according to the security status and defense configuration of the equipment within the security domain, to ensure efficient utilization of resources within the security domain and threat protection capabilities within the security domain.
[0200] According to an embodiment of the present disclosure, the state vector of the security domain may include device features , edge attribute features , current defense configuration status, security event detection information within this security domain, etc., among which the current defense configuration status may include firewall rules, intrusion detection system (IDS) activation status, etc., and the security event detection information within this security domain may include the frequency of threat occurrence in the recent period, etc.
[0201] According to embodiments of the present disclosure, the state vector of a security domain can be input into an intra-domain agent, thereby deriving actions for the intra-domain agent. These actions can include policy adjustments to defense devices within the security domain, including activating or deactivating specific IDS rules, adjusting firewall policies such as allowing or disabling certain types of traffic, reallocating device security resources, changing bandwidth priorities, or enabling isolation policies. Derivation of the intra-domain agent's actions from the security domain's state vector requires iterative training of the intra-domain agent, a training process described later.
[0202] According to an embodiment of the present disclosure, the security domain The state vector Input to the agent in the domain, the action space set in the domain , get the network defense action information of security domain i at time t , as shown in formula (36):
[0203] (36)
[0204] Where t represents the tth time step, Indicates the A collection of devices within a security domain, For devices actions, such as setting specific parameters for a defense measure.
[0205] According to the embodiments of the present disclosure, an inter-domain intelligent agent can be set up between global domains. The inter-domain intelligent agent is used for the overall coordination and optimization of cross-domain defense. By observing the status of each security domain, security policy adjustments for inter-domain communications are implemented, such as inter-domain traffic control, firewall rule configuration, and topology structure optimization.
[0206] According to an embodiment of the present disclosure, the state vector of the global domain It can be used to describe the global status of the entire network, which may include the comprehensive feature vector of each security domain, the state vector of the security domain, inter-domain communication characteristics, etc. Among them, inter-domain communication characteristics may include inter-domain communication frequency, importance and historical security event statistics.
[0207] According to embodiments of the present disclosure, the global domain state vector can be input into an inter-domain agent to derive the inter-domain agent's actions. These actions can be used to adjust cross-domain defense strategies, including configuring inter-domain firewall rules, dynamically adjusting inter-domain communication path priorities, and enabling or disabling high-risk cross-domain communications. Derivation of the inter-domain agent's actions from the global domain state vector requires iterative training of the inter-domain agent, a training process described later.
[0208] According to an embodiment of the present disclosure, the state vector of the global domain is Input to the inter-domain agent, the inter-domain action space set , get the inter-domain network defense action information of the global domain , as shown in formula (37):
[0209] (37)
[0210] in, is the set of inter-domain edges, For inter-domain edges specific policies, such as limiting communication traffic or enabling isolation mechanisms.
[0211] According to the embodiments of the present disclosure, intra-domain network defense action information and inter-domain network defense action information may be spliced and fused to obtain network defense action information.
[0212] According to the embodiments of the present disclosure, intra-domain network defense action information and inter-domain network defense action information are determined for the security domain and the global domain respectively, and the network defense action information is further determined so that the network defense action information takes into account the status information between multiple devices within the domain and the communication status information between multiple domains, thereby making the network defense action more comprehensive and accurate, and the network can be optimized more effectively based on the network defense action.
[0213] Figure 4 The flowchart of a network defense method according to another embodiment of the present disclosure is schematically shown.
[0214] like Figure 4 As shown, after determining a device topology based on multiple devices in the power grid system, the devices corresponding to each of the multiple devices in the device topology are divided according to the characteristics of the device topology, resulting in a security domain division result. Based on the division results, multiple security domains can be determined, namely security domain 1, security domain 2, and security domain 3. Based on the communication relationships between multiple global domains, a global domain can be obtained. By analyzing and processing the state vector of the global domain using an inter-domain agent, inter-domain network defense action information for the global domain can be obtained. By analyzing and processing the state vectors of each of the multiple security domains using their respective intra-domain agents, inter-domain network defense action information for the security domains can be obtained. Through the collaborative work of the inter-domain agent and the intra-domain agent, network defense action information can be obtained based on the intra-domain network defense action information and the inter-domain network defense action information.
[0215] According to an embodiment of the present disclosure, the network defense method further includes: obtaining a state vector of the security domain based on device attribute characteristics, edge attribute characteristics, network configuration status characteristics, and network abnormal event characteristics.
[0216] According to an embodiment of the present disclosure, the state vector of the security domain can be determined by formula (38): :
[0217] (38)
[0218] in, is the device attribute characteristic of device c, The network configuration status feature of device c is used to indicate the current defense configuration status. is the network abnormal event feature of device c, which is used to represent the security event detection information within this security domain. is the set of devices in security domain i, is the edge set in security domain i.
[0219] According to an embodiment of the present disclosure, the network defense method further includes: obtaining a state vector of the global domain based on inter-domain communication characteristics and state vectors of each of the multiple security domains.
[0220] According to an embodiment of the present disclosure, the state vector of the global domain can be determined by formula (39): :
[0221] (39)
[0222] in, Security Domain The state vector of is the inter-domain communication characteristic, D is the number of security domains in the power grid system, is the set of inter-domain edges.
[0223] According to an embodiment of the present disclosure, the state vectors of the security domain and the global domain are determined respectively based on the characteristics in the security domain and the characteristics in the global domain, so as to describe the local environmental characteristics of the security domain and the global state of the entire power grid system, so as to determine the strategy for defense optimization within and between domains based on the state vector of the security domain and the state vector of the global domain.
[0224] According to an embodiment of the present disclosure, the network defense method further includes: updating the network defense action information based on historical experience rules to obtain target network defense action information.
[0225] According to an embodiment of the present disclosure, historical experience rules may be used to represent feedback from the power grid system after executing a specific network defense action based on network defense action information for different states of the power grid system.
[0226] Therefore, according to the current state of the power grid system, we can determine the specific network defense action that can obtain better feedback in the current state from historical experience rules, and use this network defense action to update the network defense action information to obtain the target network defense action information.
[0227] According to the embodiments of the present disclosure, network defense action information is updated based on historical experience rules, which can select more effective target network defense action information and improve the effect of network defense.
[0228] According to an embodiment of the present disclosure, a reward function can be used to determine the feedback information obtained by the intelligent agent after taking specific network defense actions in various states of the power grid system, and the feedback information can be used to guide the intelligent agent to update and optimize the network defense action information.
[0229] According to an embodiment of the present disclosure, reward functions may be designed separately for intra-domain agents and inter-domain agents.
[0230] According to an embodiment of the present disclosure, the reward function of the intelligent agent in the domain may include security rewards and resource consumption penalties, wherein the security rewards can be used to measure the reduction of security incidents in this security domain, including threat interception effect, protection coverage, etc.
[0231] According to the embodiment of the present disclosure, the threat interception effect It can be determined by formula (40):
[0232] (40)
[0233] in, For threat collection, represents any threat in the threat set, Indicates that the agent successfully intercepts the threat the number of is the total number of threats detected, For threats The weight coefficient is proportional to the severity of the threat.
[0234] According to the embodiment of the present disclosure, the protection coverage It can be determined by formula (41):
[0235] (41)
[0236] in, Indicates the device before the action is executed Sub-device attribute characteristics of the vulnerability type, Indicates that the device Sub-device attribute characteristics of the vulnerability type, Set for device The number of devices in the .
[0237] According to the embodiment of the present disclosure, after determining the threat interception effect and protection coverage, the security reward can be determined by formula (42): :
[0238] (42)
[0239] in, and are weight coefficients, and their sum is 1.
[0240] According to an embodiment of the present disclosure, resource consumption penalty It is used to express the penalty for exceeding the limit of resource usage (such as CPU and memory) of the defense device, as shown in formula (43):
[0241] (43)
[0242] in, Device before action execution The sub-device attribute characteristics of the workload type, Device after action execution The sub-device attribute characteristics of the workload type.
[0243] According to an embodiment of the present disclosure, the reward function of the agent in the domain can be determined by formula (44): :
[0244] (44)
[0245] in, and is the weight parameter.
[0246] According to an embodiment of the present disclosure, the reward function of the inter-domain intelligent agent can be used to measure the effectiveness of the intelligent agent in cross-domain defense coordination, including communication efficiency reward and security risk reward, wherein the communication efficiency reward can be used to measure the smoothness of inter-domain communication, and the security risk reward can be used to reflect the degree of reduction of inter-domain security risks.
[0247] According to an embodiment of the present disclosure, the communication efficiency reward can be determined by formula (45): :
[0248] (45)
[0249] in, and are the inter-domain communication intensity before and after the action is executed.
[0250] According to an embodiment of the present disclosure, the security risk reward can be determined by formula (46): :
[0251] (46)
[0252] in, and are the inter-domain security risk indicators before and after the action is executed, respectively. The calculation method is shown in formula (31).
[0253] According to an embodiment of the present disclosure, the reward function of the inter-domain agent can be determined by formula (47): :
[0254] (47)
[0255] in, and is the weight parameter, and the sum is 1.
[0256] According to the embodiment of the present disclosure, the reinforcement learning method can be used to learn the strategy Maximize the cumulative reward to ensure that the network defense action information is updated in a better direction.
[0257] According to an embodiment of the present disclosure, the domain goal can be determined by formula (48) to maximize the cumulative reward of the agent in the domain :
[0258] (48)
[0259] in, is the discount factor, represents the policy function within the domain, Indicates that the reinforcement learning strategy is Next, we solve for the expectation of f(x).
[0260] According to an embodiment of the present disclosure, the inter-domain objective can be determined by formula (49) to maximize the cumulative reward of the inter-domain agent. :
[0261] (49)
[0262] Where t represents the tth time step, Represents the inter-domain policy function.
[0263] According to the embodiments of the present disclosure, the reinforcement learning module of the domain agent focuses on optimizing the defense strategy of a single security domain, which is implemented through a multi-agent deep Q network. This module assigns an independent agent to each security domain, which adjusts the defense strategy within the security domain based on the state within the security domain. The module input is the state vector , including device characteristics , edge features , Current Defense Configuration and security incident information .
[0264] According to the embodiment of the present disclosure, the design domain Value Function , used to estimate the state Next select action Specifically, The value is represented in the state Take action After that, the agent is expected to obtain a cumulative reward, which includes the immediate reward ( ) and the additional rewards that can be obtained through continuous actions in the future. Therefore, The value not only takes into account the direct effect of the current action, but also reflects the effect of the action in the future multiple time steps. The long-term benefits in , as shown in formula (50):
[0265] (50)
[0266] in, is the discount factor, is the mathematical expectation function. The network adopts a multi-layer perceptron (MLP) structure and is trained through the experience replay mechanism.
[0267] According to an embodiment of the present disclosure, the domain reinforcement learning module is based on - Greedy strategy, achieving a balance between exploration and exploitation, as shown in formula (51):
[0268] (51)
[0269] in, Indicates taking the action a that maximizes f(a), Expressed as probability Get , represents uniformly random selection of an action.
[0270] According to the embodiment of the present disclosure, the above operations can be used to configure the reinforcement learning module of the agent in the domain. When learning using the reinforcement learning module, the Value Function Network and target network , and set the initial weight to enter the iteration. In each round of iteration, the agent starts from the current state Select Action , get rewards after interacting with the environment and the next state .Will( ) is stored in the experience replay pool, and small batches of data are sampled for network training. During the network training process, the experience quadruple is taken out from the experience pool ,in Represents state, action, reward and next state respectively, used to update The value function network is expressed as a loss function as shown in formula (52). Optimize the network:
[0271] (52)
[0272] According to the embodiments of the present disclosure, the reinforcement learning module of the inter-domain agent can optimize the global defense strategy using the policy gradient method for defense coordination between multiple security domains. By observing the global state, the agent outputs a strategy adjustment plan for cross-domain communication.
[0273] According to an embodiment of the present disclosure, the input of the reinforcement learning module of the inter-domain agent is the global state vector , including the state vector of security domain i and inter-domain communication characteristics The learning strategy of the reinforcement learning module of the inter-domain agent is shown in formula (53):
[0274] (53)
[0275] in, To parameterize the network Generated inter-domain actions The probability distribution of is the policy network output, is the set of inter-domain action spaces, for Any action in .
[0276] According to an embodiment of the present disclosure, the value function shown in formula (54) can be used , estimated state and use it for strategy optimization.
[0277] (54)
[0278] According to the embodiment of the present disclosure, the above operations can be used to configure the reinforcement learning module of the inter-domain agent, and when learning using the reinforcement learning module, the policy network can be initialized. and value function network , and set the initial weight parameters 、 , enter the iteration. In each round of iteration, the agent Select Action , get rewards after interacting with the environment and the next state Using the advantage function shown in formula (55) Optimization strategy network:
[0279] (55)
[0280] According to the embodiment of the present disclosure, the policy network parameters can be calculated using formula (56): To update:
[0281] (56)
[0282] in, Indicated by the parameter The control strategy function, Indicates that the strategy Perform logarithmic operations on the given probabilities, Represents the objective function About parameters Gradient operation.
[0283] According to an embodiment of the present disclosure, the loss function shown in formula (57) can be used , for the value function network parameters To update:
[0284] (57)
[0285] According to the embodiments of the present disclosure, in order to achieve effective collaboration between intra-domain and inter-domain intelligent agents in the deep collaborative reinforcement learning framework, a collaborative mechanism based on shared reward signals and experience pools can be utilized to enhance the coordination between strategies through information interaction and experience sharing between intelligent agents, accelerate the learning process, and avoid the degradation of global defense performance caused by local optimization.
[0286] According to the embodiments of the present disclosure, the optimization objectives of intra-domain and inter-domain agents during reinforcement learning differ. Intra-domain agents focus on optimizing the defense strategy of a single security domain, while inter-domain agents focus on coordinating global defense strategies. Therefore, to ensure inter-strategy synergy, a shared reward signal coordination mechanism can be utilized to design the reward signals of intra-domain and inter-domain agents to be interconnected, thereby achieving global optimization and adjustment.
[0287] According to the embodiment of the present disclosure, the shared reward of the intelligent agents in the domain Not only the defense effect within the security domain should be considered, but also the contribution to the global defense should be considered, as shown in formula (58):
[0288] (58)
[0289] in, For rewards within the original security domain, Rewards for global contributions, A coordination factor for global contribution rewards, The inter-domain security risk index can be calculated using formula (59):
[0290] (59)
[0291] in, Respectively represent the security domains before and after the action is executed and security domains security risk indicators between them.
[0292] According to the embodiment of the present disclosure, the shared reward between the inter-domain agents It is necessary to consider the inter-domain communication efficiency and the defense effect within each domain, as shown in formula (60):
[0293] (60)
[0294] in, is the original inter-domain reward, For the The local contribution of each security domain, is the local contribution coordination factor, is the security domain weight coefficient, The defense effect within the domain can be calculated using formula (61):
[0295] (61)
[0296] in, and are the security rewards of the i-th security domain before and after the action is executed, It is proportional to the number of key devices in the security domain, as shown in formula (62):
[0297] (62)
[0298] Where K(c) represents the device indicator type feature of device c. The larger the K(c) value, the more critical the device c is. According to the embodiment of the present disclosure, when the intelligent agent in the domain performs a defense action, it not only generates a defense reward for the security domain, but also affects the security risk index between domains. To change the global contribution reward , thus affecting the shared reward within the domain. This shared reward will guide the agents within the domain to consider the impact on other security domains while optimizing the defense of their own security domain. At the same time, the defense effect of the agents within the domain is calculated through local contributions. Included in inter-domain shared rewards , affecting the decision-making of inter-domain agents. Therefore, this two-way reward sharing mechanism enables intra-domain and inter-domain agents to perceive and influence each other, thus naturally forming synergy in their respective optimization processes and jointly improving the overall network defense effect.
[0299] According to the embodiments of the present disclosure, the experience sharing mechanism can share key state-action-reward sequences between intra-domain and inter-domain agents by establishing a unified experience pool, thereby improving data utilization efficiency and accelerating strategy convergence.
[0300] According to the embodiment of the present disclosure, in each training iteration, the intra-domain and inter-domain agents generate experience samples respectively. ,in, and Represent the current and next round states respectively, Indicates the action to be performed currently. Represents the reward corresponding to the current action. The experience samples are stored in the corresponding sub-experience pool, where the sub-experience pool can include the intra-domain experience pool, the inter-domain experience pool and the global shared experience pool. The intra-domain experience pool Used to store samples of agents within a domain and experience pools between domains Used to store samples of inter-domain agents and a globally shared experience pool Used to store the above two types of experience and for collaborative training.
[0301] According to an embodiment of the present disclosure, the experience pool uses a priority experience replay technique to determine the key state-action-reward sequence and defines its priority according to the TD error of the experience sample. , as shown in formula (63):
[0302] (63)
[0303] in, is the discount factor, is the Q-value function of the corresponding agent.
[0304] According to the embodiments of the present disclosure, during training, small batches of data are sampled from the experience pool according to priority to improve the efficiency of utilizing key experience. Through this collaborative mechanism, agents within a domain can learn from the defense experience of agents in other domains, as well as from inter-domain defense experience. Inter-domain agents can also learn from the defense strategies within their domain, thereby accelerating strategy convergence and improving overall defense effectiveness.
[0305] According to the embodiment of the present disclosure, in the process of training and updating the collaborative mechanism and experience sharing, the policy network parameters of the intra-domain agent and the inter-domain agent are first initialized, and the intra-domain experience pool is created. , inter-domain experience pool And the global shared experience pool Then determine the action to be performed by the agent, that is, determine the agent in the domain in the current state Select Action , get rewarded , and determine the current state of the inter-domain agent Select Action , get rewarded .
[0306] According to the embodiment of the present disclosure, the agent in the domain is rewarded Corrected to , the inter-domain agent reward Corrected to . The domain experience ( ) and inter-domain experience ( ) stored in the domain experience pool and inter-domain experience pools At the same time, the above experience is stored in the global shared experience pool , thereby achieving the sharing of reward signals.
[0307] According to an embodiment of the present disclosure, the agent in the domain is drawn from the experience pool in the domain. Sample a portion of data according to priority, and at the same time extract data from the shared experience pool In the training process, a portion of the data is sampled according to priority, and the ratio of the two parts of experience is dynamically adjusted according to the actual training situation. The in-domain policy network is then updated based on the collected experience to complete the policy optimization of the in-domain agent.
[0308] According to an embodiment of the present disclosure, the inter-domain agent draws from the inter-domain experience pool Sample a portion of data according to priority, and at the same time extract data from the shared experience pool In the process of sampling a portion of data according to priority, the ratio of the two parts of experience is dynamically adjusted according to the actual training situation. The inter-domain policy network is then updated based on the collected experience to complete the policy optimization of the inter-domain agent.
[0309] Based on the above network defense method, the present disclosure also provides a network defense device. Figure 5 The device is described in detail.
[0310] Figure 5 The structural block diagram of the network defense device according to an embodiment of the present disclosure is schematically shown.
[0311] like Figure 5 As shown, the network defense device 500 of this embodiment includes a topology map determination module 510 , a feature extraction module 520 , a device classification module 530 and an action determination module 540 .
[0312] The topology determination module 510 is configured to determine a device topology based on multiple devices in the power grid system. The device topology includes devices and edges, where an edge represents the communication relationship between multiple devices. In one embodiment, the topology determination module 510 can be configured to perform operation S210 described above, which will not be further described here.
[0313] The feature extraction module 520 is used to input the device topology map into the topology encoder to obtain the device topology features. In one embodiment, the feature extraction module 520 can be used to perform the operation S220 described above, which will not be repeated here.
[0314] Device partitioning module 530 is configured to partition multiple devices in the power grid system based on device topology characteristics, device attribute characteristics of multiple devices in the device topology diagram, and edge attribute characteristics of multiple edges, thereby obtaining multiple security domains and global domains. A security domain includes at least one device with an intra-domain communication relationship, and a global domain includes multiple security domains with inter-domain communication relationships. In one embodiment, device partitioning module 530 can be configured to perform operation S230 described above and will not be further described here.
[0315] The action determination module 540 is used for the action determination module to obtain network defense action information for each device based on the state vectors of each of the multiple security domains and the state vector of the global domain. In one embodiment, the action determination module 540 can be used to perform the operation S240 described above, which will not be repeated here.
[0316] According to an embodiment of the present disclosure, the network defense device 500 further includes a first information acquisition module, a first information encoding module, and a feature determination module.
[0317] The first information acquisition module is used to acquire multiple attribute information sets of different types for each device.
[0318] The first information encoding module is configured to encode different types of attribute information sets using an encoding strategy that matches the type of the attribute information set to obtain sub-device attribute characteristics.
[0319] The feature determination module is used to obtain device attribute features based on multiple sub-device attribute features.
[0320] According to an embodiment of the present disclosure, the first information encoding module includes a weight determination submodule, a first feature determination submodule, a position determination submodule, a second feature determination submodule, and a third feature determination submodule.
[0321] The weight determination submodule is used to determine the weights of the multiple device indicators based on the correlation between the multiple device indicators in the attribute information set when the type of the attribute information set is a device indicator type.
[0322] The first feature determination submodule is configured to perform weighted summation of the respective index values of the plurality of device indexes based on a plurality of weights to obtain a sub-device attribute feature of the device index type.
[0323] The location determination submodule is used to determine the network location information of the device in the power grid system based on the pre-divided network hierarchy when the type of the attribute information set is the device location type.
[0324] The second feature determination submodule is used to encode the network location information to obtain the sub-device attribute features of the device location type.
[0325] The third feature determination submodule is configured to, when the type of the attribute information set is a device type, encode the device type information in the attribute information set to obtain a sub-device attribute feature of the device type.
[0326] According to an embodiment of the present disclosure, the network defense device 500 further includes a second information acquisition module and a second information encoding module.
[0327] The second information acquisition module is used to obtain a communication attribute information set between two connected devices in the device topology diagram, wherein the communication attribute information set includes communication type information, communication frequency information and communication importance information.
[0328] The second information encoding module is used to encode the communication attribute information set to obtain edge attribute features.
[0329] According to an embodiment of the present disclosure, the device attribute characteristics include sub-device attribute characteristics of the device location type, and the device division module 530 includes an embedded feature determination submodule, a target feature determination submodule, a fusion feature determination submodule, an identification determination submodule and a device division submodule.
[0330] The embedding feature determination submodule is used to obtain the device embedding features of the multiple devices based on the device topology features, the device attribute features of the multiple devices and the sub-device attribute features of the device location types.
[0331] The target feature determination submodule is used to extract and update the attention features of each device embedding feature in turn to obtain multiple target device features.
[0332] The fusion feature determination submodule is used to fuse multiple target device features, edge attribute features of multiple edges and device topology features to obtain target fusion features.
[0333] The identification determination submodule is used to determine the domain identification of each of the multiple devices used to characterize the security domain to which they belong based on the target fusion feature.
[0334] The device division submodule is used to divide multiple devices in the power grid system based on their respective domain identifiers to obtain multiple security domains and a global domain.
[0335] According to an embodiment of the present disclosure, the action determination module 540 includes a first action determination submodule, a second action determination submodule, and a third action determination submodule.
[0336] The first action determination submodule is used to input the state vector of the security domain into the intra-domain intelligent agent to obtain the intra-domain network defense action information of the security domain.
[0337] The second action determination submodule is used to input the state vector of the global domain into the inter-domain intelligent agent to obtain the inter-domain network defense action information of the global domain.
[0338] The third action determination submodule is used to obtain network defense action information based on the intra-domain network defense action information and the inter-domain network defense action information.
[0339] According to an embodiment of the present disclosure, the network defense device 500 further includes a first vector determination module.
[0340] The first vector determination module is used to obtain the state vector of the security domain based on device attribute characteristics, edge attribute characteristics, network configuration state characteristics and network abnormal event characteristics.
[0341] According to an embodiment of the present disclosure, the network defense device 500 further includes a second vector determination module.
[0342] The second vector determination module is configured to obtain a state vector of the global domain based on inter-domain communication characteristics and state vectors of each of the plurality of security domains.
[0343] According to an embodiment of the present disclosure, the network defense device 500 further includes an action update module.
[0344] The action update module is used to update the network defense action information based on historical experience rules to obtain the target network defense action information.
[0345] According to embodiments of the present disclosure, any multiple modules among the topology determination module 510, feature extraction module 520, device classification module 530, and action determination module 540 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present disclosure, at least one of the topology determination module 510, feature extraction module 520, device classification module 530, and action determination module 540 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of these. Alternatively, at least one of the topology map determination module 510, the feature extraction module 520, the device classification module 530, and the action determination module 540 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.
[0346] Figure 6 The block diagram schematically shows an electronic device suitable for implementing the network defense method according to an embodiment of the present disclosure.
[0347] like Figure 6 As shown, the electronic device 600 according to an embodiment of the present disclosure includes a processor 601, which can perform various appropriate actions and processes according to the program stored in the ROM 602 or the program loaded from the storage unit 608 into the RAM 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present disclosure.
[0348] Various programs and data required for the operation of the electronic device 600 are stored in the RAM 603. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The processor 601 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 602 and / or RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and RAM 603. The processor 601 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0349] According to an embodiment of the present disclosure, electronic device 600 may further include an I / O interface 605, which is also connected to bus 604. Electronic device 600 may also include one or more of the following components connected to I / O interface 605: an input unit 606 including a keyboard, mouse, etc.; an output unit 607 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage unit 608 including a hard disk; and a communication unit 609 including a network interface card such as a LAN card or modem. Communication unit 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 610 as needed, so that computer programs read from the removable media can be installed into storage unit 608 as needed.
[0350] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.
[0351] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but is not limited to: a portable computer disk, a hard disk, RAM, ROM, an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 602 and / or RAM 603 described above, and / or one or more memories other than ROM 602 and RAM 603.
[0352] The embodiments of the present disclosure also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiments of the present disclosure.
[0353] The computer program executes the above functions defined in the system / device of the embodiment of the present disclosure when the processor 601 executes the computer program. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0354] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0355] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from a removable medium 611. When the computer program is executed by the processor 601, the above-described functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0356] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0357] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0358] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways, even if such combinations and / or couplings are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or couplings are intended to fall within the scope of this disclosure.
[0359] The embodiments of the present disclosure are described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be used in combination to advantage. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. A network defense method, characterized in that: The method comprises: Determine a device topology graph based on a plurality of devices in the power grid system, wherein the device topology graph includes devices and edges, and the edges represent communication relationships between the plurality of devices; Inputting the device topology map into a topology encoder to obtain device topology features; Based on the device topology characteristics, the device attribute characteristics of the plurality of devices in the device topology graph, and the edge attribute characteristics of the plurality of edges, the plurality of devices in the power grid system are divided to obtain a plurality of security domains and a global domain, wherein the security domain includes at least one of the devices having an intra-domain communication relationship, and the global domain includes a plurality of the security domains having an inter-domain communication relationship; Inputting the state vector of the security domain into the intra-domain agent to obtain intra-domain network defense action information of the security domain; Inputting the state vector of the global domain into the inter-domain agent to obtain inter-domain network defense action information of the global domain; Obtaining network defense action information based on the intra-domain network defense action information and the inter-domain network defense action information; The method further comprises: Acquire a communication attribute information set between two devices connected in the device topology graph, wherein the communication attribute information set includes communication type information, communication frequency information, and communication importance information; and encode the communication attribute information set to obtain edge attribute features; The device attribute characteristics include sub-device attribute characteristics of the device location type. Based on the device topology characteristics, the device attribute characteristics of the plurality of devices in the device topology graph, and the edge attribute characteristics of the plurality of edges, the plurality of devices in the power grid system are divided to obtain a plurality of security domains and a global domain, including: Obtaining device embedding features of each of the plurality of devices based on the device topology feature, the device attribute features of the plurality of devices, and the sub-device attribute features of the device location type; Performing attention feature extraction and feature update on each of the device embedded features in sequence to obtain multiple target device features; Performing feature fusion on a plurality of target device features, a plurality of edge attribute features, and the device topology feature to obtain a target fusion feature; Based on the target fusion feature, domain identifiers of the multiple devices are determined to represent the security domains to which they belong; and based on the domain identifiers of the multiple devices, the multiple devices in the power grid system are divided to obtain multiple security domains and a global domain.
2. The method according to claim 1, characterized in that The method further comprises: Acquire multiple sets of attribute information of different types for each of the devices; Using a coding strategy that matches the type of the attribute information set, different types of attribute information sets are encoded to obtain sub-device attribute characteristics; and based on the multiple sub-device attribute characteristics, the device attribute characteristics are obtained.
3. The method according to claim 2, characterized in that The encoding strategy that matches the type of the attribute information set is used to encode different types of attribute information sets to obtain sub-device attribute characteristics, including: In a case where the type of the attribute information set is a device indicator type, determining a weight of each of the plurality of device indicators based on correlations between the plurality of device indicators in the attribute information set; Based on the plurality of weights, performing weighted summation on the respective index values of the plurality of device indexes to obtain a sub-device attribute feature of the device index type; In a case where the type of the attribute information set is a device location type, determining network location information of the device in the power grid system based on a pre-divided network hierarchy; Encoding the network location information to obtain sub-device attribute characteristics of the device location type; In the case where the type of the attribute information set is a device type, the attribute information of the device type is encoded to obtain the sub-device attribute characteristics of the device type.
4. The method according to claim 1, wherein The method further comprises: A state vector of the security domain is obtained based on the device attribute characteristics, the edge attribute characteristics, the network configuration state characteristics, and the network abnormal event characteristics.
5. The method according to claim 4, characterized in that The method further comprises: Based on the inter-domain communication characteristics and the state vectors of each of the plurality of security domains, a state vector of the global domain is obtained.
6. The method according to claim 1, characterized in that The method further comprises: Based on historical experience rules, the network defense action information is updated to obtain target network defense action information.
7. A network defense device, characterized in that: The device comprises: A topology determination module is configured to determine a device topology based on a plurality of devices in the power grid system, wherein the device topology includes devices and edges, and the edges represent communication relationships between the plurality of devices; A feature extraction module, configured to input the device topology map into a topology encoder to obtain device topology features; a device partitioning module, configured to partition a plurality of devices in the power grid system based on the device topology characteristics, the device attribute characteristics of the plurality of devices in the device topology graph, and the edge attribute characteristics of the plurality of edges, to obtain a plurality of security domains and a global domain, wherein the security domain includes at least one of the devices having an intra-domain communication relationship, and the global domain includes a plurality of the security domains having an inter-domain communication relationship; An action determination module, the action determination module including a first action determination submodule, a second action determination submodule and a third action determination submodule; The first action determination submodule is used to input the state vector of the security domain into the intra-domain agent to obtain intra-domain network defense action information of the security domain; The second action determination submodule is used to input the state vector of the global domain into the inter-domain agent to obtain the inter-domain network defense action information of the global domain; a third action determination submodule, configured to obtain network defense action information based on the intra-domain network defense action information and the inter-domain network defense action information; The network defense device also includes a second information acquisition module and a second information encoding module; A second information acquisition module is configured to acquire a communication attribute information set between two devices connected in the device topology diagram, wherein the communication attribute information set includes communication type information, communication frequency information, and communication importance information; The second information encoding module is used to encode the communication attribute information set to obtain edge attribute features; The device attribute feature includes a sub-device attribute feature of a device location type, and the device classification module includes an embedded feature determination submodule, a target feature determination submodule, a fusion feature determination submodule, an identification determination submodule, and a device classification submodule; an embedding feature determination submodule, configured to obtain device embedding features of each of the plurality of devices based on device topology features, device attribute features of the plurality of devices, and sub-device attribute features of the device location types; The target feature determination submodule is used to extract and update the attention features of each device embedding feature in turn to obtain multiple target device features; The fusion feature determination submodule is used to fuse multiple target device features, edge attribute features of multiple edges, and device topology features to obtain target fusion features; An identification determination submodule, configured to determine, based on target fusion features, domain identifications of multiple devices used to characterize their respective security domains; The device division submodule is used to divide multiple devices in the power grid system based on their respective domain identifiers to obtain multiple security domains and a global domain.
8. The device according to claim 7, characterized in that The device also includes a first information acquisition module, a first information encoding module and a feature determination module; A first information acquisition module is used to acquire a plurality of attribute information sets of different types for each device; A first information encoding module is configured to encode different types of attribute information sets using an encoding strategy that matches the type of the attribute information set to obtain sub-device attribute characteristics; The feature determination module is used to obtain device attribute features based on multiple sub-device attribute features.
Citation Information
Patent Citations
Network information security protection method and system based on power grid information system
CN117527360A
Self-adaptive path selection system based on GNN and multi-agent DRL
CN120301812A