Power grid network security anti-fact defense method based on multi-agent cooperation
By employing a multi-agent collaborative defense method, the problem of insufficient identification of unknown attacks in power and industrial control networks has been solved, enabling real-time and precise defense decisions and cross-site linkage, thereby improving the overall protection effect of network security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-04-14
AI Technical Summary
Existing cybersecurity defense methods in the power and industrial control fields are insufficient in their ability to identify and handle unknown attacks, variant attacks, and multi-stage coordinated attacks. They also struggle to dynamically depict the operational status of services and the impact of cross-site linkages, leading to false alarms that trigger service interruptions or an inability to balance local optimization with global coordination.
A multi-agent collaborative counterfactual defense approach for cybersecurity is adopted. By deploying local defense agents and core collaborative defense agents in the power grid, network topology, business constraints and state definition parameters are constructed, local candidate defense actions are generated, and simulation evaluation is carried out in a digital twin environment to achieve real-time risk assessment and defense strategy optimization.
It enables precise defense decisions against unknown attacks, reduces false alarms, improves the real-time and collaborative nature of network security protection, and ensures the continuity and operational security of critical businesses.
Smart Images

Figure CN121864376A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of network and information security technology, and more specifically, relates to a counterfactual defense method for power grid network security based on multi-agent collaboration. Background Technology
[0002] As the informatization and networking of power systems, industrial internet, and various critical infrastructures continue to improve, dispatch control systems, production control systems, and management information systems achieve data interaction and business collaboration through private networks, metropolitan area networks, and even cross-domain interconnections. While improving business efficiency and automation levels, the attack surface has also expanded significantly. Threats such as ransomware attacks, supply chain attacks, lateral movement penetration, and sophisticated attacks targeting industrial control protocols are characterized by high frequency, strong concealment, and complex propagation paths. Inadequate protection measures can easily lead to widespread business disruptions or even power grid operation accidents. Therefore, cybersecurity defense must not only focus on the attacks themselves but also consider their impact on critical business operations and power grid operational safety.
[0003] The existing network security defense methods in the power and industrial control fields mainly include the following: The first is the static defense method based on boundary isolation and access control, which usually restricts the scope of communication and access through means such as internal and external network isolation, whitelist policies, ACL / firewall rules, and one-way isolation devices; The second is the intrusion detection and alarm linkage method based on feature matching, which usually uses intrusion detection / prevention systems (IDS / IPS), rule bases or feature bases to identify abnormal traffic, malicious behavior and known attack patterns, and performs linkage actions such as blocking, rate limiting or isolation after alarms are triggered; The third is the security operation method based on situational awareness and centralized orchestration, which usually aggregates logs, alarms and traffic data on the central side, relies on the security operation platform or orchestration system to perform correlation analysis on events, and issues unified handling policies; The fourth is the protection method based on simulation or contingency plan drills, which usually relies on offline simulation, attack and defense drills and human experience to form handling plans, and executes defense actions according to the plans or operation and maintenance procedures in actual operation.
[0004] However, the aforementioned existing network security defense methods all have some significant drawbacks: First, existing static boundary isolation and access control methods lack the ability to dynamically depict the operational status of services, the evolution of attacks, and the impact of cross-site linkages, making it difficult to make timely, precise, and interpretable defense decisions when attacks spread rapidly or business topologies change frequently. Second, existing detection linkage methods based on feature matching are insufficient in their ability to identify and handle unknown attacks, variant attacks, and multi-stage collaborative attacks. Furthermore, the linkage actions are often mainly based on single-point blocking, which can easily lead to false alarms triggering business interruptions or suboptimal handling of "trading security for availability". Third, existing centralized situational awareness and orchestration methods usually rely on the central side to collect and distribute all data. In power and industrial control scenarios, they face constraints such as limited data collection, limited cross-domain collaboration, and high real-time requirements, which result in delays in strategy generation and distribution, and make it difficult to balance local optimization and global collaboration. Fourth, existing simulation and contingency plan exercises are mostly offline and manually driven, making it difficult to conduct online quantitative assessments of "what risk reduction and business impact will take a certain defensive action" when an actual attack occurs. They lack counterfactual comparisons and verifiable security margin determination mechanisms, making it difficult to form a joint defense plan that can effectively reduce risks and constrain negative business impacts. Summary of the Invention
[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a counterfactual defense method for power grid network security based on multi-agent collaboration. Its purpose is to solve the technical problems of existing static boundary isolation and access control methods, which lack the ability to dynamically characterize the operational status, attack evolution process, and cross-site linkage impact, making it difficult to achieve timely, precise, and interpretable defense decisions when attacks spread rapidly or service topologies change frequently. Furthermore, it addresses the shortcomings of existing feature-matching-based detection and linkage methods in identifying and handling unknown attacks, variant attacks, and multi-stage collaborative attacks, and the fact that linkage actions often focus on single-point blocking, easily leading to false alarms triggering service interruptions or suboptimal "security for availability" trade-offs. Technical issues include the fact that existing centralized situational awareness and orchestration methods typically rely on the central side to collect and distribute all data. In power and industrial control scenarios, these methods face constraints such as limited data collection, limited cross-domain collaboration, and high real-time requirements. This results in delays in strategy generation and distribution, and makes it difficult to balance local optimization with global collaboration. Furthermore, existing simulation and contingency planning methods are mostly offline and manually driven, making it difficult to conduct online quantitative assessments of "what risk reduction and business impact will take a certain defensive action" when an actual attack occurs. There is a lack of counterfactual comparison and verifiable security margin determination mechanisms, making it difficult to form a joint defense solution that can effectively reduce risks and constrain negative business impacts.
[0006] To achieve the above objectives, according to one aspect of the present invention, a network security counterfactual defense method based on multi-agent collaboration is provided, comprising the following steps: (1) Obtain the initial state information from network equipment, business systems and security audit equipment in the power grid, and construct a basic modeling result set based on the initial information. ; (2) Obtain the local raw data within time period t from the site to which each node belongs in the original network topology data, and based on the set of basic modeling results obtained in step (1). Construct a family of local candidate defense actions for time period t using all local raw data within time period t. ; (3) The set of local candidate defense actions obtained from step (2) within time period t Obtain site-level candidate actions and construct a set of joint defense schemes for time period t based on these site-level candidate actions. ; (4) The set of joint defense schemes within time period t obtained from steps (3-4) Obtain a joint defense scheme, and based on this joint defense scheme, obtain a set of utility evaluation results for time period t. ; (5) Set of joint defense scheme effectiveness evaluation results obtained in step (4-5) within time period t. Construct a set of secure and feasible joint defense solutions for time period t. And based on this set of secure and feasible joint defense schemes To obtain the final joint defense plan.
[0007] Preferably, step (1) includes the following sub-steps: (1-1) Obtain raw network topology data from network devices, business systems, and security auditing devices, and preprocess the obtained raw network topology data to obtain a set of network topology data. ; (1-2) Based on the network topology data set The node list and link list in the data are used to construct node sets respectively. and link set The network topology data set is constructed using an adjacency list-based graph construction method. The corresponding site and node information are mapped to a site set. and the subset of nodes within the i-th site Finally, the node set With Link Set Combine them to obtain the network topology diagram. And the station with the most adjacent nodes in the graph is used as the central station, where i∈[1,J]; (1-3) Obtain the original business configuration data from the database of the business system, and obtain the business flow set based on the obtained original business configuration data. and business constraint parameter set ; (1-4) Obtain historical operation data and raw attack data from the database of the security audit equipment, and obtain the set of intelligent agent design parameters based on the obtained historical operation data and raw attack data. and state definition parameter set The historical operation data includes low-risk log and alarm data, while the original attack data includes medium- or high-risk scenario data. Both the historical operation data and the original attack data contain the source and destination assets of the alarm, alarm type, danger level, link status, service type, and attack fields. (1-5) The attack feature vector and the agent design parameter set obtained in step (1-4) are respectively used. State definition parameter set and site index collection Obtain the input feature vectors, link state components, service state components, and attack state components required for agent configuration. Based on the obtained input feature vectors, link state components, service state components, and attack state components, obtain a set of local defense agents. and core collaborative defense intelligent agents ,in This represents the i-th local defense agent in the set of local defense agents; (1-6) Network topology diagram obtained from step (1-2) The business flow set obtained in steps (1-3) With business constraint parameter set The set of intelligent agent design parameters obtained in steps (1-4) and state definition parameter set The set of local defense agents obtained in steps (1-5) and core collaborative defense intelligent agents Construct a basic modeling result set .
[0008] Preferably, step (1-1) specifically involves obtaining multiple raw network topology data by polling the interfaces of network devices and reading the databases of business systems and security audit devices. The device list and link connection relationships in the raw network topology data are then converted into node lists and link lists, respectively. The site identifier of each node in the raw network topology data is obtained. All node lists, link lists, and site identifiers constitute the preprocessed network topology data set. ; The raw data for service configuration includes communication service configuration information, as well as procedure and engineering configuration information; Steps (1-3) are as follows: First, obtain communication service configuration information from the database of the business system, and parse the communication service configuration information to obtain the service flow set composed of the source node, destination node, and protocol information of the service configuration. Then, the latency thresholds for all business flows are obtained from the procedures and engineering configuration information stored in the database. Packet loss rate threshold and business importance weight The set of business constraint parameters constituted ; Steps (1-4) specifically involve: first, retrieving historical operational data and raw attack data from the database of the security audit equipment; then, normalizing and vectorizing the historical operational data and raw attack data to obtain multiple attack feature vectors, each of which is... ,in It is the dimension of the attack feature vector. Represents the first element in the attack feature vector. Each attack feature vector is first identified, and then k-means clustering is performed on all obtained attack feature vectors to obtain multiple security event clusters. Next, the cluster center vector corresponding to each security event cluster is obtained. The difference between the maximum and minimum values of each feature dimension across all cluster center vectors in the historical running data and the original attack data is calculated. All feature dimensions whose individual differences account for more than 25% of the total sum of differences are constructed into a feature index set. , feature index set Its corresponding normalization parameter pair Organized into a set of intelligent agent design parameters Subsequently, a vector set is constructed based on historical operational data and original attack data. The feature components of all vectors in this vector set are classified according to their physical meaning. That is, the indices of the link state components representing the link state in this vector set constitute a link state index set. The indices of the business state components representing the business state in this vector set constitute the business state index set. The indices of the attack state components representing the attack states in this vector set constitute the attack state index set. These three elements together constitute the state definition parameter set. ; The following formula is used to normalize and vectorize historical operational data to obtain attack feature vectors: ; Where m∈[1, the total number of historical running data obtained], Let m be the attack feature vector corresponding to the m-th historical running data. For attack feature mapping function, For the m-th historical running data, The attack type to which the m-th historical data belongs. For the attack severity level feature of the m-th historical data, For the m-th historical data segment, based on the source and destination assets in the network topology diagram The quantity characteristics of affected assets are calculated based on their positional relationships. This field represents the number of alarm triggers in the m-th historical data segment. The normalized alarm trigger count characteristic obtained through mean-variance normalization. This indicates the attack duration field in the m-th historical data segment. The normalized attack duration feature is obtained after mean-variance normalization, and the calculation methods for both are as follows: ; ; in For the m-th historical running data The number of alarms triggered within the preset statistics window. For the m-th historical running data The duration of the attack, and The number of alarms triggered are respectively The mean and standard deviation over historical operating data, and The duration of the attack The mean and standard deviation of historical operating data.
[0009] Preferably, step (1-5) specifically involves first, setting up the intelligent agent design parameters. The corresponding index belongs to the feature index set. The features of each dimension are processed according to the mean-variance normalization rule and concatenated in index order to obtain the first feature. The input feature vector of a local defense agent : ; in, Indicates the first The attack feature vector at the th ... Values in a dimension and , respectively, are the mean and standard deviation of the d-th dimension of the attack feature vector corresponding to the historical running data and the original attack data; The input feature vector of a local defense agent Input feature dimension And d∈[1,D]; Then, define the state parameter set. Link-state index set Business status index set and attack state index set Concatenate them sequentially to obtain the local defense agent for the i-th site. Indexes of the required link status, service status, and attack status components; Subsequently, a shallow neural network model is used to develop the local defense agent for the i-th site. Perform offline training to obtain the risk assessment parameter vector for the i-th site. and bias parameters ; Finally, the feature index set Normalized parameters and the risk assessment parameter vector of the i-th site and bias parameters Combining local defense agents for the i-th site Among all the sites, the local defense agent corresponding to the central site is used as the core collaborative defense agent CDA; Steps (1-6) specifically involve first defining the parameter set based on the state. Link-state index set From the network topology diagram Link set Select all link state index sets The links associated with one or more indices are used to form the set of links to be monitored. Obtain the j-th link from the network device. Link utilization Average latency Packet loss rate monitoring value The link utilization, average latency, and packet loss rate monitoring values of all links are concatenated into a link state vector. Then, define the parameter set according to the state. Business status index set From business flow set Select the set with the highest business importance weight and business status index. The ten business flows associated with the index are used as a set of key business flows. Obtain the set of key business flows Each key business flow end-to-end delay End-to-end packet loss rate monitoring value and set up key business flows All corresponding end-to-end latency and end-to-end packet loss rate monitoring values are concatenated sequentially according to the service flow number to form a service state vector. Subsequently, the parameter set is defined according to the state. attack state index set The index counts and statistically analyzes relevant features in historical execution data and raw attack data, and concatenates them in a fixed order to form an attack state vector. ; Finally, define the parameter set based on the state. Link state vector Business state vector and attack state vector Concatenate into a joint state vector Then, the network topology diagram is drawn. Business Flow Set Business constraint parameter set Local defense intelligent agent set Core collaborative defense intelligent agent and the resulting joint state vector Integrate the results to obtain a set of basic modeling results. .
[0010] Preferably, step (2) includes the following sub-steps: (2-1) Obtain the local raw data set of the i-th site, which consists of traffic logs, host logs and security alarm information within the t-period, from the i-th site to which the i-th node belongs in the original network topology data. For the acquired local raw data set Perform format processing to obtain the normalized local raw data set family of the i-th site within time period t. ); (2-2) The set of basic modeling results obtained in step (1-6) By integrating the network topology graph G and the service flow set F, the set of locally controllable links for the i-th site is obtained. and business link set and the resulting set of locally controllable links and business link set Discretize and merge to obtain the local action space set of the i-th station. ; (2-3) The normalized local raw data set of the i-th site obtained in step (2-1) Feature extraction is performed to obtain the local security feature vector of the i-th site within time period t. ; (2-4) The local security feature vector of the i-th site obtained in step (2-3) Perform local risk assessment to obtain the local risk score of the i-th site within time period t. ; (2-5) The normalized local raw data set of the i-th station obtained in step (2-1) within time period t. Perform service quality statistical processing to obtain the service quality vector of the i-th site within time period t. ; (2-6) The local security feature vector of the i-th site obtained in step (2-3) within time period t. The service quality vector of the i-th site obtained in step (2-5) within time period t and the local action space of the i-th site obtained in step (2-2). All candidate actions within the time period t are processed for action effect estimation to obtain the effect of the i-th station on the k-th candidate action within the time period t. Risk reduction estimate Impact of business quality on estimated value ; (2-7) The local risk score of the i-th site obtained in step (2-4) within time period t. Compared with the i-th station obtained in step (2-6) for the k-th candidate action Risk reduction estimate Impact of business quality on estimated value Comparison processing is performed to obtain a family of local candidate defense actions within time period t. .
[0011] Preferably, step (2-3) specifically involves processing the normalized local raw data set of the i-th station within time period t. The system counts and ratios of traffic logs, host logs, and security alerts. It then concatenates security features such as local alert counts, abnormal connection ratios, abnormal login counts, suspicious port access ratios, and average connection latency in a fixed order to obtain the local security feature vector of the i-th site within time period t. : ; in, For the site The number of local alarms within time period t. For the site The proportion of abnormal connections during time period t Let be the number of abnormal logins to the i-th site within time period t. Let be the percentage of suspicious port accesses to the i-th site during time period t. Let be the average connection latency of the i-th site during time period t. Let be the local security feature vector of the i-th site during time period t; Step (2-4) specifically involves using the local security feature vector of the i-th site. Input local defense agent Using a pre-defined linear scoring and non-linear activation structure, the local risk score of the i-th site during time period t is calculated. : ; in, Let i be the risk assessment weight vector for the i-th site. Let be the bias parameter for the i-th station. For the Sigmoid function, Give the local risk score for the i-th site during time period t; Steps (2-5) specifically involve first calculating the local raw data set of the i-th site within time period t. The traffic log records the average of all critical link utilization, average latency of all critical service flows, and average packet loss rate of all critical service flows. Then, the critical link utilization, average latency of critical service flows, and average packet loss rate are sorted according to their average values to obtain the service quality vector of the i-th site within time period t. ; Steps (2-6) specifically involve, firstly, defining the local action space of the i-th site. The k-th candidate action Encode the feature vector of the k-th candidate action at the i-th station. Then, the local security feature vector of the i-th site during time period t is... The obtained k-th candidate action feature vector of the i-th station After concatenation, input the data into a linear estimation model to obtain the relationship between the i-th station and the k-th candidate action. Risk reduction estimate Simultaneously, the service quality vector of the i-th site during time period t is... With the obtained candidate action feature vector After concatenation, input another linear estimation model to obtain the i-th station for the k-th candidate action. Business quality impact estimate : ; ; in, To estimate the parameter vector for risk reduction at the i-th site, Let be the parameter vector for estimating the service quality impact of the i-th site. This is the concatenated vector of the local security feature vector of the i-th site and the feature vector of the k-th candidate action. It is the concatenated vector of the service quality vector of the i-th station and the feature vector of the k-th candidate action; Step (2-7) specifically involves first comparing the local risk score of the i-th site with the local risk threshold. If the risk level is greater than a threshold, the site is considered a high-risk site; otherwise, it is considered a low-risk site. Then, the i-th site is used as a candidate action. Risk reduction estimate Impact of business quality on estimated value By comparison, those with a risk reduction estimate greater than the risk reduction threshold are selected. Furthermore, the estimated impact of business quality on the loss is not lower than the preset lower limit. All candidate actions are selected, and finally, all the selected candidate actions are combined to obtain a set of local candidate defense actions for all sites within time period t. .
[0012] Preferably, step (3) includes the following sub-steps: (3-1) Based on the site index set The set of local candidate defense actions within time period t obtained in step (2) Get the local candidate defense action set for the i-th site. The local candidate defense action sets of all obtained sites are merged to obtain the global candidate action set for time period t. ; (3-2) For the global candidate action set within time period t obtained in step (3-1) For each candidate action, the local risk score of the i-th site during time period t is obtained by combining steps (2-3). and the i-th station with respect to the k-th candidate action obtained in steps (2-5). Risk reduction estimate Impact of business quality on estimated value A comprehensive scoring process is performed to obtain the comprehensive score set of the candidate action within time period t. ; Specifically, this step involves first obtaining the global candidate action set within time period t. Each candidate action in The associated first The first site, then obtain the first Local risk score for each site and regarding candidate actions Risk reduction estimate Impact of business quality on estimated value Then, the three factors are weighted and summed according to preset weights to obtain the result for the i-th station in relation to the k-th candidate action. Overall rating Finally, the combined scores of all stations for all candidate actions are merged into a combined score set for candidate actions. : ; in, The three factors are weighted coefficients for the overall score, and their sum is 1. The value is 0.4. The value is 0.3. The value is 0.3; To obtain the maximum of the two; (3-3) The comprehensive score set of candidate actions obtained in step (3-2) Sorting and truncation are performed to obtain a set of selected candidate actions within time period t. ; Specifically, this step involves first, obtaining a comprehensive score set for the candidate actions. All candidate actions are sorted by their comprehensive scores from largest to smallest, and then the top-ranked actions in the sorting results are obtained. Each candidate action corresponds to a specific candidate action. Finally, all candidate actions are combined into a set of selected candidate actions for the time period t. ,in This is the preset maximum number of selected actions; (3-4) The selected candidate action set within time period t obtained in step (3-3) A combined construction process is performed to obtain a set of joint defense schemes for time period t. ; Specifically, this step involves first using a greedy algorithm to traverse the set of selected candidate actions within time period t. To obtain a comprehensive score From the top 30% of candidate actions, all the resulting candidate actions are then combined according to a preset strategy to form a set of joint defense schemes for time period t. .
[0013] Preferably, step (4) includes the following sub-steps: (4-1) The set of joint defense schemes for time period t obtained from step (3-4) Obtain the r-th joint defense scheme For the r-th joint defense scheme obtained Perform control sequence mapping to obtain the set of defense control sequences for time period t. and baseline control sequence set Where r∈[1, the total number of joint defense schemes]; (4-2) The set of defense control sequences for executing the r-th joint defense scheme within time period t obtained in step (4-1). and baseline control sequence set and the set of basic modeling results obtained in steps (1-6). Joint state vector Perform joint simulation to obtain the set of state trajectories for executing the r-th joint defense scheme within time period t. The set of state trajectories for not implementing the joint defense plan ; (4-3) For each of the state trajectory sets obtained in step (4-2) during time period t, execute the r-th joint defense scheme. The set of attack state variables and state trajectories that do not execute the joint defense scheme. The attack status variables are aggregated using time-weighted and node-weighted methods to obtain the attack risk indicator set for the execution of the r-th joint defense scheme within time period t. The set of attack risk indicators for when the joint defense plan is not implemented during time period t. ; (4-4) The set of state trajectories for executing the r-th joint defense scheme during time period t obtained in step (4-2). The latency and packet loss rate of critical business flows are aggregated using time-weighted and business-weighted methods to obtain a set of business / operational security indicators for executing the r-th joint defense scheme within time period t. And the set of state trajectories for which the joint defense scheme is not implemented during time period t. The latency and packet loss rate of critical business flows are aggregated using time-weighted and business-weighted methods to obtain a set of business / operational security indicators for scenarios where the joint defense scheme is not executed within time period t. ; (4-5) Execute the attack risk index of the r-th joint defense scheme within the time period t obtained in step (4-3). A set of attack risk indicators for failure to implement joint defense strategies And the business / operational security indicators for implementing the r-th joint defense scheme within time period t obtained in step (4-4). Business / Operational Security Indicators for Failure to Implement Joint Defense Plan Differential processing is performed to obtain the net defense effect of executing the r-th joint defense scheme within time period t. Net business impact And the net defense effectiveness of all joint defense schemes within the obtained time period t. Net business impact and joint defense plan The combination is the set of joint defense strategy effectiveness evaluation results within time period t. .
[0014] Preferably, step (4-1) specifically involves first, setting the r-th joint defense scheme... The site-level candidate actions of the i-th site are discretized in time to obtain a candidate action list. Then, the obtained candidate action list is mapped using one-hot encoding to obtain the defense control vector sequence of the i-th site executing the r-th joint defense scheme within time period t. and baseline control vector sequence Then, the defense control vector sequences of all stations executing the r-th joint defense scheme within time period t are combined into a set of defense control sequences executing the r-th joint defense scheme within time period t. The baseline control vector sequences of all stations are combined into a set of baseline control sequences for time period t. ; Step (4-2) specifically involves, under the same joint state vector In a digital twin environment, the defense control vector sequence for executing the r-th joint defense scheme within time period t is respectively... and baseline control vector sequence Discrete-time simulations are performed, and the system state trajectory of each joint defense scheme is recorded throughout the simulation process to obtain the set of state trajectories for executing the r-th joint defense scheme within time period t. The set of state trajectories for not implementing the joint defense plan ; Step (4-3) specifically involves, for the r-th joint defense scheme By performing the same normalization and vectorization processing as in step (1) on the newly generated running data and attack data of each node in its state trajectory, a new attack state vector is obtained. Then, the modulo of the obtained attack state vector is taken to obtain the attack state indicator. Finally, the obtained attack state indicator is accumulated according to the time and the proportion of key business flow of the node to obtain the attack risk indicator of the execution of the r-th joint defense scheme in time period t. Attack risk indicators for not implementing the joint defense plan during time period t : ; ; in, and They are nodes Executing the r-th joint defense plan and failure to implement joint defense plan time The attack status indicator, and has ∈[1, ], ∈[1, t]; For nodes The proportion of key business flows; For a moment The time decay value, and equal to: ; in, Let be the set of discrete times within time interval t. Let t be the end time point of the time period. The time decay coefficient, This indicates that no time discounting is applied, and all time points are weighted equally. The larger the value, the stronger the discount on time. 0.3, ∈ ; In step (4-4), the operational security indicators for executing the r-th joint defense scheme within time period t are obtained. and business / operational security indicators that are not implemented in each plan The following formula is used: ; ; in, For critical business flows Importance weight; and At time respectively Execute the r-th joint defense plan The time delay and the failure to execute the r-th joint defense scheme Time delay, and At time respectively Execute the r-th joint defense plan Packet loss rate and failure to execute the r-th joint defense scheme Packet loss rate at that time; For a moment The time decay value is calculated in the same way as in step (4-3); For business flow The security utility function is specifically... ; in and These are the latency threshold and packet loss rate threshold in the set of business constraint parameters; In step (4-5), the net defense effect of executing the r-th joint defense scheme within time period t is obtained. Net business impact The formula is as follows: ; ;
[0015] Preferably, step (5) includes the following sub-steps: (5-1) Set of joint defense scheme utility evaluation results obtained from step (4-5) within time period t We obtain all joint defense schemes with positive net defense effectiveness and net business impact, acquire the simulation trajectory corresponding to all links and all critical business flows under each joint defense scheme, and perform extreme value processing on the simulation trajectory to obtain the set of security margin indicators for all links and critical business flows within time period t. ; Specifically, this step involves first analyzing the set of joint defense strategy effectiveness evaluation results within time period t. First, identify all joint defense schemes where both net defense effectiveness and net business impact are positive. Then, obtain the latency and packet loss rate of each link in each joint defense scheme within time period t. Extract the maximum latency from all latency values and the maximum packet loss rate from all packet loss rates. Next, obtain all critical business flows included in the critical business flow set within each joint defense scheme, calculate the number of available paths for each critical business flow within time period t, and extract the minimum number of available paths from all available paths. Finally, combine the maximum latency, maximum packet loss rate, and minimum number of available paths as extreme value indicators to obtain the set of security margin indicators within time period t. ; (5-2) From the set of basic modeling results The maximum allowable latency and maximum allowable packet loss rate of each link are obtained, and the set of security margin indicators for the time period t obtained in step (5-1) is adjusted according to these maximum allowable latency and maximum allowable packet loss rate. Filtering is performed to obtain multiple joint defense schemes corresponding to multiple security margin indicators within time period t, which are then used as a set of feasible joint defense schemes. ; Specifically, this step involves, firstly, analyzing the set of basic modeling results. The maximum allowable latency and maximum allowable packet loss rate for each link are obtained, and the set of security margin indicators for the obtained time period t is adjusted based on these maximum allowable latency and maximum allowable packet loss rate. Filtering is performed, which involves checking whether the maximum link latency and maximum packet loss rate in each combination of the security margin indicator set do not exceed the maximum allowable latency and packet loss rate. Simultaneously, the results are filtered from the basic modeling result set. The process involves obtaining a set of business flows and checking whether the minimum number of available paths for critical business flows in each combination within the security margin indicator set meets the business requirements based on the number of services in the business flow set. Finally, the joint defense schemes corresponding to all combinations that satisfy all the above constraints are integrated to obtain a set of secure and feasible joint defense schemes for time period t. ; (5-3) Obtain the set of safe and feasible joint defense schemes within time period t obtained in step (5-2). The net defense effect and net business impact of each secure and feasible joint defense scheme are calculated, and the two are added together. The secure and feasible joint defense scheme with the largest sum is selected as the final joint defense scheme.
[0016] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: 1. Because the present invention adopts steps (1) to (2), it collects operation and alarm data from power secondary equipment, network equipment and security equipment and constructs a set of network topology, business constraints and state definition parameters. At the same time, it forms a state representation and local risk assessment mechanism that can be used for real-time perception at the site side. Therefore, it can solve the technical problem that the existing static boundary isolation and access control methods lack the ability to dynamically depict the business operation status, attack evolution process and cross-site linkage impact, and are difficult to achieve timely and precise defense decisions. 2. Since the present invention adopts steps (1) to (4), it forms an attack situation feature representation and local candidate action generation based on historical operation and security event statistics, and compares and evaluates the risk reduction and business impact of candidate actions under the framework of digital twin and counterfactual assessment. It does not rely on a fixed feature library or single rule matching. Therefore, it can solve the technical problems of existing feature matching-based detection linkage methods being insufficient in identifying and handling unknown attacks, variant attacks and multi-stage collaborative attacks, and easily causing false alarms that lead to business interruption. 3. Since the present invention adopts steps (2) to (3), it adopts a multi-agent collaborative mechanism of "local generation and core collaboration". The local defense agents of each site complete the candidate action screening and summary reporting locally, while the core collaborative defense agents only combine and coordinate the candidate actions. This avoids the time delay and cross-domain limitation problems caused by the central side's full data aggregation and centralized real-time reasoning. Therefore, it can solve the technical problems of existing centralized situational awareness and orchestration methods that rely on full data aggregation, have time delays in strategy generation and distribution, and are difficult to balance local optimization and global collaboration. 4. Because the present invention employs steps (4) to (5), it constructs the comparison trajectories of "executing the joint defense scheme" and "not executing the joint defense scheme" in the same simulation window, calculates the security margin index, and further constructs the set of operational security envelope parameters for boundary judgment and screening, thereby realizing online quantification and verifiable constraints on "whether the risk reduction is positive and whether the negative impact on business is under control". Therefore, it can solve the technical problems that existing simulation and simulation or contingency plan exercise methods are mostly offline and manually driven, lack counterfactual comparison and verifiable security margin judgment, and are difficult to form a joint defense scheme that takes into account both risk reduction and business availability.
[0017] 5. The counterfactual defense method and system for network security based on multi-agent collaboration proposed in this invention can achieve assessable, constrainable, and optimizable joint defense decisions in complex multi-site environments, and significantly improve the overall effectiveness and intelligence level of network security protection while ensuring the continuity of critical business and operational security. Attached Figure Description
[0018] Figure 1 This is a block diagram of the overall structure of the network security counterfactual defense method based on multi-agent collaboration of the present invention; Figure 2 This is a flowchart illustrating the counterfactual defense method for network security based on multi-agent collaboration according to the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0020] It should be noted that in the description of the embodiments of the present invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The terms "upper," "lower," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. Those skilled in the art can understand the specific meaning of the above terms in the present invention according to the specific circumstances.
[0021] Furthermore, the technical solutions of the various embodiments of the present invention can be combined with each other, but only if they are feasible for those skilled in the art. If the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.
[0022] The basic idea of this invention is to provide a counterfactual defense method for network security based on multi-agent collaboration. This method deploys local defense agents with autonomous decision-making capabilities in different areas of the network and sets up a core defense agent for global collaboration, thereby achieving unified coordination and control of multi-point linked security strategies and realizing cross-regional and cross-security domain collaborative defense. This invention is particularly applicable to scenarios in power secondary systems, industrial control systems, industrial internet, and other critical infrastructure networks where joint analysis and proactive response to multi-site and multi-zone network attacks are required. It can also be extended to network environments with complex topologies and strict business continuity requirements, such as cloud computing data centers and government private networks. The key technologies involved in this invention include multi-agent collaborative decision-making, network security situation assessment, power grid / industrial control business characteristic modeling, digital twin simulation, and defense strategy optimization based on counterfactual assessment, and can be categorized into the fields of network security protection, intelligent network operation and maintenance, and integrated application technology of industrial information security.
[0023] like Figure 1 and Figure 2 As shown, this invention provides a network security counterfactual defense method based on multi-agent collaboration, comprising the following steps: (1) Obtain the initial state information from network equipment, business systems and security audit equipment in the power grid, and construct a basic modeling result set based on the initial information. ; The purpose of this step is to obtain raw information from network devices, business systems, and security auditing devices in the power grid, including raw network topology data, raw business configuration data, historical operational data, and raw attack data, and to construct a basic modeling result set based on this raw information. This provides unified network and service modeling results, as well as multi-agent and joint state definitions, for subsequent steps.
[0024] This step includes the following sub-steps: (1-1) Obtain raw network topology data from network devices, business systems, and security auditing devices, and preprocess the obtained raw network topology data to obtain a set of network topology data. ; Specifically, this step involves obtaining multiple sets of raw network topology data by polling network device interfaces and reading databases from business systems and security audit devices. The device lists and link connections in the raw network topology data are then converted into node lists and link lists, respectively. The site identifier for each node in the raw network topology data is also obtained. All node lists, link lists, and site identifiers constitute the preprocessed network topology data set. ; (1-2) Based on the network topology data set The node list and link list in the data are used to construct node sets respectively. and link set The network topology data set is constructed using an adjacency list-based graph construction method. The corresponding site and node information are mapped to a site set. and the subset of nodes within the i-th site Finally, the node set With Link Set Combine them to obtain the network topology diagram. And the station with the most adjacent nodes in the graph is used as the central station, where i∈[1,J]; (1-3) Obtain the original business configuration data from the database of the business system, and obtain the business flow set based on the obtained original business configuration data. and business constraint parameter set ; Specifically, the raw data for service configuration includes configuration information for communication services (including key services such as scheduling, protection, measurement and control, and remote control), as well as configuration information for procedures and engineering.
[0025] Specifically, this step involves first retrieving communication service configuration information from the business system's database, then parsing this information to obtain a set of service flows consisting of the source node, destination node, and protocol information of the service configuration. Then, the latency thresholds for all business flows are obtained from the procedures and engineering configuration information stored in the database. Packet loss rate threshold and business importance weight The set of business constraint parameters constituted .
[0026] (1-4) Obtain historical operational data and raw attack data from the database of the security audit equipment (historical operational data includes low-risk log and alarm data, and raw attack data includes medium- or high-risk scenario data; both historical operational data and raw attack data contain the source and destination assets of the alarms, alarm type, danger level, link status, service type, and attack fields), and obtain the intelligent agent design parameter set based on the obtained historical operational data and raw attack data. and state definition parameter set ; Specifically, this step involves first retrieving historical operational data and raw attack data from the security audit equipment's database. Then, the historical operational data and raw attack data are normalized and vectorized to obtain multiple attack feature vectors. Each attack feature vector is... ,in It is the dimension of the attack feature vector. Represents the first element in the attack feature vector. Each attack feature vector is first identified, and then k-means clustering is performed on all obtained attack feature vectors to obtain multiple security event clusters. Next, the cluster center vector corresponding to each security event cluster is obtained. The difference between the maximum and minimum values of each feature dimension across all cluster center vectors in the historical running data and the original attack data is calculated. All feature dimensions whose individual differences account for more than 25% of the total sum of differences are constructed into a feature index set. , feature index set Its corresponding normalization parameter pair Organized into a set of intelligent agent design parameters Subsequently, a vector set is constructed based on historical operational data and original attack data. The feature components of all vectors in this vector set are classified according to their physical meaning. That is, the indices of the link state components representing the link state in this vector set constitute a link state index set. The indices of the business state components representing the business state in this vector set constitute the business state index set. The indices of the attack state components representing the attack states in this vector set constitute the attack state index set. These three elements together constitute the state definition parameter set. ; More specifically, the normalization and vectorization of historical operational data to obtain attack feature vectors is performed using the following formula (the process of normalizing and vectorizing the original attack data is exactly the same): ; Where m∈[1, the total number of historical running data obtained], Let m be the attack feature vector corresponding to the m-th historical running data. For attack feature mapping function, For the m-th historical running data, The attack type to which the m-th historical data belongs (represented by binary encoding from 0 to 14, such as port scanning, brute force, denial-of-service attack, SQL injection, cross-site scripting attack, command injection, buffer overflow attack, man-in-the-middle attack, ARP spoofing, DNS spoofing, phishing attack, malware delivery, privilege escalation attack, lateral movement, and backdoor control). The attack severity level feature of the m-th historical data (divided into high-risk, medium-risk, and low-risk levels as integers 1, 2, and 3). For the m-th historical data segment, based on the source and destination assets in the network topology diagram The quantity characteristics of affected assets are calculated based on their positional relationships. This field represents the number of alarm triggers in the m-th historical data segment. The normalized alarm trigger count characteristic obtained through mean-variance normalization. This indicates the attack duration field in the m-th historical data segment. The normalized attack duration feature is obtained after mean-variance normalization, and the calculation methods for both are as follows: ; ; in For the m-th historical running data The number of alarms triggered within the preset statistics window. For the m-th historical running data The duration of the attack, and The number of alarms triggered are respectively The mean and standard deviation over historical operating data, and The duration of the attack The mean and standard deviation of historical operating data.
[0027] (1-5) The attack feature vector and the agent design parameter set obtained in step (1-4) are respectively used. State definition parameter set and site index collection Obtain the input feature vectors, link state components, service state components, and attack state components required for agent configuration. Based on the obtained input feature vectors, link state components, service state components, and attack state components, obtain a set of local defense agents. and core collaborative defense intelligent agents ,in This represents the i-th local defense agent in the set of local defense agents; Specifically, this step involves first setting up the intelligent agent design parameters. The corresponding index belongs to the feature index set. The features of each dimension are processed according to the mean-variance normalization rule and concatenated in index order to obtain the first feature. The input feature vector of a local defense agent : ; in, Indicates the first The attack feature vector at the th ... Values in a dimension and , respectively, are the mean and standard deviation of the d-th dimension of the attack feature vector corresponding to the historical running data and the original attack data; The input feature vector of a local defense agent Input feature dimension And d∈[1,D]; Then, define the state parameter set. Link-state index set Business status index set and attack state index set Concatenate them sequentially to obtain the local defense agent for the i-th site. Indexes of the required link status, service status, and attack status components; Subsequently, a shallow neural network model is used to develop the local defense agent for the i-th site. Perform offline training to obtain the risk assessment parameter vector for the i-th site. and bias parameters ; Finally, the feature index set Normalized parameters and the risk assessment parameter vector of the i-th site and bias parameters Combining local defense agents for the i-th site Among all the sites, the local defense agent corresponding to the central site is used as the core collaborative defense agent CDA.
[0028] (1-6) Network topology diagram obtained from step (1-2) The business flow set obtained in steps (1-3) With business constraint parameter set The set of intelligent agent design parameters obtained in steps (1-4) and state definition parameter set The set of local defense agents obtained in steps (1-5) and core collaborative defense intelligent agents Construct a basic modeling result set ; Specifically, this step involves first defining a parameter set based on the state. Link-state index set From the network topology diagram Link set Select all link state index sets The links associated with one or more indices are used to form the set of links to be monitored. Obtain each link (i.e., the j-th link, where j∈[1, the total number of links in the network device]) from the network device. Link utilization Average latency Packet loss rate monitoring value The link utilization, average latency, and packet loss rate monitoring values of all links are concatenated into a link state vector. Then, define the parameter set according to the state. Business status index set From business flow set Select the set with the highest business importance weight and business status index. The ten business flows associated with the index are used as a set of key business flows. Obtain the set of key business flows Each key business flow end-to-end delay End-to-end packet loss rate monitoring value and set up key business flows All corresponding end-to-end latency and end-to-end packet loss rate monitoring values are concatenated sequentially according to the service flow number to form a service state vector. Subsequently, the parameter set is defined according to the state. attack state index set The index counts and statistically analyzes relevant features in historical execution data and raw attack data, and concatenates them in a fixed order to form an attack state vector. ; Finally, define the parameter set based on the state. Link state vector Business state vector and attack state vector Concatenate into a joint state vector Then, the network topology diagram... Business Flow Set Business constraint parameter set Local defense intelligent agent set Core collaborative defense intelligent agent and the resulting joint state vector Integrate the results to obtain a set of basic modeling results. .
[0029] (2) Obtain the local raw data within time period t from the site to which each node belongs in the original network topology data, and based on the set of basic modeling results obtained in step (1). Construct a family of local candidate defense actions for time period t using all local raw data within time period t. ; The purpose of this step is to: extract the results from the basic modeling dataset. The system acquires local raw data within the current time period t through a real-time acquisition system, and constructs a set of local candidate defense actions based on the local raw data. This provides site-level candidate action inputs for constructing a global joint defense scheme.
[0030] This step includes the following sub-steps: (2-1) Obtain the local raw data set of the i-th site, which consists of traffic logs, host logs and security alarm information within the t-period, from the i-th site to which the i-th node belongs in the original network topology data. For the acquired local raw data set Perform format processing (i.e., align the timestamp of each local raw data with the site identifier and remove records with abnormal formats) to obtain the normalized local raw data set family of the i-th site within time period t. ), where i∈[1, J]; (2-2) The set of basic modeling results obtained in step (1-6) By integrating the network topology graph G and the service flow set F, the set of locally controllable links for the i-th site is obtained. and business link set and the resulting set of locally controllable links and business link set Discretize and merge to obtain the local action space set of the i-th station. ; (2-3) The normalized local raw data set of the i-th site obtained in step (2-1) Feature extraction is performed to obtain the local security feature vector of the i-th site within time period t. ; Specifically, this step involves processing the normalized local raw data set of the i-th site within time period t. The system counts and ratios of traffic logs, host logs, and security alerts. It then concatenates security features such as local alert counts, abnormal connection ratios, abnormal login counts, suspicious port access ratios, and average connection latency in a fixed order to obtain the local security feature vector of the i-th site within time period t. .
[0031] Specifically, the local security feature vector of the i-th site during time period t is obtained. The formula is as follows: ; in, For the site The number of local alarms within time period t. For the site The proportion of abnormal connections during time period t Let be the number of abnormal logins to the i-th site within time period t. Let be the percentage of suspicious port accesses to the i-th site during time period t. Let be the average connection latency of the i-th site during time period t. Let be the local security feature vector of the i-th site during time period t.
[0032] (2-4) The local security feature vector of the i-th site obtained in step (2-3) Perform local risk assessment to obtain the local risk score of the i-th site within time period t. ; Specifically, this step involves using the local security feature vector of the i-th site... Input local defense agent Using a pre-defined linear scoring and non-linear activation structure, the local risk score of the i-th site during time period t is calculated. .
[0033] Specifically, the local risk score of the i-th site during time period t is obtained. The formula is as follows: ; in, Let i be the risk assessment weight vector for the i-th site. Let be the bias parameter for the i-th station. For the Sigmoid function, The local risk score for the i-th site during time period t.
[0034] (2-5) The normalized local raw data set of the i-th station obtained in step (2-1) within time period t. Perform service quality statistical processing to obtain the service quality vector of the i-th site within time period t. ; Specifically, this step involves first calculating the local raw data set of the i-th site within time period t. The traffic log records the average of all critical link utilization, average latency of all critical service flows, and average packet loss rate of all critical service flows. Then, the critical link utilization, average latency of critical service flows, and average packet loss rate are sorted according to their average values to obtain the service quality vector of the i-th site within time period t. .
[0035] (2-6) The local security feature vector of the i-th site obtained in step (2-3) within time period t. The service quality vector of the i-th site obtained in step (2-5) within time period t and the local action space of the i-th site obtained in step (2-2). All candidate actions within the time period t are processed for action effect estimation to obtain the effect of the i-th station on the k-th candidate action within the time period t. Risk reduction estimate Impact of business quality on estimated value ; Specifically, this step involves first, by analyzing the local action space of the i-th site. The k-th candidate action Encode the feature vector of the k-th candidate action at the i-th station. Then, the local security feature vector of the i-th site during time period t is... The obtained k-th candidate action feature vector of the i-th station After concatenation, input the data into a linear estimation model to obtain the relationship between the i-th station and the k-th candidate action. Risk reduction estimate Simultaneously, the service quality vector of the i-th site during time period t is... With the obtained candidate action feature vector After concatenation, input another linear estimation model to obtain the i-th station for the k-th candidate action. Business quality impact estimate .
[0036] Specifically, the estimated risk reduction in acquiring the i-th site And the i-th station for the k-th candidate action Business quality impact estimate The following formula is used: ; ; in, To estimate the parameter vector for risk reduction at the i-th site, Let be the parameter vector for estimating the service quality impact of the i-th site. This is the concatenated vector of the local security feature vector of the i-th site and the feature vector of the k-th candidate action. It is the concatenated vector of the service quality vector of the i-th station and the feature vector of the k-th candidate action; (2-7) The local risk score of the i-th site obtained in step (2-4) within time period t. Compared with the i-th station obtained in step (2-6) for the k-th candidate action Risk reduction estimate Impact of business quality on estimated value Comparison processing is performed to obtain a family of local candidate defense actions within time period t. ; Specifically, this step involves first comparing the local risk score of the i-th site with the local risk threshold. (Value range 0.5-0.9, preferably 0.7) Compare the values. If the value is greater than the threshold, the site is considered a high-risk site; otherwise, it is considered a low-risk site. Then, the i-th site is used as a candidate action. Risk reduction estimate Impact of business quality on estimated value By comparison, those with a risk reduction estimate greater than the risk reduction threshold are selected. (For high-risk sites, the threshold value ranges from 0.18 to 0.3, preferably 0.3; for low-risk sites, the threshold value ranges from 0.05 to 0.15, preferably 0.05.) Furthermore, the estimated impact on business quality is not lower than the preset lower limit of loss. (For high-risk sites, the lower limit of the loss ranges from 0.04 to 0.1, preferably 0.1; for low-risk sites, the lower limit of the loss ranges from 0.02 to 0.04, preferably 0.03.) All candidate actions are selected, and finally, all the selected candidate actions are combined to obtain a set of local candidate defense actions for all sites within time period t. ; The advantage of step (2) is that it simultaneously quantifies the current risk level, the expected risk and benefit of local defensive actions, and the cost to the business within a local scope of a single site, ensuring that subsequent collaborative defense is only implemented within a set of high-quality candidate actions. Combining these elements significantly reduces the complexity of joint decision-making.
[0037] (3) The set of local candidate defense actions obtained from step (2) within time period t Obtain site-level candidate actions and construct a set of joint defense schemes for time period t based on these site-level candidate actions. ; The purpose of this step is to: select the local candidate defense action set family within time period t. All site-level candidate actions are obtained, and a global comprehensive score and combination construction are performed based on the site-level candidate actions to obtain a set of joint defense schemes. This provides input for joint defense strategies in subsequent digital twin simulations.
[0038] This step includes the following sub-steps: (3-1) Based on the site index set The set of local candidate defense actions within time period t obtained in step (2) Get the local candidate defense action set for the i-th site. The local candidate defense action sets of all obtained sites are merged to obtain the global candidate action set for time period t. ; (3-2) For the global candidate action set within time period t obtained in step (3-1) For each candidate action, the local risk score of the i-th site during time period t is obtained by combining steps (2-3). and the i-th station with respect to the k-th candidate action obtained in steps (2-5). Risk reduction estimate Impact of business quality on estimated value A comprehensive scoring process is performed to obtain the comprehensive score set of the candidate action within time period t. ; Specifically, this step involves first obtaining the global candidate action set within time period t. Each candidate action in The associated first The first site, then obtain the first Local risk score for each site and regarding candidate actions Risk reduction estimate Impact of business quality on estimated value Then, the three factors are weighted and summed according to preset weights to obtain the result for the i-th station in relation to the k-th candidate action. Overall rating Finally, the combined scores of all stations for all candidate actions are merged into a combined score set for candidate actions. .
[0039] Specifically, obtain the i-th station for the k-th candidate action. Overall rating The following formula is used: ; in, The three factors are weighted coefficients for the overall score, and their sum is 1. The value is 0.4. The value is 0.3. The value is 0.3; To obtain the maximum value of the two.
[0040] (3-3) The comprehensive score set of candidate actions obtained in step (3-2) Sorting and truncation are performed to obtain a set of selected candidate actions within time period t. .
[0041] Specifically, this step involves first, obtaining a comprehensive score set for the candidate actions. All candidate actions are sorted by their comprehensive scores from largest to smallest, and then the top-ranked actions in the sorting results are obtained. Each candidate action corresponds to a specific candidate action. Finally, all candidate actions are combined into a set of selected candidate actions for the time period t. ,in The preset maximum number of selected actions, with a preferred value of 5; (3-4) The selected candidate action set within time period t obtained in step (3-3) A combined construction process is performed to obtain a set of joint defense schemes for time period t. ; Specifically, this step involves first using a greedy algorithm to traverse the set of selected candidate actions within time period t. To obtain a comprehensive score From the top 30% of candidate actions, all the resulting candidate actions are then combined according to a preset strategy to form a set of joint defense schemes for time period t. .
[0042] The advantage of this step (3) is that it allows for the selection of a set of local candidate defense actions within time period t. Starting from this point, and through comprehensive global evaluation and the construction of a limited-scale action combination, a set of joint defense schemes for time period t is formed. This approach takes into account multi-site collaboration while controlling the scale of the combined space, thus reducing the computational complexity for subsequent digital twin simulations.
[0043] (4) The set of joint defense schemes within time period t obtained from steps (3-4) Obtain a joint defense scheme, and based on this joint defense scheme, obtain a set of utility evaluation results for time period t. ; The purpose of this step is to: select from the set of joint defense strategies Each joint defense scheme is acquired, and realistic and counterfactual simulations are performed in a digital twin environment based on these schemes to obtain the attack risk and business / operational security indicators of each scheme. Furthermore, the net defense effectiveness and net business impact are calculated to construct a set of joint defense scheme utility evaluation results for time period t. .
[0044] This step includes the following sub-steps: (4-1) The set of joint defense schemes for time period t obtained from step (3-4) Obtain the r-th joint defense scheme For the r-th joint defense scheme obtained Perform control sequence mapping to obtain the set of defense control sequences for time period t. and baseline control sequence set Where r∈[1, the total number of joint defense schemes]; Specifically, this step involves first, configuring the r-th joint defense scheme. The site-level candidate actions of the i-th site are discretized in time to obtain a candidate action list. Then, the obtained candidate action list is mapped using one-hot encoding to obtain the defense control vector sequence of the i-th site executing the r-th joint defense scheme within time period t. and baseline control vector sequence Then, the defense control vector sequences of all stations executing the r-th joint defense scheme within time period t are combined into a set of defense control sequences executing the r-th joint defense scheme within time period t. The baseline control vector sequences of all stations are combined into a set of baseline control sequences for time period t. .
[0045] (4-2) The set of defense control sequences for executing the r-th joint defense scheme within time period t obtained in step (4-1). and baseline control sequence set and the set of basic modeling results obtained in steps (1-6). Joint state vector Perform joint simulation to obtain the set of state trajectories for executing the r-th joint defense scheme within time period t. The set of state trajectories for not implementing the joint defense plan ; Specifically, this step involves, under the same joint state vector... In a digital twin environment, the defense control vector sequence for executing the r-th joint defense scheme within time period t is respectively... and baseline control vector sequence Discrete-time simulations are performed, and the system state trajectory of each joint defense scheme is recorded throughout the simulation process to obtain the set of state trajectories for executing the r-th joint defense scheme within time period t. The set of state trajectories for not implementing the joint defense plan ; (4-3) For each of the state trajectory sets obtained in step (4-2) during time period t, execute the r-th joint defense scheme. The set of attack state variables and state trajectories that do not execute the joint defense scheme. The attack status variables are aggregated using time-weighted and node-weighted methods to obtain the attack risk indicator set for the execution of the r-th joint defense scheme within time period t. The set of attack risk indicators for when the joint defense plan is not implemented during time period t. ; Specifically, this step involves considering the r-th joint defense scheme. By performing the same normalization and vectorization processing as in step (1) on the newly generated running data and attack data of each node in its state trajectory, a new attack state vector is obtained. Then, the modulo of the obtained attack state vector is taken to obtain the attack state indicator. Finally, the obtained attack state indicator is accumulated according to the time and the proportion of key business flow of the node to obtain the attack risk indicator of the execution of the r-th joint defense scheme in time period t. Attack risk indicators for not implementing the joint defense plan during time period t .
[0046] Specifically, the attack risk index for executing the r-th joint defense scheme within time period t is obtained. The joint defense plan is not implemented during time period t. The formula is as follows: ; ; in, and They are nodes Executing the r-th joint defense plan and failure to implement joint defense plan time The attack status indicator, and has ∈[1, ], ∈[1, t]; For nodes The proportion of key business flows; For a moment The time decay value, and equal to: ; in, Let be the set of discrete times within time interval t. Let t be the end time point of the time period. The time decay coefficient, This indicates that no time discounting is applied, and all time points are weighted equally. The larger the value, the stronger the discount on time. 0.3, ∈ .
[0047] (4-4) The set of state trajectories for executing the r-th joint defense scheme during time period t obtained in step (4-2). The latency and packet loss rate of critical business flows are aggregated using time-weighted and business-weighted methods to obtain a set of business / operational security indicators for executing the r-th joint defense scheme within time period t. And the set of state trajectories for which the joint defense scheme is not implemented during time period t. The latency and packet loss rate of critical business flows are aggregated using time-weighted and business-weighted methods to obtain a set of business / operational security indicators for scenarios where the joint defense scheme is not executed within time period t. ; Specifically, it involves obtaining the operational security metrics for executing the r-th joint defense scheme within time period t. and business / operational security indicators that are not implemented in each plan The following formula is used: ; ; in, For critical business flows Importance weight; and At time respectively Execute the r-th joint defense plan The time delay and the failure to execute the r-th joint defense scheme Time delay, and At time respectively Execute the r-th joint defense plan Packet loss rate and failure to execute the r-th joint defense scheme Packet loss rate at that time; For a moment The time decay value is calculated in the same way as in step (4-3); For business flow The security utility function is specifically... ; in and These are the latency threshold and packet loss rate threshold, respectively, from the set of business constraint parameters.
[0048] (4-5) Execute the attack risk index of the r-th joint defense scheme within the time period t obtained in step (4-3). A set of attack risk indicators for failure to implement joint defense strategies And the business / operational security indicators for implementing the r-th joint defense scheme within time period t obtained in step (4-4). Business / Operational Security Indicators for Failure to Implement Joint Defense Plan Differential processing is performed to obtain the net defense effect of executing the r-th joint defense scheme within time period t. Net business impact And the net defense effectiveness of all joint defense schemes within the obtained time period t. Net business impact and joint defense plan The combination is the set of joint defense strategy effectiveness evaluation results within time period t. ; Specifically, the net defense effect of implementing the r-th joint defense scheme within time period t is obtained. Net business impact The formula is as follows: ; ; The advantage of this step (4) is that by conducting pairwise simulations of the "execution plan" and "non-execution plan" under the same attack and business scenario in the digital twin environment, a utility evaluation framework with counterfactual comparison significance is established, which enables the offensive and defensive benefits and business costs of the joint defense plan to be quantitatively compared.
[0049] (5) Set of joint defense scheme effectiveness evaluation results obtained in step (4-5) within time period t. Construct a set of secure and feasible joint defense solutions for time period t. And based on this set of secure and feasible joint defense schemes To obtain the final joint defense plan; The purpose of this step is to: extract the set of joint defense strategy effectiveness evaluation results for time period t. Obtain the evaluation results of each scheme, and base them on the set of pre-set operation safety envelope parameters in the power grid operation procedures. Constraints are applied to the simulated trajectory to obtain a set of safe and feasible joint defense schemes within time period t. This provides candidate solutions within a safe boundary for subsequent utility optimization.
[0050] This step includes the following sub-steps: (5-1) Set of joint defense scheme utility evaluation results obtained from step (4-5) within time period t We obtain all joint defense schemes with positive net defense effectiveness and net business impact, acquire the simulation trajectory corresponding to all links and all critical business flows under each joint defense scheme, and perform extreme value processing on the simulation trajectory to obtain the set of security margin indicators for all links and critical business flows within time period t. ; Specifically, this step involves first analyzing the set of joint defense strategy effectiveness evaluation results within time period t. First, identify all joint defense schemes where both net defense effectiveness and net business impact are positive. Then, obtain the latency and packet loss rate of each link in each joint defense scheme within time period t. Extract the maximum latency from all latency values and the maximum packet loss rate from all packet loss rates. Next, obtain all critical business flows included in the critical business flow set within each joint defense scheme, calculate the number of available paths for each critical business flow within time period t, and extract the minimum number of available paths from all available paths. Finally, combine the maximum latency, maximum packet loss rate, and minimum number of available paths as extreme value indicators to obtain the set of security margin indicators within time period t. .
[0051] (5-2) From the set of basic modeling results The maximum allowable latency and maximum allowable packet loss rate of each link are obtained, and the set of security margin indicators for the time period t obtained in step (5-1) is adjusted according to these maximum allowable latency and maximum allowable packet loss rate. Filtering is performed to obtain multiple joint defense schemes corresponding to multiple security margin indicators within time period t, which are then used as a set of feasible joint defense schemes. ; Specifically, this step involves, firstly, analyzing the set of basic modeling results. The maximum allowable latency and maximum allowable packet loss rate for each link are obtained, and the set of security margin indicators for the obtained time period t is adjusted based on these maximum allowable latency and maximum allowable packet loss rate. Filtering is performed, which involves checking whether the maximum link latency and maximum packet loss rate in each combination of the security margin indicator set do not exceed the maximum allowable latency and packet loss rate. Simultaneously, the results are filtered from the basic modeling result set. The process involves obtaining a set of business flows and checking whether the minimum number of available paths for critical business flows in each combination within the security margin indicator set meets the business requirements based on the number of services in the business flow set. Finally, the joint defense schemes corresponding to all combinations that satisfy all the above constraints are integrated to obtain a set of secure and feasible joint defense schemes for time period t. .
[0052] (5-3) Obtain the set of safe and feasible joint defense schemes within time period t obtained in step (5-2). The net defense effect and net business impact of each secure and feasible joint defense scheme are calculated, and the two are added together. The secure and feasible joint defense scheme with the largest sum is selected as the final joint defense scheme (if two sums are equal, the secure and feasible joint defense scheme with the better net defense effect is selected as the final joint defense scheme).
[0053] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A counterfactual defense method for network security based on multi-agent collaboration, characterized in that, Includes the following steps: (1) Obtain the initial state information from network equipment, business systems and security audit equipment in the power grid, and construct a basic modeling result set based on the initial information. ; (2) Obtain the local raw data within time period t from the site to which each node belongs in the original network topology data, and based on the set of basic modeling results obtained in step (1). Construct a family of local candidate defense actions for time period t using all local raw data within time period t. ; (3) The set of local candidate defense actions obtained from step (2) within time period t Obtain site-level candidate actions and construct a set of joint defense schemes for time period t based on these site-level candidate actions. ; (4) The set of joint defense schemes within time period t obtained from steps (3-4) Obtain a joint defense scheme, and based on this joint defense scheme, obtain a set of utility evaluation results for time period t. ; (5) Set of joint defense scheme effectiveness evaluation results obtained in step (4-5) within time period t. Construct a set of secure and feasible joint defense solutions for time period t. And based on this set of secure and feasible joint defense schemes To obtain the final joint defense plan.
2. The network security counterfactual defense method based on multi-agent collaboration according to claim 1, characterized in that, Step (1) includes the following sub-steps: (1-1) Obtain raw network topology data from network devices, business systems, and security auditing devices, and preprocess the obtained raw network topology data to obtain a set of network topology data. ; (1-2) Based on the network topology data set The node list and link list in the data are used to construct node sets respectively. and link set The network topology data set is constructed using an adjacency list-based graph construction method. The corresponding site and node information are mapped to a site set. and the subset of nodes within the i-th site Finally, the node set With Link Set Combine them to obtain the network topology diagram. And the station with the most adjacent nodes in the graph is used as the central station, where i∈[1,J]; (1-3) Obtain the original business configuration data from the database of the business system, and obtain the business flow set based on the obtained original business configuration data. and business constraint parameter set ; (1-4) Obtain historical operation data and raw attack data from the database of the security audit equipment, and obtain the set of intelligent agent design parameters based on the obtained historical operation data and raw attack data. and state definition parameter set The historical operation data includes low-risk log and alarm data, while the original attack data includes medium- or high-risk scenario data. Both the historical operation data and the original attack data contain the source and destination assets of the alarm, alarm type, danger level, link status, service type, and attack fields. (1-5) The attack feature vector and the agent design parameter set obtained in step (1-4) are respectively used. State definition parameter set and site index collection Obtain the input feature vectors, link state components, service state components, and attack state components required for agent configuration. Based on the obtained input feature vectors, link state components, service state components, and attack state components, obtain a set of local defense agents. and core collaborative defense intelligent agents ,in This represents the i-th local defense agent in the set of local defense agents; (1-6) Network topology diagram obtained from step (1-2) The business flow set obtained in steps (1-3) With business constraint parameter set The set of intelligent agent design parameters obtained in steps (1-4) and state definition parameter set The set of local defense agents obtained in steps (1-5) and core collaborative defense intelligent agents Construct a basic modeling result set .
3. The network security counterfactual defense method based on multi-agent collaboration according to claim 2, characterized in that, Step (1-1) specifically involves obtaining multiple sets of raw network topology data by polling the interfaces of network devices and reading the databases of business systems and security audit devices. The device lists and link connection relationships in the raw network topology data are then converted into node lists and link lists, respectively. The site identifier of each node in the raw network topology data is also obtained. All node lists, link lists, and site identifiers constitute the preprocessed network topology data set. ; The raw data for service configuration includes communication service configuration information, as well as procedure and engineering configuration information; Steps (1-3) are as follows: First, obtain communication service configuration information from the database of the business system, and parse the communication service configuration information to obtain the service flow set composed of the source node, destination node, and protocol information of the service configuration. Then, the latency thresholds for all business flows are obtained from the procedures and engineering configuration information stored in the database. Packet loss rate threshold and business importance weight The set of business constraint parameters constituted ; Steps (1-4) specifically involve: first, retrieving historical operational data and raw attack data from the database of the security audit equipment; then, normalizing and vectorizing the historical operational data and raw attack data to obtain multiple attack feature vectors, each of which is... ,in It is the dimension of the attack feature vector. Represents the first element in the attack feature vector. Each attack feature vector is first identified, and then k-means clustering is performed on all obtained attack feature vectors to obtain multiple security event clusters. Next, the cluster center vector corresponding to each security event cluster is obtained. The difference between the maximum and minimum values of each feature dimension across all cluster center vectors in the historical running data and the original attack data is calculated. All feature dimensions whose individual differences account for more than 25% of the total sum of differences are constructed into a feature index set. , feature index set Its corresponding normalization parameter pair Organized into a set of intelligent agent design parameters Subsequently, a vector set is constructed based on historical operational data and original attack data. The feature components of all vectors in this vector set are classified according to their physical meaning. That is, the indices of the link state components representing the link state in this vector set constitute a link state index set. The indices of the business state components representing the business state in this vector set constitute the business state index set. The indices of the attack state components representing the attack states in this vector set constitute the attack state index set. These three elements together constitute the state definition parameter set. ; The following formula is used to normalize and vectorize historical operational data to obtain attack feature vectors: ; Where m∈[1, the total number of historical running data obtained], Let m be the attack feature vector corresponding to the m-th historical running data. For attack feature mapping function, For the m-th historical running data, The attack type to which the m-th historical data belongs. For the attack severity level feature of the m-th historical data, For the m-th historical data segment, based on the source and destination assets in the network topology diagram The quantity characteristics of affected assets are calculated based on their positional relationships. This field represents the number of alarm triggers in the m-th historical data segment. The normalized alarm trigger count characteristic obtained through mean-variance normalization. This indicates the attack duration field in the m-th historical data segment. The normalized attack duration feature is obtained after mean-variance normalization, and the calculation methods for both are as follows: ; ; in For the m-th historical running data The number of alarms triggered within the preset statistics window. For the m-th historical running data The duration of the attack, and The number of alarms triggered are respectively The mean and standard deviation over historical operating data, and The duration of the attack The mean and standard deviation of historical operating data.
4. The network security counterfactual defense method based on multi-agent collaboration according to claim 3, characterized in that, Steps (1-5) specifically involve, firstly, setting up the intelligent agent design parameters. The corresponding index belongs to the feature index set. The features of each dimension are processed according to the mean-variance normalization rule and concatenated in index order to obtain the first feature. The input feature vector of a local defense agent : ; in, Indicates the first The attack feature vector at the th ... Values in a dimension and , respectively, are the mean and standard deviation of the d-th dimension of the attack feature vector corresponding to the historical running data and the original attack data; The input feature vector of a local defense agent Input feature dimension And d∈[1,D]; Then, define the state parameter set. Link-state index set Business status index set and attack state index set Concatenate them sequentially to obtain the local defense agent for the i-th site. Indexes of the required link status, service status, and attack status components; Subsequently, a shallow neural network model is used to develop the local defense agent for the i-th site. Perform offline training to obtain the risk assessment parameter vector for the i-th site. and bias parameters ; Finally, the feature index set Normalized parameters and the risk assessment parameter vector of the i-th site and bias parameters Combining local defense agents for the i-th site Among all the sites, the local defense agent corresponding to the central site is used as the core collaborative defense agent CDA; Steps (1-6) specifically involve first defining the parameter set based on the state. Link-state index set From the network topology diagram Link set Select all link state index sets The links associated with one or more indices are used to form the set of links to be monitored. Obtain the j-th link from the network device. Link utilization Average latency Packet loss rate monitoring value The link utilization, average latency, and packet loss rate monitoring values of all links are concatenated into a link state vector. Then, define the parameter set according to the state. Business status index set From business flow set Select the set with the highest business importance weight and business status index. The ten business flows associated with the index are used as a set of key business flows. Obtain the set of key business flows Each key business flow end-to-end delay End-to-end packet loss rate monitoring value and set up key business flows All corresponding end-to-end latency and end-to-end packet loss rate monitoring values are concatenated sequentially according to the service flow number to form a service state vector. Subsequently, the parameter set is defined according to the state. attack state index set The index counts and statistically analyzes relevant features in historical execution data and raw attack data, and concatenates them in a fixed order to form an attack state vector. ; Finally, define the parameter set based on the state. Link state vector Business state vector and attack state vector Concatenate into a joint state vector Then, the network topology diagram is drawn. Business Flow Set Business constraint parameter set Local defense intelligent agent set Core collaborative defense intelligent agent and the resulting joint state vector Integrate the results to obtain a set of basic modeling results. .
5. The network security counterfactual defense method based on multi-agent collaboration according to claim 4, characterized in that, Step (2) includes the following sub-steps: (2-1) Obtain the local raw data set of the i-th site, which consists of traffic logs, host logs and security alarm information within the t-period, from the i-th site to which the i-th node belongs in the original network topology data. For the acquired local raw data set Perform format processing to obtain the normalized local raw data set family of the i-th site within time period t. ); (2-2) The set of basic modeling results obtained in step (1-6) By integrating the network topology graph G and the service flow set F, the set of locally controllable links for the i-th site is obtained. and business link set and the resulting set of locally controllable links and business link set Discretize and merge to obtain the local action space set of the i-th station. ; (2-3) The normalized local raw data set of the i-th site obtained in step (2-1) Feature extraction is performed to obtain the local security feature vector of the i-th site within time period t. ; (2-4) The local security feature vector of the i-th site obtained in step (2-3) Perform local risk assessment to obtain the local risk score of the i-th site within time period t. ; (2-5) The normalized local raw data set of the i-th station obtained in step (2-1) within time period t. Perform service quality statistical processing to obtain the service quality vector of the i-th site within time period t. ; (2-6) The local security feature vector of the i-th site obtained in step (2-3) within time period t. The service quality vector of the i-th site obtained in step (2-5) within time period t and the local action space of the i-th site obtained in step (2-2). All candidate actions within the time period t are processed for action effect estimation to obtain the effect of the i-th station on the k-th candidate action within the time period t. Risk reduction estimate Impact of business quality on estimated value ; (2-7) The local risk score of the i-th site obtained in step (2-4) within time period t. Compared with the i-th station obtained in step (2-6) for the k-th candidate action Risk reduction estimate Impact of business quality on estimated value Comparison processing is performed to obtain a family of local candidate defense actions within time period t. .
6. The network security counterfactual defense method based on multi-agent collaboration according to claim 5, characterized in that, Steps (2-3) specifically involve processing the normalized local raw data set of the i-th station within time period t. The system counts and ratios of traffic logs, host logs, and security alerts. It then concatenates security features such as local alert counts, abnormal connection ratios, abnormal login counts, suspicious port access ratios, and average connection latency in a fixed order to obtain the local security feature vector of the i-th site within time period t. : ; in, For the site The number of local alarms within time period t. For the site The proportion of abnormal connections during time period t Let be the number of abnormal logins to the i-th site within time period t. Let be the percentage of suspicious port accesses to the i-th site during time period t. Let be the average connection latency of the i-th site during time period t. Let be the local security feature vector of the i-th site during time period t; Step (2-4) specifically involves using the local security feature vector of the i-th site. Input local defense agent Using a pre-defined linear scoring and non-linear activation structure, the local risk score of the i-th site during time period t is calculated. : ; in, Let i be the risk assessment weight vector for the i-th site. Let be the bias parameter for the i-th station. For the Sigmoid function, Give the local risk score for the i-th site during time period t; Steps (2-5) specifically involve first calculating the local raw data set of the i-th site within time period t. The traffic log records the average of all critical link utilization, average latency of all critical service flows, and average packet loss rate of all critical service flows. Then, the critical link utilization, average latency of critical service flows, and average packet loss rate are sorted according to their average values to obtain the service quality vector of the i-th site within time period t. ; Steps (2-6) specifically involve, firstly, defining the local action space of the i-th site. The k-th candidate action Encode the feature vector of the k-th candidate action at the i-th station. Then, the local security feature vector of the i-th site during time period t is... The obtained k-th candidate action feature vector of the i-th station After concatenation, input the data into a linear estimation model to obtain the relationship between the i-th station and the k-th candidate action. Risk reduction estimate Simultaneously, the service quality vector of the i-th site during time period t is... With the obtained candidate action feature vector After concatenation, input another linear estimation model to obtain the i-th station for the k-th candidate action. Business quality impact estimate : ; ; in, To estimate the parameter vector for risk reduction at the i-th site, Let be the parameter vector for estimating the service quality impact of the i-th site. This is the concatenated vector of the local security feature vector of the i-th site and the feature vector of the k-th candidate action. It is the concatenated vector of the service quality vector of the i-th station and the feature vector of the k-th candidate action; Step (2-7) specifically involves first comparing the local risk score of the i-th site with the local risk threshold. If the risk level is greater than a threshold, the site is considered a high-risk site; otherwise, it is considered a low-risk site. Then, the i-th site is used as a candidate action. Risk reduction estimate Impact of business quality on estimated value By comparison, those with a risk reduction estimate greater than the risk reduction threshold are selected. Furthermore, the estimated impact of business quality on the loss is not lower than the preset lower limit. All candidate actions are selected, and finally, all the selected candidate actions are combined to obtain a set of local candidate defense actions for all sites within time period t. .
7. The network security counterfactual defense method based on multi-agent collaboration according to claim 6, characterized in that, Step (3) includes the following sub-steps: (3-1) Based on the site index set The set of local candidate defense actions within time period t obtained in step (2) Get the local candidate defense action set for the i-th site. The local candidate defense action sets of all obtained sites are merged to obtain the global candidate action set for time period t. ; (3-2) For the global candidate action set within time period t obtained in step (3-1) For each candidate action, the local risk score of the i-th site during time period t is obtained by combining steps (2-3). and the i-th station with respect to the k-th candidate action obtained in steps (2-5). Risk reduction estimate Impact of business quality on estimated value A comprehensive scoring process is performed to obtain the comprehensive score set of the candidate action within time period t. ; Specifically, this step involves first obtaining the global candidate action set within time period t. Each candidate action in The associated first The first site, then obtain the first Local risk score for each site and regarding candidate actions Risk reduction estimate Impact of business quality on estimated value Then, the three factors are weighted and summed according to preset weights to obtain the result for the i-th station in relation to the k-th candidate action. Overall rating Finally, the combined scores of all stations for all candidate actions are merged into a combined score set for candidate actions. : ; in, The three factors are weighted coefficients for the overall score, and their sum is 1. The value is 0.
4. The value is 0.
3. The value is 0.3; To obtain the maximum of the two; (3-3) The comprehensive score set of candidate actions obtained in step (3-2) Sorting and truncation are performed to obtain a set of selected candidate actions within time period t. ; Specifically, this step involves first, obtaining a comprehensive score set for the candidate actions. All candidate actions are sorted by their comprehensive scores from largest to smallest, and then the top-ranked actions in the sorting results are obtained. Each candidate action corresponds to a specific candidate action. Finally, all candidate actions are combined into a set of selected candidate actions for the time period t. ,in This is the preset maximum number of selected actions; (3-4) The selected candidate action set within time period t obtained in step (3-3) A combined construction process is performed to obtain a set of joint defense schemes for time period t. ; Specifically, this step involves first using a greedy algorithm to traverse the set of selected candidate actions within time period t. To obtain a comprehensive score From the top 30% of candidate actions, all the resulting candidate actions are then combined according to a preset strategy to form a set of joint defense schemes for time period t. .
8. The network security counterfactual defense method based on multi-agent collaboration according to claim 7, characterized in that, Step (4) includes the following sub-steps: (4-1) The set of joint defense schemes for time period t obtained from step (3-4) Obtain the r-th joint defense scheme For the r-th joint defense scheme obtained Perform control sequence mapping to obtain the set of defense control sequences for time period t. and baseline control sequence set Where r∈[1, the total number of joint defense schemes]; (4-2) The set of defense control sequences for executing the r-th joint defense scheme within time period t obtained in step (4-1). and baseline control sequence set and the set of basic modeling results obtained in steps (1-6). Joint state vector Perform joint simulation to obtain the set of state trajectories for executing the r-th joint defense scheme within time period t. The set of state trajectories for not implementing the joint defense plan ; (4-3) For each of the state trajectory sets obtained in step (4-2) during time period t, execute the r-th joint defense scheme. The set of attack state variables and state trajectories that do not execute the joint defense scheme. The attack status variables are aggregated using time-weighted and node-weighted methods to obtain the attack risk indicator set for the execution of the r-th joint defense scheme within time period t. The set of attack risk indicators for when the joint defense plan is not implemented during time period t. ; (4-4) The set of state trajectories for executing the r-th joint defense scheme during time period t obtained in step (4-2). The latency and packet loss rate of critical business flows are aggregated using time-weighted and business-weighted methods to obtain a set of business / operational security indicators for executing the r-th joint defense scheme within time period t. And the set of state trajectories for which the joint defense scheme is not implemented during time period t. The latency and packet loss rate of critical business flows are aggregated using time-weighted and business-weighted methods to obtain a set of business / operational security indicators for scenarios where the joint defense scheme is not executed within time period t. ; (4-5) Execute the attack risk index of the r-th joint defense scheme within the time period t obtained in step (4-3). A set of attack risk indicators for failure to implement joint defense strategies And the business / operational security indicators for implementing the r-th joint defense scheme within time period t obtained in step (4-4). Business / Operational Security Indicators for Failure to Implement Joint Defense Plan Differential processing is performed to obtain the net defense effect of executing the r-th joint defense scheme within time period t. Net business impact And the net defense effectiveness of all joint defense schemes within the obtained time period t. Net business impact and joint defense plan The combination is the set of joint defense strategy effectiveness evaluation results within time period t. .
9. The network security counterfactual defense method based on multi-agent collaboration according to claim 8, characterized in that, Step (4-1) specifically involves first, determining the r-th joint defense scheme. The site-level candidate actions of the i-th site are discretized in time to obtain a candidate action list. Then, the obtained candidate action list is mapped using one-hot encoding to obtain the defense control vector sequence of the i-th site executing the r-th joint defense scheme within time period t. and baseline control vector sequence Then, the defense control vector sequences of all stations executing the r-th joint defense scheme within time period t are combined into a set of defense control sequences executing the r-th joint defense scheme within time period t. The baseline control vector sequences of all stations are combined into a set of baseline control sequences for time period t. ; Step (4-2) specifically involves, under the same joint state vector In a digital twin environment, the defense control vector sequence for executing the r-th joint defense scheme within time period t is respectively... and baseline control vector sequence Discrete-time simulations are performed, and the system state trajectory of each joint defense scheme is recorded throughout the simulation process to obtain the set of state trajectories for executing the r-th joint defense scheme within time period t. The set of state trajectories for not implementing the joint defense plan ; Step (4-3) specifically involves, for the r-th joint defense scheme By performing the same normalization and vectorization processing as in step (1) on the newly generated running data and attack data of each node in its state trajectory, a new attack state vector is obtained. Then, the modulo of the obtained attack state vector is taken to obtain the attack state indicator. Finally, the obtained attack state indicator is accumulated according to the time and the proportion of key business flow of the node to obtain the attack risk indicator of the execution of the r-th joint defense scheme in time period t. Attack risk indicators for not implementing the joint defense plan during time period t : ; ; in, and They are nodes Executing the r-th joint defense plan and failure to implement joint defense plan time The attack status indicator, and has ∈[1, ], ∈[1, t]; For nodes The proportion of key business flows; For a moment The time decay value, and equal to: ; in, Let be the set of discrete times within time interval t. Let t be the end time point of the time period. The time decay coefficient, This indicates that no time discounting is applied, and all time points are weighted equally. The larger the value, the stronger the discount on time. 0.3, ∈ ; In step (4-4), the operational security indicators for executing the r-th joint defense scheme within time period t are obtained. and business / operational security indicators that are not implemented in each plan The following formula is used: ; ; in, For critical business flows Importance weight; and At time respectively Execute the r-th joint defense plan The time delay and the failure to execute the r-th joint defense scheme Time delay, and At time respectively Execute the r-th joint defense plan Packet loss rate and failure to execute the r-th joint defense scheme Packet loss rate at that time; For a moment The time decay value is calculated in the same way as in step (4-3); For business flow The security utility function is specifically... ; in and These are the latency threshold and packet loss rate threshold in the set of business constraint parameters; In step (4-5), the net defense effect of executing the r-th joint defense scheme within time period t is obtained. Net business impact The formula is as follows: ; ; 10. The network security counterfactual defense method based on multi-agent collaboration according to claim 9, characterized in that, Step (5) includes the following sub-steps: (5-1) Set of joint defense scheme utility evaluation results obtained from step (4-5) within time period t We obtain all joint defense schemes with positive net defense effectiveness and net business impact, acquire the simulation trajectory corresponding to all links and all critical business flows under each joint defense scheme, and perform extreme value processing on the simulation trajectory to obtain the set of security margin indicators for all links and critical business flows within time period t. ; Specifically, this step involves first analyzing the set of joint defense strategy effectiveness evaluation results within time period t. First, identify all joint defense schemes where both net defense effectiveness and net business impact are positive. Then, obtain the latency and packet loss rate of each link in each joint defense scheme within time period t. Extract the maximum latency from all latency values and the maximum packet loss rate from all packet loss rates. Next, obtain all critical business flows included in the critical business flow set within each joint defense scheme, calculate the number of available paths for each critical business flow within time period t, and extract the minimum number of available paths from all available paths. Finally, combine the maximum latency, maximum packet loss rate, and minimum number of available paths as extreme value indicators to obtain the set of security margin indicators within time period t. ; (5-2) From the set of basic modeling results The maximum allowable latency and maximum allowable packet loss rate of each link are obtained, and the set of security margin indicators for the time period t obtained in step (5-1) is adjusted according to these maximum allowable latency and maximum allowable packet loss rate. Filtering is performed to obtain multiple joint defense schemes corresponding to multiple security margin indicators within time period t, which are then used as a set of feasible joint defense schemes. ; Specifically, this step involves, firstly, analyzing the set of basic modeling results. The maximum allowable latency and maximum allowable packet loss rate for each link are obtained, and the set of security margin indicators for the obtained time period t is adjusted based on these maximum allowable latency and maximum allowable packet loss rate. Filtering is performed, which involves checking whether the maximum link latency and maximum packet loss rate in each combination of the security margin indicator set do not exceed the maximum allowable latency and packet loss rate. Simultaneously, the results are filtered from the basic modeling result set. The process involves obtaining a set of business flows and checking whether the minimum number of available paths for critical business flows in each combination within the security margin indicator set meets the business requirements based on the number of services in the business flow set. Finally, the joint defense schemes corresponding to all combinations that satisfy all the above constraints are integrated to obtain a set of secure and feasible joint defense schemes for time period t. ; (5-3) Obtain the set of safe and feasible joint defense schemes within time period t obtained in step (5-2). The net defense effect and net business impact of each secure and feasible joint defense scheme are calculated, and the two are added together. The secure and feasible joint defense scheme with the largest sum is selected as the final joint defense scheme.