Power system cross-domain network security modeling method and system based on federal cognitive collaboration

By adopting a federated cognitive collaborative approach to cross-domain cybersecurity modeling of power systems, the problems of data privacy leakage and insufficient model generalization are solved. This approach improves the accuracy and robustness of threat identification in highly non-independent and co-distributed scenarios, and adapts to the changing network threat environment of multi-regional power grids.

CN121547368AActive Publication Date: 2026-02-17HUNAN KUANGAN NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610065416.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-19
Publication Date
2026-02-17
Estimated Expiration
2046-01-19

AI Technical Summary

Technical Problem

Existing cross-domain cybersecurity modeling methods for power systems suffer from data privacy leakage risks, insufficient model generalization ability, and distortion of threat detection results. In particular, they are difficult to effectively identify and filter abnormal nodes in highly non-independent and identically distributed scenarios.

Method used

By adopting a federated cognitive collaboration approach, through local data preprocessing and model training, combined with distribution similarity calculation, hierarchical clustering and parameter directional distance filtering, we can achieve adaptive aggregation and abnormal node filtering of cross-domain threat identification models, avoid cross-domain data transmission and improve the model's generalization ability and robustness.

Benefits of technology

It achieves privacy protection without violating the compliance of cross-domain power data flow, significantly improves the accuracy and robustness of threat identification, and adapts to the changing network threat environment of multi-regional power grids.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121547368A_ABST
    Figure CN121547368A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power system cross-domain network security modeling method based on federal cognitive collaboration, and the method comprises the steps: introducing a federal cognitive collaboration mechanism, and carrying out the cross-domain network security modeling of an electric power system on the premise of protecting data privacy and complying with an electric power system privacy protocol; the problem that the generalization performance of the model is reduced due to data non-independent identically distributed and abnormal nodes in cross-domain network security modeling of the power system is solved; according to the method disclosed by the invention, a hierarchical clustering and model fusion strategy is adopted, so that a client can use knowledge of global and other clusters for reference during local training, and the limitation of data distribution of the client is broken through; a server side screens reliable clients and generates a cluster threat identification model and a global threat identification model which are adaptive to different data distributions through outlier filtering and clustering aggregation, so that the accuracy and robustness of threat identification are improved, and dynamic adaptation to a network threat environment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of power system network security and artificial intelligence, and more particularly relates to a power system cross-domain network security modeling method and system based on federal cognitive collaboration. BACKGROUND

[0002] With the continuous improvement of informationization and intelligentization level of the power system, the attack threat faced by the power system cyberspace is becoming increasingly serious. Attackers can use cross-domain penetration, horizontal movement and other methods to implement network attacks on business systems and management systems of power generation, power transmission, power transformation, power distribution and power consumption. Therefore, power system cross-domain network security modeling and threat identification for multi-level dispatch centers and multi-regional power grids have become an urgent need.

[0003] The commonly used power system cross-domain network security modeling method at present is to collect network traffic, system logs and other security data of each provincial dispatch center, regional dispatch center and related business domain to a central server, and use a unified machine learning or deep learning model on the central server for modeling analysis to identify typical network attack behaviors such as denial of service attacks and data injection attacks. In order to improve the identification accuracy, the existing method often needs to collect large-scale raw data of each domain for a long time, and store and train offline at the central node. After training is completed, the model is distributed to each level of dispatch center for deployment and use.

[0004] However, the above existing power system cross-domain network security modeling method still has some defects that cannot be ignored: First, centralized modeling needs to transmit and store a large amount of raw security data across domains, and power system operation data often involves critical infrastructure and user privacy information. Such cross-domain aggregation not only increases the risk of data leakage, but also is difficult to meet the compliance requirements of existing power data security management, privacy protection and cross-regional data flow; Second, the centralized method usually assumes that the data of different dispatch regions satisfies independent and identically distributed, but in reality, there are significant differences in power grid structure, load characteristics, device types and attack behavior patterns among regions, resulting in insufficient generalization ability of the model in highly non-independent and identically distributed scenarios; Third, the centralized modeling lacks an effective mechanism to identify and filter malicious or abnormal clients, and is easily affected by malicious data updates uploaded by abnormal nodes, causing deviation or distortion of the identification results of the threat model. SUMMARY

[0005] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a method and system for cross-domain cybersecurity modeling of power systems based on federated cognitive collaboration. Its purpose is to solve the technical problems of existing cross-domain cybersecurity modeling methods for power systems, which, due to their centralized aggregation approach, result in the cross-domain transmission of large amounts of raw security data, leading to data privacy risks and difficulty in meeting compliance requirements for cross-domain power data flow; the technical problem that, due to significant differences in network structure, load characteristics, and attack patterns among different dispatch centers, existing methods cannot effectively handle highly non-independent and identically distributed data, resulting in insufficient model generalization performance; and the technical problem that, due to the lack of effective identification and protection mechanisms for malicious or abnormal nodes, the model is easily contaminated by abnormal updates during the aggregation process, leading to distorted threat detection results.

[0006] To achieve the above objectives, according to one aspect of the present invention, a cross-domain network security modeling method for power systems based on federated cognitive collaboration is provided, which is applied to applications including... In a power system for both clients and servers, and For any natural number, the cross-domain network security modeling method for this power system includes the following steps: (1) No. Each client obtains network traffic, system log data, and multi-dimensional features reflecting communication behavior and equipment status from its local network devices and Supervisory Control and Data Acquisition / Energy Management System (SCADA / EMS). One-hot encoding is used to process the categorical features within these multi-dimensional features, while Z-score normalization is used to process the numerical features. The results are then used to construct a local dataset. ,in ; (2) Server initializes global threat identification model Clustering Threat Identification Model Set And initialize the j-th cluster threat identification model in the cluster threat identification model set for the k-th client. and the initialized global threat identification model and the Threat identification model for clusters to which each client belongs Issued to the One client; among which The number of clusters is preset, and j represents the number of clusters. The index of the cluster threat identification model to which each client belongs in the cluster threat identification model set; (3) No. Each client sets the training round counter t=1; (4) No. Each client determines whether t is greater than the preset training round threshold T. If it is, proceed to step (8); otherwise, set t=t+1 and proceed to step (5). (5) No. The client obtains the first Global threat identification model trained in rounds and its associated cluster threat identification model And based on the acquired global threat identification model Clustering Threat Identification Model Initialize the first Local threat identification model trained in rounds : (6) No. Each client uses its local dataset The first one obtained in step (5) Local initial model during round training Training is performed to obtain a trained local threat identification model. Calculate the trained local threat identification model With the Global threat identification model trained in rounds The distribution similarity between them is used to construct the first distribution similarity. Similarity distance matrix of round training And the parameters of the trained local threat identification model and similarity distance matrix Uploaded to the server; (7) The server uses the parameters of the local threat identification model uploaded by all clients { All clients are identified and filtered to exclude potentially abnormal clients, and multiple filtered clients are obtained. These filtered clients are then compared with the global threat identification model. The similarity distance matrix between them is used, and a hierarchical clustering algorithm is employed to divide all filtered clients into S clusters. For each cluster Perform intra-cluster weighted average aggregation to generate the cluster. The corresponding threat recognition model trained in the (t+1)th round. And for all clusters, the threat identification model trained in round t+1 is { Perform average aggregation to generate the global threat identification model trained in the (t+1)th round. And this global threat identification model and the threat recognition model of the (t+1)th round of training corresponding to all the clustering clusters is sent to the ith client, and the process returns to step (4), where i∈[1, S]; (8) All clients use the global threat recognition model of the Tth round of training to perform threat recognition on local real-time network traffic and system log data to output attack categories and threat levels.

[0007] Preferably, the value of the training round threshold T is in the range of 100 to 500.

[0008] Preferably, step (5) is to initialize the local threat recognition model using the following formula : .

[0009] Preferably, in step (6), the kth client constructs a similarity distance matrix between the trained local threat recognition model and the global threat recognition model of the Tth round of training based on the distribution similarity, which specifically includes the following steps:First, the kth client randomly selects samples from each category in the local dataset; then, for each category, the kth client inputs each sample of the category into the updated local threat recognition model parameters to obtain a first output vector, and inputs each sample of the category into the global threat recognition model to obtain a second output vector, and calculates the Euclidean distance between the first output vector and the second output vector corresponding to the sample, and then takes the average of the Euclidean distances corresponding to all samples in the category to obtain the average Euclidean distance corresponding to the category; finally, all the average Euclidean distances corresponding to all categories are collected, which constitutes the similarity distance matrix between the kth client and the global threat recognition model . .

[0010] Preferably, in step (7), the server performs identification and screening on all clients based on the parameters of the local threat recognition models uploaded by all clients to exclude potential abnormal clients and obtain filtered multiple clients, which specifically includes the following sub-steps: (7-1) The server calculates the parameter directional distance MDD of the kth client . (7-2) The server obtains the mean and standard deviation of the parameter directional distance MDD of all clients;​​​ (7-3) The server obtains the first step based on step (7-1). The parameter direction distance of each client The mean of the directional distances of all clients obtained in step (7-2) to the parameter direction. and standard deviation Get the The standard deviation radius of each client; (7-4) Server filters out standard deviation radius Clients with a value greater than or equal to a preset threshold are selected to obtain a filtered list of clients.

[0011] Preferably, step (7-1) uses the following calculation formula: ; in, , Global threat identification model The total number of parameters, A function representing statistical quantities.

[0012] Preferably, step (7-3) specifically uses the following formula to obtain the first... Radius of standard deviation for each client: .

[0013] Preferably, the cluster is obtained in step (7-4). The corresponding cluster threat identification model trained in the (t+1)th round The following formula is used: ; in Indicates the first Number of local data samples per client.

[0014] Preferably, in step (7), the server trains the cluster threat identification model for all clusters in the (t+1)th round. Perform average aggregation to obtain the global threat identification model trained in round t+1. The following formula is used: .

[0015] According to another aspect of the present invention, a cross-domain network security modeling system for power systems based on federated cognitive collaboration is provided, which is applied in applications including... In a power system for both clients and servers, and For any natural number, the power system cross-domain network security modeling system includes the following steps: The first module, which is set in the first... A client is used to acquire network traffic, system log data, and multi-dimensional features reflecting communication behavior and equipment status from network devices and Supervisory Control and Data Acquisition / Energy Management System (SCADA / EMS) in its local domain. One-hot encoding is used to process the categorical features of these multi-dimensional features, and Z-score normalization is used to process the numerical features. The processing results are then used to construct a local dataset. ,in ; The second module, located on the server, is used to initialize the global threat identification model. Clustering Threat Identification Model Set And initialize the j-th cluster threat identification model in the cluster threat identification model set for the k-th client. and the initialized global threat identification model and the Threat identification model for clusters to which each client belongs Issued to the One client; among which The number of clusters is preset, and j represents the number of clusters. The index of the cluster threat identification model to which each client belongs in the cluster threat identification model set; The third module, which is set in the first... One client is used to set the training round counter t=1; The fourth module, which is set in the first... One client is used to determine whether t is greater than the preset training round threshold T. If it is, the process proceeds to the eighth module; otherwise, t is set to t+1 and the process proceeds to the fifth module. The fifth module, which is set in the first... The client is used to obtain the first one. Global threat identification model trained in rounds and its associated cluster threat identification model And based on the acquired global threat identification model Clustering Threat Identification Model Initialize the first Local threat identification model trained in rounds : The sixth module, which is set in the first... One client, used to use its local dataset The fifth module yielded the first Local initial model during round training Training is performed to obtain a trained local threat identification model. Calculate the trained local threat identification model With the Global threat identification model trained in rounds The distribution similarity between them is used to construct the first distribution similarity. Similarity distance matrix of round training And the parameters of the trained local threat identification model and similarity distance matrix Uploaded to the server; The seventh module, located on the server, is used to determine the parameters of the local threat identification model uploaded by all clients. All clients are identified and filtered to exclude potentially abnormal clients, and multiple filtered clients are obtained. These filtered clients are then compared with the global threat identification model. The similarity distance matrix between them is used, and a hierarchical clustering algorithm is employed to divide all filtered clients into S clusters. For each cluster Perform intra-cluster weighted average aggregation to generate the cluster. The corresponding threat recognition model trained in the (t+1)th round. And for all clusters, the threat identification model trained in round t+1 is { Perform average aggregation to generate the global threat identification model trained in the (t+1)th round. And this global threat identification model and the threat identification model trained in the (t+1)th round for all clusters { } Send it to the i-th client and return to step (4), where i∈[1,S]; The eighth module, which is set in the first... One client is used to utilize the global threat identification model trained in round T. The network traffic and system log data obtained from the first module are used for threat identification to output attack categories and threat levels.

[0016] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: 1. This invention adopts the local data construction and preprocessing mechanism in step (1) and the local training and uploading of only model parameters and similarity distance matrix in step (6), which avoids the transmission of sensitive data such as raw network traffic and system logs across dispatch centers and realizes the privacy protection mode of "usable but invisible" data. Therefore, it can complete cross-domain threat modeling without touching the compliance red line of cross-domain transmission of power data, and solves the technical problems of high risk of data privacy leakage and insufficient compliance in the existing centralized method. 2. Because the present invention adopts the distribution similarity calculation mechanism in step (6) and the hierarchical clustering and cross-cluster customized aggregation method based on the similarity distance matrix in step (7), it can automatically identify the data distribution differences between various scheduling centers and divide clients with similar distributions into the same cluster for differentiated model training and aggregation. Therefore, it can significantly improve the generalization and adaptive capabilities of the threat identification model in highly non-independent identically distributed (Non-IID) scenarios, thereby effectively solving the technical problem of low identification accuracy of existing centralized methods under data heterogeneity. 3. This invention employs the parameter direction distance calculation and standard deviation radius outlier filtering mechanism in steps (7-1) to (7-4), combined with a filtering and then aggregation strategy. It can effectively identify and exclude malicious or low-quality nodes that upload abnormal model parameters, thereby ensuring that the updates of the cluster threat identification model and the global threat identification model come from reliable clients. Therefore, it can significantly improve the robustness of model aggregation and solve the technical problem in existing methods that the model is easily contaminated by abnormal nodes, leading to distortion of threat detection results. 4. This invention adopts the dual-source parameter fusion initialization mechanism of the global threat identification model combined with the cluster threat identification model in step (5), and the multi-round federation iteration and update strategy in step (7). It can integrate knowledge from different regions before local training on the client and continuously adapt and evolve during the training process, thereby breaking through the bottleneck of a single client being limited by its own data and realizing cross-domain cognitive collaboration and knowledge sharing between multi-region and multi-level scheduling centers. Attached Figure Description

[0017] Figure 1 This is a flowchart of the cross-domain network security modeling method for power systems based on federated cognitive collaboration, as described in this invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0019] It should be noted that in the description of the embodiments of the present application, the terms "comprising", "containing" or any other variants thereof are intended to cover the non-exclusive inclusion, so that the process, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes the elements inherent to such process, article or device. Without more limitations, the element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, article or device including the element. The terms "upper", "lower" and the like indicate the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0020] In addition, the technical solutions of each embodiment of the present application can be combined with each other, but it must be based on the fact that a person of ordinary skill in the art can realize it, and when the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist, nor is it within the scope of protection required by the present application.

[0021] The basic idea of the present application is to provide a power system cross-domain network security modeling method based on federal cognitive collaboration, which introduces a federal cognitive collaboration mechanism to solve the problem of model generalization performance decline caused by data Non-Independent and Identically Distributed (Non-IID) and abnormal nodes in power system cross-domain network security modeling under the premise of protecting data privacy and complying with power system privacy protocols. Specifically, this method uses hierarchical clustering and model fusion strategy, so that the client can learn from the global and other cluster knowledge when training locally, breaking the limitations of its own data distribution; the server side filters reliable clients and generates clustering cluster threat identification models and global threat identification models adaptive to different data distributions through outlier filtering and clustering aggregation, thereby improving the accuracy and robustness of threat identification, and realizing dynamic adaptation to network threat environment.

[0022] As shown in Figure 1 The present application provides a power system cross-domain network security modeling method based on federal cognitive collaboration, which is applied in a power system including A client (which is set in each level of power dispatching center) and a server, and is any natural number, the power system cross-domain network security modeling method includes the following steps: (1) No. Each client obtains network traffic, system log data, and multi-dimensional features reflecting communication behavior and equipment status from its local network devices and Supervisory Control and Data Acquisition / Energy Management System (SCADA / EMS). One-hot encoding is used to process the categorical features within these multi-dimensional features, while Z-score normalization is used to process the numerical features. The results are then used to construct a local dataset. ,in ; (2) Server initializes global threat identification model Clustering Threat Identification Model Set And initialize the j-th cluster threat identification model in the cluster threat identification model set for the k-th client. and the initialized global threat identification model and the Threat identification model for clusters to which each client belongs Issued to the One client; among which The preset number of clusters is defined, ranging from 3 to 10, preferably 5, where j represents the number of clusters. The index of the cluster threat identification model to which each client belongs in the cluster threat identification model set; During the initialization process of this step, the threat identification model for all client clusters is the same; The advantage of this step (2) is that by uniformly initializing the global threat identification model and the cluster threat identification model on the server side, the consistency and comparability of the initial threat identification models of each client are ensured, providing a unified starting point for subsequent clustering and differential evolution based on distribution similarity, and improving the stability of the aggregation results.

[0023] (3) No. Each client sets the training round counter t=1; (4) No. Each client determines whether t is greater than the preset training round threshold T. If it is, proceed to step (8); otherwise, set t=t+1 and proceed to step (5). Specifically, the training round threshold T in this step ranges from 100 to 500, preferably 200.

[0024] (5) No. The client obtains the first Global threat identification model trained in rounds and its associated cluster threat identification model And based on the acquired global threat identification model Clustering Threat Identification Model Initialize the first Local threat identification model trained in rounds : Specifically, this step initializes the local threat identification model using the following formula. : ; The advantage of this step (5) is that by using the global threat identification model and the cluster threat identification model of round t to initialize the local model, the local training can take into account both global knowledge and local knowledge under the same data distribution, thereby improving the convergence speed and identification performance of the model on local data.

[0025] (6) No. Each client uses its local dataset The first one obtained in step (5) Local initial model during round training Training is performed to obtain a trained local threat identification model. Calculate the trained local threat identification model With the Global threat identification model trained in rounds The distribution similarity between them is used to construct the first distribution similarity. Similarity distance matrix of round training And the parameters of the trained local threat identification model and similarity distance matrix Uploaded to the server; The advantage of this step (6) is that by completing the threat identification model training and distribution similarity calculation locally, and only uploading the model parameters and similarity distance matrix to the server, the original data is not leaked, and refined and effective statistical information is provided for clustering and robust aggregation on the server side.

[0026] In this step, the first Each client uses a trained local threat identification model. With the Global threat identification model trained in rounds The process of constructing a similarity distance matrix based on the distributional similarity between threats involves the following steps: First, the k-th client randomly selects samples from each category in the local dataset; then, for each category, each sample from that category is input into the updated local threat identification model parameters. To obtain the first output vector, each sample of that category is fed into the global threat identification model. To obtain the second output vector, the Euclidean distance between the first and second output vectors corresponding to the sample is calculated. Then, the average Euclidean distances of all samples in that category are averaged to obtain the average Euclidean distance for that category. Finally, the average Euclidean distances of all categories are aggregated to form the first output vector. Individual Client and Global Threat Identification Model Similarity distance matrix between .

[0027] (7) The server uses the parameters of the local threat identification model uploaded by all clients { All clients are identified and filtered to exclude potentially abnormal clients, and multiple filtered clients are obtained. These filtered clients are then compared with the global threat identification model. The similarity distance matrix between them is used, and a hierarchical clustering algorithm is employed to divide all filtered clients into S clusters. For each cluster Perform intra-cluster weighted average aggregation to generate the cluster. The corresponding threat recognition model trained in the (t+1)th round. And for all clusters, the threat identification model trained in round t+1 is { Perform average aggregation to generate the global threat identification model trained in the (t+1)th round. And this global threat identification model and the threat identification model trained in the (t+1)th round for all clusters { } Send it to the i-th client and return to step (4), where i∈[1,S]; In this step (7), the server uses the parameters of the local threat identification model uploaded by all clients. The process of identifying and filtering all clients to exclude potentially abnormal clients and obtaining multiple filtered clients specifically includes the following sub-steps: (7-1) The server calculates the first... The parameter direction distance of each client ; This step uses the following calculation formula:

[0028] in, , Global threat identification model The total number of parameters, The parameter direction distance reflects the consistency of the parameter update direction of the client threat identification model with the global threat identification model. (7-2) The server obtains the mean value of the parameter direction distance MDD of all clients and the standard deviation ; (7-3) The server obtains the standard deviation radius of the first client according to the parameter direction distance MDD of the first client obtained in step (7-1) and the mean value and the standard deviation of the parameter direction distance MDD of all clients obtained in step (7-2) . This step specifically obtains the standard deviation radius of the first client by using the following formula:

[0029] (7-4) The server filters out the clients whose standard deviation radius is greater than or equal to a preset threshold value to obtain a plurality of filtered clients.

[0030] Specifically, the preset threshold value is in the range of 2 to 4, and is preferably 3.

[0031] The above sub-steps (7-1) to (7-4) have the advantage that by introducing the statistical indicators of the parameter direction distance and the standard deviation radius thereof, the abnormal nodes whose parameter update direction deviates obviously from most clients are effectively identified and removed before aggregation, thereby avoiding the pollution of malicious or low-quality model updates to the global threat identification model and the cluster threat identification model, and improving the security and robustness of the model aggregation process.

[0032] In this step (7), the cluster corresponding to the cluster threat identification model of the t+1th round of training is obtained by using the following formula:

[0033] wherein represents the number of local data samples of the first client.

[0034] In this step (7), the server performs average aggregation processing on the cluster threat identification models of the t+1th round of training corresponding to all clusters to obtain the global threat identification model of the t+1th round of training is obtained by using the following formula: ​​​

[0035] (8) the Tth round of training of the global threat identification model Step (1) obtained by the network traffic and system log data for threat identification, to output attack category and threat level.

[0036] The advantage of this step (8) is that by deploying the global threat identification model converged after the Tth round of training in the client of each level of dispatch center, the threat identification is carried out in real time without increasing additional communication burden, and the discovery speed and accuracy of the power system to various cross-domain network attacks are improved.

[0037] Those skilled in the art will readily understand that the above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.​

Claims

1. A federated cognitive collaboration based power system cross-domain network security modeling method, applied in a power system comprising a plurality of clients and a server, and n is any natural number, characterized in that, The power system cross-domain network security modeling method comprises the following steps: (1) the first The network flow, system log data, and multi-dimensional features reflecting communication behavior and device status are obtained by the first client from the network equipment and the data acquisition and monitoring control system / energy management system (SCADA / EMS) in the local region of the client, the categorical features in the multi-dimensional features are processed by using a one-hot encoding method, the numerical features in the multi-dimensional features are processed by using a Z-score standardization method, and the processed results are constructed as a local data set , wherein ; (2) Server initializes global threat identification model Clustering Threat Identification Model Set And initialize the j-th cluster threat identification model in the cluster threat identification model set for the k-th client. and the initialized global threat identification model and the Threat identification model for clusters to which each client belongs Issued to the One client; among which The number of clusters is preset, and j represents the number of clusters. The index of the cluster threat identification model to which each client belongs in the cluster threat identification model set; (3) the first client sets a training round counter t = 1; (4) the first client judges whether t is greater than a preset training round threshold T, if yes, it goes to step (8), otherwise, it sets t = t + 1 and goes to step (5); (4) the first client judges whether t is greater than a preset training round threshold T, if yes, it goes to step (8), otherwise, it sets t = t + 1 and goes to step (5); (5) the first client obtains the global threat identification model of the first round of training and the cluster threat identification model to which it belongs, and initializes the local threat identification model of the first round of training according to the obtained global threat identification model and cluster threat identification model : (6) No. Each client uses its local dataset The first one obtained in step (5) Local initial model during round training Training is performed to obtain a trained local threat identification model. Calculate the trained local threat identification model With the Global threat identification model trained in rounds The distribution similarity between them is used to construct the first distribution similarity. Similarity distance matrix of round training And the parameters of the trained local threat identification model and similarity distance matrix Uploaded to the server; (7) The server uses the parameters of the local threat identification model uploaded by all clients { All clients are identified and filtered to exclude potentially abnormal clients, and multiple filtered clients are obtained. These filtered clients are then compared with the global threat identification model. The similarity distance matrix between them is used, and a hierarchical clustering algorithm is employed to divide all filtered clients into S clusters. For each cluster Perform intra-cluster weighted average aggregation to generate the cluster. The corresponding threat recognition model trained in the (t+1)th round. And for all clusters, the threat identification model trained in round t+1 is { Perform average aggregation to generate the global threat identification model trained in the (t+1)th round. And this global threat identification model and the threat identification model trained in the (t+1)th round for all clusters { } Send it to the i-th client and return to step (4), where i∈[1,S]; (8) All clients utilize the global threat identification model trained in round T Threat identification on local real-time network traffic and system log data to output attack categories and threat levels.

2. The federated cognitive collaboration based power system cross-domain network security modeling method of claim 1, wherein, The value range of the training round threshold T is 100 to 500.

3. The federated cognitive collaboration based power system cross-domain network security modeling method of claim 2, wherein, Step (5) is to initialize the local threat identification model using the following equation : 。 4. The federated cognitive collaboration based power system cross-domain network security modeling method of claim 3, wherein, The process of constructing the similarity distance matrix between the kth client and the global threat identification model trained in the (t-1)th round is as follows: first, the kth client randomly selects samples from each category in the local data set; then, for each category, each sample of the category is input into the updated local threat identification model parameters to obtain a first output vector, and each sample of the category is input into the global threat identification model to obtain a second output vector, and the Euclidean distance between the first output vector and the second output vector corresponding to the sample is calculated, then the average Euclidean distance corresponding to the category is obtained by averaging the Euclidean distances corresponding to all samples in the category; finally, the average Euclidean distances corresponding to all categories are collected, that is, the similarity distance matrix between the kth client and the global threat identification model is constructed. ​​​​​​​​​ 5. The federated cognitive collaborative based power system cross-domain network security modeling method according to claim 4, characterized in that, In step (7), the server uses the parameters of the local threat identification model uploaded by all clients { The process of identifying and filtering all clients to exclude potentially abnormal clients and obtaining multiple filtered clients specifically includes the following sub-steps: (7-1) The server calculates the parameter direction distance of each client ;​ (7-2) The server obtains the mean of the parameter directional distance MDD of all clients and the standard deviation ; (7-3) The server obtains the standard deviation radius of the first client according to the parameter direction distance MDD of the first client obtained in step (7-1) and the mean value and the standard deviation of the parameter direction distances MDD of all the clients obtained in step (7-2) (7-4) The server obtains the parameter direction distance MDD of the second client according to the standard deviation radius of the first client obtained in step (7-3) and the mean value and the standard deviation of the parameter direction distances MDD of all the clients obtained in step (7-2)​​ (7-4) The server filters out the clients with a standard deviation radius clients with a client number greater than or equal to a preset threshold value, to obtain the filtered clients.

6. The federated cognitive collaborative based power system cross-domain network security modeling method according to claim 5, characterized in that, Step (7-1) is to use the following calculation formula: ; wherein , is a global threat identification model the total number of parameters, is a function representing a statistical quantity.

7. The federated cognitive collaborative based power system cross-domain network security modeling method according to claim 6, characterized in that, Step (7-3) is specifically to obtain the standard deviation radius of the i-th client by using the following formula: ​ 。 8. The federated cognitive collaborative based power system cross-domain network security modeling method according to claim 7, characterized in that, The clustering cluster is obtained in step (7-4) The clustering cluster threat identification model corresponding to the t+1th round of training is calculated by the following formula: ; wherein represents the number of local data samples of the th client.

9. The federated cognitive collaborative based power system cross-domain network security modeling method according to claim 8, characterized in that, The server in step (7) performs clustering on the threat identification model of the t+1th round of training corresponding to all clustering clusters An average aggregation process is performed to obtain a global threat identification model of the t+1th round of training The following formula is used: 。 10. A federated cognitive collaborative based power system cross-domain network security modeling system, which is applied in a power system comprising a plurality of clients and a server, and n is any natural number, characterized in that, The power system cross-domain network security modeling system comprises the following steps: The first module, which is set in the first... A client is used to acquire network traffic, system log data, and multi-dimensional features reflecting communication behavior and equipment status from network devices and data acquisition and monitoring control / energy management systems (SCADA / EMS) in its local domain. One-hot encoding is used to process the categorical features of these multi-dimensional features, and Z-score normalization is used to process the numerical features. The processing results are then used to construct a local dataset. ,in ; The second module, located on the server, is used to initialize the global threat identification model. Clustering Threat Identification Model Set And initialize the j-th cluster threat identification model in the cluster threat identification model set for the k-th client. and the initialized global threat identification model and the Threat identification model for clusters to which each client belongs Issued to the One client; among which The number of clusters is preset, and j represents the number of clusters. The index of the cluster threat identification model to which each client belongs in the cluster threat identification model set; a third module configured to set a training round counter t = 1 at the first client a third module configured to set a training round counter t = 1 at the first client A fourth module is arranged in the first client, used for judging whether t is greater than a preset training round threshold T, if yes, entering an eighth module, otherwise setting t=t+1 and entering a fifth module. A fourth module is arranged in the first client, used for judging whether t is greater than a preset training round threshold T, if yes, entering an eighth module, otherwise setting t=t+1 and entering a fifth module. A fifth module is arranged in the first client, configured to acquire the global threat identification model of the first round of training and the cluster threat identification model to which the global threat identification model belongs, and initialize the local threat identification model of the first round of training according to the acquired global threat identification model and cluster threat identification model. ​​​​​​​​ a sixth module configured to train, using the local dataset of the sixth module, the local initial model obtained by the fifth module to obtain a trained local threat identification model ​​​​​​​​​​​​ The seventh module, located on the server, is used to determine the parameters of the local threat identification model uploaded by all clients. All clients are identified and filtered to exclude potentially abnormal clients, and multiple filtered clients are obtained. These filtered clients are then compared with the global threat identification model. The similarity distance matrix between them is used, and a hierarchical clustering algorithm is employed to divide all filtered clients into S clusters. For each cluster Perform intra-cluster weighted average aggregation to generate the cluster. The corresponding threat recognition model trained in the (t+1)th round. And for all clusters, the threat identification model trained in round t+1 is { Perform average aggregation to generate the global threat identification model trained in the (t+1)th round. And this global threat identification model and the threat identification model trained in the (t+1)th round for all clusters { } Send it to the i-th client and return to step (4), where i∈[1,S]; The eighth module, which is set in the first... One client is used to utilize the global threat identification model trained in round T. The network traffic and system log data obtained from the first module are used for threat identification to output attack categories and threat levels.

Citation Information

Patent Citations

  • Network security access method and device, equipment and storage medium

    CN116405262A

  • Determination of cybersecurity recommendations

    US20190098039A1