A power system cross-domain network security modeling method and system based on federal cognitive collaboration
By adopting a federated cognitive collaborative approach to cross-domain cybersecurity modeling of power systems, the problems of data privacy leakage and insufficient model generalization were solved, and the accuracy and robustness of threat identification were improved in highly independent and co-distributed scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-19
- Publication Date
- 2026-03-27
AI Technical Summary
Existing cross-domain cybersecurity modeling methods for power systems suffer from risks of data privacy leakage, insufficient model generalization ability, and distortion of threat detection results. In particular, they are difficult to effectively identify and filter abnormal nodes in highly independent and identically distributed scenarios.
By adopting a federated cognitive collaboration approach, local data preprocessing and model training are combined with distribution similarity calculation, hierarchical clustering and parameter directional distance filtering to achieve distribution adaptation and robustness improvement of the cross-domain threat identification model.
It enables threat modeling without cross-domain transmission of sensitive data, improves the model's generalization ability and robustness in non-independent and identically distributed scenarios, and ensures identification accuracy and security.
Smart Images

Figure CN121547368B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of power system network security and artificial intelligence, and more particularly relates to a power system cross-domain network security modeling method and system based on federal cognitive collaboration. BACKGROUND
[0002] With the continuous improvement of informationization and intelligentization level of the power system, the attack threat faced by the power system cyberspace is becoming increasingly serious. Attackers can use cross-domain penetration, horizontal movement and other methods to implement network attacks on business systems and management systems of power generation, power transmission, power transformation, power distribution and power consumption. Therefore, power system cross-domain network security modeling and threat identification for multi-level dispatch centers and multi-regional power grids have become an urgent need.
[0003] The commonly used power system cross-domain network security modeling method at present is to centralize the network flow, system log and other security data of each provincial dispatch center, regional dispatch center and related business domain to a central server, and the central server uses a unified machine learning or deep learning model for modeling analysis to identify typical network attack behaviors such as denial of service attacks and data injection attacks. In order to improve the identification accuracy, the existing method often needs to collect large-scale raw data of each domain for a long time, and store and train offline at the central node. After training is completed, the model is issued to each level of dispatch center for deployment and use.
[0004] However, the above existing power system cross-domain network security modeling method still has some defects that cannot be ignored:
[0005] First, centralized modeling needs to transmit and store a large amount of raw security data across domains, and power system operation data often involves critical infrastructure and user privacy information. Such cross-domain aggregation not only increases the risk of data leakage, but also is difficult to meet the compliance requirements of existing power data security management, privacy protection and cross-regional data flow;
[0006] Second, the centralized method usually assumes that the data of different dispatch regions satisfies independent and identically distributed, but in reality, there are significant differences in power grid structure, load characteristics, device types and attack behavior patterns among regions, resulting in insufficient generalization ability of the model in the highly non-independent and identically distributed scenario;
[0007] Third, the centralized modeling lacks an effective mechanism to identify and filter malicious or abnormal clients, and is easily affected by malicious data updates uploaded by abnormal nodes, causing deviation or distortion of the identification results of the threat model. SUMMARY
[0008] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a method and system for cross-domain cybersecurity modeling of power systems based on federated cognitive collaboration. Its purpose is to solve the technical problems of existing cross-domain cybersecurity modeling methods for power systems, which, due to their centralized aggregation approach, result in the cross-domain transmission of large amounts of raw security data, leading to data privacy risks and difficulty in meeting compliance requirements for cross-domain power data flow; the technical problem that, due to significant differences in network structure, load characteristics, and attack patterns among different dispatch centers, existing methods cannot effectively handle highly non-independent and identically distributed data, resulting in insufficient model generalization performance; and the technical problem that, due to the lack of effective identification and protection mechanisms for malicious or abnormal nodes, the model is easily contaminated by abnormal updates during the aggregation process, leading to distorted threat detection results.
[0009] To achieve the above objectives, according to one aspect of the present invention, a cross-domain network security modeling method for power systems based on federated cognitive collaboration is provided, which is applied to applications including... In a power system for both clients and servers, and For any natural number, the cross-domain network security modeling method for this power system includes the following steps:
[0010] (1) No. Each client obtains network traffic, system log data, and multi-dimensional features reflecting communication behavior and equipment status from its local network devices and Supervisory Control and Data Acquisition / Energy Management System (SCADA / EMS). One-hot encoding is used to process the categorical features within these multi-dimensional features, while Z-score normalization is used to process the numerical features. The results are then used to construct a local dataset. ,in ;
[0011] (2) Server initializes global threat identification model Clustering Threat Identification Model Set And initialize the j-th cluster threat identification model in the cluster threat identification model set for the k-th client. and the initialized global threat identification model and the Threat identification model for clusters to which each client belongs Issued to the One client; among which The number of clusters is preset, and j represents the number of clusters. The index of the cluster threat identification model to which each client belongs in the cluster threat identification model set;
[0012] (3) No. Each client sets the training round counter t=1;
[0013] (4) No. Each client determines whether t is greater than the preset training round threshold T. If it is, proceed to step (8); otherwise, set t=t+1 and proceed to step (5).
[0014] (5) No. The client obtains the first Global threat identification model trained in rounds and its associated cluster threat identification model And based on the acquired global threat identification model Clustering Threat Identification Model Initialize the first Local threat identification model trained in rounds :
[0015] (6) No. Each client uses its local dataset The first one obtained in step (5) Local initial model during round training Training is performed to obtain a well-trained local threat identification model. Calculate the trained local threat identification model With the Global threat identification model trained in rounds The distribution similarity between them is used to construct the first distribution similarity. Similarity distance matrix during round training And the trained local threat identification model Parameters and similarity distance matrix Uploaded to the server;
[0016] (7) The server identifies the local threat model based on all client uploads. The parameters of} are used to identify and filter all clients to exclude potentially abnormal clients, and multiple filtered clients are obtained. Based on all filtered clients and the global threat identification model... The similarity distance matrix between them is used, and a hierarchical clustering algorithm is employed to divide all filtered clients into S clusters. For each cluster Perform intra-cluster weighted average aggregation to generate the cluster. The corresponding threat recognition model trained in the (t+1)th round. And for all clusters, the threat identification model trained in round t+1 is { Perform average aggregation to generate the global threat identification model trained in the (t+1)th round. And this global threat identification model and the threat identification model trained in the (t+1)th round for all clusters { } Send it to the i-th client and return to step (4), where i∈[1,S];
[0017] (8) All clients utilize the global threat identification model trained in round T. Perform threat identification on local real-time network traffic and system log data to output attack categories and threat levels.
[0018] Preferably, the training round threshold T ranges from 100 to 500.
[0019] Preferably, step (5) involves initializing the local threat identification model using the following formula. :
[0020] Preferably, in step (6) the first... Each client uses a trained local threat identification model. With the Global threat identification model trained in rounds The process of constructing a similarity distance matrix based on the distributional similarity between threats involves the following steps: First, the k-th client randomly selects samples from each category in the local dataset; then, for each category, each sample from that category is input into the updated local threat identification model. To obtain the first output vector, each sample of that category is fed into the global threat identification model. To obtain the second output vector, the Euclidean distance between the first and second output vectors corresponding to the sample is calculated. Then, the average Euclidean distances of all samples in that category are averaged to obtain the average Euclidean distance for that category. Finally, the average Euclidean distances of all categories are aggregated to form the first output vector. Individual Client and Global Threat Identification Model Similarity distance matrix between .
[0021] Preferably, in step (7), the server uses the local threat identification model uploaded by all clients { The parameters of} are used to identify and filter all clients to exclude potentially abnormal clients, and the process of obtaining multiple filtered clients includes the following sub-steps:
[0022] (7-1) The server calculates the parameter direction distance MDD of the i-th client ;
[0023] (7-2) The server obtains the mean value and the standard deviation of the parameter direction distance MDD of all clients ;
[0024] (7-3) The server obtains the standard deviation radius of the i-th client according to the parameter direction distance MDD of the i-th client obtained in step (7-1) and the mean value and the standard deviation of the parameter direction distance MDD of all clients obtained in step (7-2)
[0025] (7-4) The server filters out the clients whose standard deviation radius is greater than or equal to a preset threshold value to obtain a plurality of filtered clients.
[0026] Preferably, step (7-1) is calculated by the following formula:
[0027] ;
[0028] wherein, n is the total number of parameters of the global threat identification model, is a function of the statistical quantity. Preferably, step (7-3) is to obtain the standard deviation radius of the i-th client by the following formula:
[0029]
[0030] . Preferably, in step (7-4), the clustering cluster threat identification model corresponding to the t+1-th round of training of the clustering cluster is obtained by the following formula:
[0031] ;
[0032] wherein, m represents the number of local data samples of the i-th client.
[0033] Preferably, in step (7), the server performs average aggregation processing on the clustering cluster threat identification models corresponding to the t+1-th round of training of all clustering clusters to obtain the global threat identification model of the t+1-th round of training The following formula is used:
[0034] According to another aspect of the present invention, a cross-domain network security modeling system for power systems based on federated cognitive collaboration is provided, which is applied in applications including... In a power system for both clients and servers, and For any natural number, the power system cross-domain network security modeling system includes the following steps:
[0035] The first module, which is set in the first... A client is used to acquire network traffic, system log data, and multi-dimensional features reflecting communication behavior and equipment status from network devices and Supervisory Control and Data Acquisition / Energy Management System (SCADA / EMS) in its local domain. One-hot encoding is used to process the categorical features of these multi-dimensional features, and Z-score normalization is used to process the numerical features. The processing results are then used to construct a local dataset. ,in ;
[0036] The second module, located on the server, is used to initialize the global threat identification model. Clustering Threat Identification Model Set And initialize the j-th cluster threat identification model in the cluster threat identification model set for the k-th client. and the initialized global threat identification model and the Threat identification model for clusters to which each client belongs Issued to the One client; among which The number of clusters is preset, and j represents the number of clusters. The index of the cluster threat identification model to which each client belongs in the cluster threat identification model set;
[0037] The third module, which is set in the first... One client is used to set the training round counter t=1;
[0038] The fourth module, which is set in the first... One client is used to determine whether t is greater than the preset training round threshold T. If it is, the process proceeds to the eighth module; otherwise, t is set to t+1 and the process proceeds to the fifth module.
[0039] The fifth module, which is set in the first... The client is used to obtain the first one. Global threat identification model trained in rounds and its associated cluster threat identification model And based on the acquired global threat identification model Clustering Threat Identification Model Initialize the first Local threat identification model trained in rounds :
[0040] The sixth module, which is set in the first... One client, used to use its local dataset The fifth module yielded the first Local initial model during round training Training is performed to obtain a trained local threat identification model. Calculate the trained local threat identification model With the Global threat identification model trained in rounds The distribution similarity between them is used to construct the first distribution similarity. Similarity distance matrix of round training And the trained local threat identification model Parameters and similarity distance matrix Uploaded to the server;
[0041] The seventh module, located on the server, is used to identify local threat models uploaded by all clients. The parameters of} are used to identify and filter all clients to exclude potentially abnormal clients, and multiple filtered clients are obtained. Based on all filtered clients and the global threat identification model... The similarity distance matrix between them is used, and a hierarchical clustering algorithm is employed to divide all filtered clients into S clusters. For each cluster Perform intra-cluster weighted average aggregation to generate the cluster. The corresponding threat recognition model trained in the (t+1)th round. And for all clusters, the threat identification model trained in round t+1 is { Perform average aggregation to generate the global threat identification model trained in the (t+1)th round. And this global threat identification model and the threat identification model trained in the (t+1)th round for all clusters { The message is sent to the i-th client and returned to the fourth module, where i∈[1,S];
[0042] The eighth module, which is set in the first... a client, configured to utilize the global threat identification model trained in the Tth round of training threat identification is performed on the network traffic and system log data obtained by the first module to output an attack category and a threat level.
[0043] Overall, compared with the prior art, the above technical solutions conceived by the present application can achieve the following beneficial effects:
[0044] 1. The present application adopts the local data construction and preprocessing mechanism in step (1) and the local training and model parameter and similarity distance matrix uploading method in step (6), which avoids the transmission of sensitive data such as raw network traffic and system logs across the dispatching center, realizes the privacy protection mode of "available but invisible", and therefore can complete cross-domain threat modeling without touching the compliance red line of cross-domain transmission of power data, solving the technical problems of large data privacy leakage risk and insufficient compliance in the existing centralized method.
[0045] 2. The present application adopts the distributed similarity calculation mechanism in step (6) and the hierarchical clustering and cross-cluster customized aggregation method based on the similarity distance matrix in step (7), which can automatically identify the data distribution differences between the dispatching centers and divide the clients with similar distribution into the same clustering cluster for differential model training and aggregation, so as to significantly improve the generalization ability and adaptive ability of the threat identification model in the highly Non-IID scene, thereby effectively solving the technical problem of low identification accuracy of the existing centralized method under data heterogeneity.
[0046] 3. The present application adopts the parameter directional distance calculation and standard deviation radius outlier filtering mechanism in steps (7-1) to (7-4), and combines the filtering and then aggregation strategy, which can effectively identify and exclude malicious nodes or low-quality nodes that upload abnormal model parameters, so as to ensure that the update of the clustering cluster threat identification model and the global threat identification model comes from reliable clients, thereby significantly improving the robustness of model aggregation and solving the technical problem that the existing method is easily polluted by abnormal nodes, leading to distortion of the threat detection result.
[0047] 4. The present application adopts the dual-source parameter fusion initialization mechanism of the global threat identification model combined with the clustering cluster threat identification model in step (5) and the multi-round federal iteration and update strategy in step (7), which can fuse knowledge from different regions before local training of the client and continuously adapt and evolve during the training process, thereby breaking through the bottleneck of single client limited by its own data, realizing cross-domain cognitive collaboration and knowledge sharing among multiple regional, multi-level dispatching centers. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 is a flow chart of the federated cognitive collaborative power system cross-domain network security modeling method of the present application. DETAILED DESCRIPTION
[0049] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. In addition, the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.
[0050] It should be noted that in the description of the embodiments of the present application, the terms "comprise", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, article or device. Without more limitations, the element defined by the statement "comprising a" does not exclude the existence of another identical element in the process, article or device comprising the element. The terms "upper", "lower" and the like indicate the orientation or positional relationship shown in the drawings, and are only used to facilitate the description of the present application and simplify the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. For those of ordinary skill in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.
[0051] In addition, the technical solutions of the various embodiments of the present application can be combined with each other, but it must be based on the fact that a person of ordinary skill in the art can implement it, and when the combination of technical solutions appears contradictory or unimplementable, it should be considered that the combination of technical solutions does not exist, nor is it within the scope of protection required by the present application.
[0052] The basic idea of the present application is to provide a power system cross-domain network security modeling method based on federal cognitive collaboration, which introduces a federal cognitive collaboration mechanism to solve the problem of model generalization performance decline caused by data Non-Independent and Identically Distributed (Non-IID) and abnormal nodes in power system cross-domain network security modeling under the premise of protecting data privacy and complying with power system privacy protocols. Specifically, this method uses hierarchical clustering and model fusion strategies to enable clients to draw on global and other cluster knowledge during local training, breaking through the limitations of their own data distribution; the server side filters reliable clients and generates clustering cluster threat identification models and global threat identification models that are adaptive to different data distributions through outlier filtering and clustering aggregation, thereby improving the accuracy and robustness of threat identification and achieving dynamic adaptation to network threat environments.
[0053] As shown in Figure 1 , the present application provides a power system cross-domain network security modeling method based on federal cognitive collaboration, which is applied in a power system including a client (which is set in each level of power dispatching center) and a server, and is any natural number, the power system cross-domain network security modeling method includes the following steps:
[0054] (1) the th client obtains network traffic, system log data, and multi-dimensional features reflecting communication behavior and device state from network devices and supervisory control and data acquisition / energy management systems (SCADA / EMS) in its local region, processes categorical features in the multi-dimensional features using one-hot encoding, processes numerical features in the multi-dimensional features using Z-score standardization, and constructs the processed results into a local dataset , wherein ;
[0055] (2) the server initializes a global threat identification model and a clustering cluster threat identification model set , initializes the kth client to belong to the jth clustering cluster threat identification model in the clustering cluster threat identification model set , and issues the initialized global threat identification model and the clustering cluster threat identification model to which the kth client belongs to the kth client One client; among which The preset number of clusters is defined, ranging from 3 to 10, preferably 5, where j represents the number of clusters. The index of the cluster threat identification model to which each client belongs in the cluster threat identification model set;
[0056] During the initialization process of this step, the threat identification model for all client clusters is the same;
[0057] The advantage of this step (2) is that by uniformly initializing the global threat identification model and the cluster threat identification model on the server side, the consistency and comparability of the initial threat identification models of each client are ensured, providing a unified starting point for subsequent clustering and differential evolution based on distribution similarity, and improving the stability of the aggregation results.
[0058] (3) No. Each client sets the training round counter t=1;
[0059] (4) No. Each client determines whether t is greater than the preset training round threshold T. If it is, proceed to step (8); otherwise, set t=t+1 and proceed to step (5).
[0060] Specifically, the training round threshold T in this step ranges from 100 to 500, preferably 200.
[0061] (5) No. The client obtains the first Global threat identification model trained in rounds and its associated cluster threat identification model And based on the acquired global threat identification model Clustering Threat Identification Model Initialize the first Local threat identification model trained in rounds :
[0062] Specifically, this step initializes the local threat identification model using the following formula. :
[0063] The advantage of this step (5) is that by using the global threat identification model and the cluster threat identification model of round t to initialize the local model, the local training can take into account both global knowledge and local knowledge under the same data distribution, thereby improving the convergence speed and identification performance of the model on local data.
[0064] (6) No. Each client uses its local dataset The first one obtained in step (5) Local initial model during round training Training is performed to obtain a trained local threat identification model. Calculate the trained local threat identification model With the Global threat identification model trained in rounds The distribution similarity between them is used to construct the first distribution similarity. Similarity distance matrix of round training And the trained local threat identification model Parameters and similarity distance matrix Uploaded to the server;
[0065] The advantage of this step (6) is that by completing the threat identification model training and distribution similarity calculation locally, and only uploading the model parameters and similarity distance matrix to the server, the original data is not leaked, and refined and effective statistical information is provided for clustering and robust aggregation on the server side.
[0066] In this step, the first Each client uses a trained local threat identification model. With the Global threat identification model trained in rounds The process of constructing a similarity distance matrix based on the distributional similarity between threats involves the following steps: First, the k-th client randomly selects samples from each category in the local dataset; then, for each category, each sample from that category is input into the updated local threat identification model. To obtain the first output vector, each sample of that category is fed into the global threat identification model. To obtain the second output vector, the Euclidean distance between the first and second output vectors corresponding to the sample is calculated. Then, the average Euclidean distances of all samples in that category are averaged to obtain the average Euclidean distance for that category. Finally, the average Euclidean distances of all categories are aggregated to form the first output vector. Individual Client and Global Threat Identification Model Similarity distance matrix between .
[0067] (7) The server identifies the local threat model based on all client uploads. The parameters of} are used to identify and filter all clients to exclude potentially abnormal clients, and multiple filtered clients are obtained. Based on all filtered clients and the global threat identification model... The similarity distance matrix between them is used, and a hierarchical clustering algorithm is employed to divide all filtered clients into S clusters. For each cluster performing intra-cluster weighted average aggregation processing to generate the clustering cluster corresponding threat identification model of the t+1th round of training , and performing average aggregation on threat identification models of the t+1th round of training corresponding to all clustering clusters to generate a global threat identification model of the t+1th round of training , and delivering the global threat identification model of the t+1th round of training and the threat identification models of the t+1th round of training corresponding to all clustering clusters to the i-th client, and returning to step (4), where i ∈ [1, S]; , and delivering the global threat identification model of the t+1th round of training and the threat identification models of the t+1th round of training corresponding to all clustering clusters to the i-th client, and returning to step (4), where i ∈ [1, S]; , and delivering the global threat identification model of the t+1th round of training and the threat identification models of the t+1th round of training corresponding to all clustering clusters to the i-th client, and returning to step (4), where i ∈ [1, S]; , and delivering the global threat identification model of the t+1th round of training and the threat identification models of the t+1th round of training corresponding to all clustering clusters to the i-th client, and returning to step (4), where i ∈ [1, S];
[0068] In this step (7), the server identifies and filters all clients according to the parameters of the local threat identification models of all clients uploaded by the clients, to exclude potential abnormal clients, and obtains a plurality of filtered clients, and this process includes the following sub-steps:
[0069] (7-1) The server calculates the parameter directional distance MDD of the i-th client;
[0070] This step uses the following calculation formula:
[0071] wherein, n is the total number of parameters of the global threat identification model, and the parameter directional distance reflects the consistency of the parameter update direction of the client threat identification model and the global threat identification model;
[0072] (7-2) The server obtains the mean value and standard deviation of the parameter directional distance MDD of all clients;
[0073] (7-3) The server obtains the standard deviation radius of the i-th client according to the parameter directional distance MDD of the i-th client obtained in step (7-1) and the mean value and standard deviation of the parameter directional distance MDD of all clients obtained in step (7-2);
[0074] This step specifically obtains the standard deviation radius of the i-th client using the following formula:
[0075] (7-4) Server filters out standard deviation radius Clients with a value greater than or equal to a preset threshold are selected to obtain a filtered list of clients.
[0076] Specifically, the preset threshold value ranges from 2 to 4, preferably 3.
[0077] The advantage of the above sub-steps (7-1) to (7-4) is that by introducing statistical indicators of parameter direction distance and its standard deviation radius, abnormal nodes whose parameter update direction deviates significantly from most clients can be effectively identified and removed before aggregation, thus avoiding malicious or low-quality model updates from polluting the global threat identification model and the cluster threat identification model, thereby improving the security and robustness of the model aggregation process.
[0078] In this step (7), the cluster is obtained. The corresponding cluster threat identification model trained in the (t+1)th round The following formula is used:
[0079] in Indicates the first Number of local data samples per client.
[0080] In step (7), the server trains the cluster threat identification model for all clusters in round t+1. Perform average aggregation to obtain the global threat identification model trained in round t+1. The following formula is used:
[0081] (8) No. The client utilizes the global threat identification model trained in round T. Threat identification is performed on the network traffic and system log data obtained in step (1) to output the attack category and threat level.
[0082] The advantage of this step (8) is that by deploying the global threat identification model after the convergence of the Tth round of training on the clients of dispatch centers at all levels, threat identification can be performed in real time without increasing the additional communication burden, thereby improving the speed and accuracy of the power system in detecting various cross-domain network attacks.
[0083] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A federated cognitive collaboration based power system cross-domain network security modeling method, applied in a power system comprising a plurality of clients and a server, and n is any natural number, characterized in that, The power system cross-domain network security modeling method comprises the following steps: (1) the first The network flow, system log data, and multi-dimensional features reflecting communication behavior and device status are obtained by the first client from the network equipment and the data acquisition and monitoring control system / energy management system (SCADA / EMS) in the local region of the client, the categorical features in the multi-dimensional features are processed by using a one-hot encoding method, the numerical features in the multi-dimensional features are processed by using a Z-score standardization method, and the processed results are constructed as a local data set , wherein ; (2) Server initializes global threat identification model Clustering Threat Identification Model Set And initialize the j-th cluster threat identification model in the cluster threat identification model set for the k-th client. and the initialized global threat identification model and the Threat identification model for clusters to which each client belongs Issued to the One client; among which The number of clusters is preset, and j represents the number of clusters. The index of the cluster threat identification model to which each client belongs in the cluster threat identification model set; (3) the first client sets a training round counter t = 1; (4) the first client judges whether t is greater than a preset training round threshold T, if yes, it goes to step (8), otherwise, it sets t = t + 1 and goes to step (5); (4) the first client judges whether t is greater than a preset training round threshold T, if yes, it goes to step (8), otherwise, it sets t = t + 1 and goes to step (5); (5) the first client obtains the global threat identification model of the first round of training and the cluster threat identification model to which it belongs, and initializes the local threat identification model of the first round of training according to the obtained global threat identification model and cluster threat identification model : (6) No. Each client uses its local dataset The first one obtained in step (5) Local initial model during round training Training is performed to obtain a well-trained local threat identification model. Calculate the trained local threat identification model With the Global threat identification model trained in rounds The distribution similarity between them is used to construct the first distribution similarity. Similarity distance matrix during round training And the trained local threat identification model Parameters and similarity distance matrix Uploaded to the server; (7) The server identifies the local threat model based on all client uploads. The parameters of} are used to identify and filter all clients to exclude potentially abnormal clients, and multiple filtered clients are obtained. Based on all filtered clients and the global threat identification model... The similarity distance matrix between them is used, and a hierarchical clustering algorithm is employed to divide all filtered clients into S clusters. For each cluster Perform intra-cluster weighted average aggregation to generate the cluster. The corresponding threat recognition model trained in the (t+1)th round. And for all clusters, the threat identification model trained in round t+1 is { Perform average aggregation to generate the global threat identification model trained in the (t+1)th round. And this global threat identification model and the threat identification model trained in the (t+1)th round for all clusters { } Send it to the i-th client and return to step (4), where i∈[1,S]; (8) All clients utilize the global threat identification model trained in round T Threat identification on local real-time network traffic and system log data to output attack categories and threat levels.
2. The federated cognitive collaboration based power system cross-domain network security modeling method of claim 1, wherein, The value range of the training round threshold T is 100 to 500.
3. The federated cognitive collaboration based power system cross-domain network security modeling method of claim 2, wherein, Step (5) is to initialize the local threat identification model using the following equation : 。 4. The federated cognitive collaboration based power system cross-domain network security modeling method of claim 3, wherein, The process of constructing the similarity distance matrix between the kth client and the global threat identification model trained in the (n-1)th round is as follows: first, the kth client randomly selects samples from each category in the local data set; then, for each category, each sample of the category is input into the updated local threat identification model to obtain a first output vector, and each sample of the category is input into the global threat identification model to obtain a second output vector, and the Euclidean distance between the first output vector and the second output vector corresponding to the sample is calculated, then the average Euclidean distance corresponding to the category is obtained by averaging the Euclidean distances corresponding to all samples in the category; finally, the average Euclidean distances corresponding to all categories are collected, that is, the similarity distance matrix between the kth client and the global threat identification model trained in the (n-1)th round is obtained. 5. The federated cognitive collaborative based power system cross-domain network security modeling method according to claim 4, characterized in that, In step (7), the server uses the local threat identification model uploaded by all clients { The parameters of} are used to identify and filter all clients to exclude potentially abnormal clients, and the process of obtaining multiple filtered clients includes the following sub-steps: (7-1) The server calculates the parameter direction distance of each client ; (7-2) The server obtains the mean of the parameter directional distance MDD of all clients and the standard deviation ; (7-3) The server obtains the parameter directional distance MDD of the first client according to the parameter directional distance MDD of the first client obtained in step (7-1) and the mean value and the standard deviation of the parameter directional distances MDD of all the clients obtained in step (7-2) obtains the standard deviation radius of the first client ; (7-4) The server filters out the clients with a standard deviation radius clients with a client number greater than or equal to a preset threshold value, to obtain the filtered clients.
6. The federated cognitive collaborative based power system cross-domain network security modeling method according to claim 5, characterized in that, Step (7-1) is to use the following calculation formula: ; wherein , is a global threat identification model the total number of parameters, denotes a function of the statistical quantity.
7. The federated cognitive collaborative based power system cross-domain network security modeling method according to claim 6, characterized in that, Step (7-3) is specifically to obtain the standard deviation radius of the i-th client by using the following formula: 。 8. The federated cognitive collaborative based power system cross-domain network security modeling method according to claim 7, characterized in that, The clustering cluster is obtained in step (7-4) The clustering cluster threat identification model corresponding to the t+1th round of training is calculated by the following formula: ; wherein represents the number of local data samples of the th client.
9. The federated cognitive collaborative based power system cross-domain network security modeling method according to claim 8, characterized in that, The server in step (7) performs clustering on the threat identification model of the t+1th round of training corresponding to all clustering clusters An average aggregation process is performed to obtain a global threat identification model of the t+1th round of training The following formula is used: 。 10. A federated cognitive collaborative based power system cross-domain network security modeling system, which is applied in a power system comprising a plurality of clients and a server, and n is any natural number, characterized in that, The power system cross-domain network security modeling system comprises the following steps: The first module, which is set in the first... A client is used to acquire network traffic, system log data, and multi-dimensional features reflecting communication behavior and equipment status from network devices and data acquisition and monitoring control / energy management systems (SCADA / EMS) in its local domain. One-hot encoding is used to process the categorical features of these multi-dimensional features, and Z-score normalization is used to process the numerical features. The processing results are then used to construct a local dataset. ,in ; The second module, located on the server, is used to initialize the global threat identification model. Clustering Threat Identification Model Set And initialize the j-th cluster threat identification model in the cluster threat identification model set for the k-th client. and the initialized global threat identification model and the Threat identification model for clusters to which each client belongs Issued to the One client; among them The number of clusters is preset, and j represents the number of clusters. The index of the cluster threat identification model to which each client belongs in the cluster threat identification model set; a third module configured to set a training round counter t = 1 at the first client a third module configured to set a training round counter t = 1 at the first client A fourth module is arranged in the first client, used for judging whether t is greater than a preset training round threshold T, if yes, entering an eighth module, otherwise setting t=t+1 and entering a fifth module. A fourth module is arranged in the first client, used for judging whether t is greater than a preset training round threshold T, if yes, entering an eighth module, otherwise setting t=t+1 and entering a fifth module. A fifth module is arranged in the first client, configured to acquire the global threat identification model of the first round of training and the cluster threat identification model to which the global threat identification model belongs, and initialize the local threat identification model of the first round of training according to the acquired global threat identification model and cluster threat identification model. The sixth module, which is set in the first... One client, used to use its local dataset The fifth module yielded the first Local initial model during round training Training is performed to obtain a trained local threat identification model. Calculate the trained local threat identification model With the Global threat identification model trained in rounds The distribution similarity between them is used to construct the first distribution similarity. Similarity distance matrix of round training And the trained local threat identification model Parameters and similarity distance matrix Uploaded to the server; The seventh module, located on the server, is used to identify local threat models uploaded by all clients. The parameters of} are used to identify and filter all clients to exclude potentially abnormal clients, and multiple filtered clients are obtained. Based on all filtered clients and the global threat identification model... The similarity distance matrix between them is used, and a hierarchical clustering algorithm is employed to divide all filtered clients into S clusters. For each cluster Perform intra-cluster weighted average aggregation to generate the cluster. The corresponding threat recognition model trained in the (t+1)th round. And for all clusters, the threat identification model trained in round t+1 is { Perform average aggregation to generate the global threat identification model trained in the (t+1)th round. And this global threat identification model and the threat identification model trained in the (t+1)th round for all clusters { The message is sent to the i-th client and returned to the fourth module, where i∈[1,S]; An eighth module is arranged in the first server, configured to use the global threat identification model trained in the Tth round to identify threats of network traffic and system log data obtained by the first module. The first module obtains network traffic and system log data, and the eighth module uses the global threat identification model trained in the Tth round to identify threats of the network traffic and system log data, to output an attack category and a threat level.
Citation Information
Patent Citations
Network security access method and device, equipment and storage medium
CN116405262A
Determination of cybersecurity recommendations
US20190098039A1