Electric power communication network security risk identification method based on graph neural network

By adopting the graph neural network method in the power communication network, modeling topology and attribute information into graph models, the problem of difficulty in integrating and utilizing network information in the existing technology is solved, and more accurate security risk identification and evaluation is achieved.

CN119940913APending Publication Date: 2025-05-06INFORMATION & COMM COMPANY OF QINGHAI ELECTRIC POWER +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411910953.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate and utilize network topology and attribute information in the identification of security risks of power communication networks, resulting in difficulty in modeling complex power communication networks, and data integration has integrity, quality and noise problems.

Method used

Using a graph neural network-based method, the topological information and attribute information of the power communication network are modeled into a graph model. Through the structural encoder and attribute encoder, the embedded representation of nodes and attributes is learned, combined with the abnormality detection model, the abnormality score is output to identify the security risks of the power communication network.

Benefits of technology

A more comprehensive and accurate power communication network security risk assessment is achieved, which can effectively capture network complexity and key attributes and improve the accuracy and efficiency of risk identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940913A_ABST
    Figure CN119940913A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power communication network security risk identification method based on a graph neural network. The method comprises the following steps: preprocessing original data to obtain an optical cable feature matrix and an adjacent matrix; inputting the optical cable characteristic matrix and the adjacent matrix into an encoder, learning the embedded representation of the node through a structure encoder, and obtaining the embedded representation of the attribute through an attribute encoder; respectively inputting the embedded representation of the node and the embedded representation of the attribute into a structure decoder and an attribute decoder to reconstruct an original feature matrix and an adjacent matrix, and respectively obtaining a reconstruction error of the feature matrix and a reconstruction error of the adjacent matrix; outputting an anomaly score through an anomaly detection model; and obtaining a safety risk identification result of the power communication network according to the abnormal score. According to the method, the complexity of the network topology is considered, and key attributes such as the performance parameters of the optical cable are integrated, so that the security risk of the electric power communication network can be evaluated more comprehensively and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of smart grids and relates to power communication network security risk identification technology, and specifically to a power communication network security risk identification method based on graph neural network. Background Art

[0002] As an important part of the transformation of smart grid, the strategic importance of power communication network is becoming increasingly prominent. Through the continuous construction of the new generation of communication management system (SG-TMS2.0), the power communication system has initially realized informatization and automation. The purpose of power communication network security risk analysis and identification is to identify and evaluate the optical cables that may affect the stability and reliability of the power communication network, expand or expand the optical cables with risks, and adjust the services they carry when conditions are not met to ensure the stable operation of the network. With the access of a large number of new energy plants and stations, the business volume continues to grow. In addition, the special requirements of the power communication profession for business arrangements and optical cable redundancy have increased the workload of communication network security risk analysis. The existing working model has been unable to cope with this challenge.

[0003] At present, the power industry has carried out extensive research work in the field of communication network security risk identification, which can be roughly divided into the following three categories. The first is the currently widely used rule-based method (such as the same optical cable for the main and standby and the same optical cable for the optical path). Although it has high accuracy, it has poor flexibility, low recall rate, and it is difficult to accurately describe the rules in complex modes. The second is the method based on indicator statistics and quantitative evaluation, such as the indicator importance quantification method based on the hierarchical analysis method. Although this method of converting attributes into quantitative risk values ​​is more scientific and objective, it often encounters problems such as unreasonable structure of the indicator system and insufficient comprehensiveness of indicators, and it cannot effectively utilize the complex topological structure information of the power communication network. The third is the method based on machine learning, such as the dynamic Bayesian network, DBSCAN clustering and convolutional neural network methods used in the field of power communication network fault prediction and diagnosis. These advanced machine learning methods show good adaptability to environmental changes and higher flexibility and accuracy, and have become the mainstream trend in solving various problems in the field of power communication networks.

[0004] Although the identification methods of power communication network security risks are gradually developing in a scientific and intelligent direction, the identification of security risks is affected by the network topology relationship, and the existing technology does not fully integrate and utilize the rich information contained in the topology and attributes, which brings two significant challenges to the modeling of complex power communication networks: (1) Effectively selecting and representing these features to capture the complexity of the network is a major difficulty. (2) When collecting data, due to the lack of a unified equipment identification standard, it is difficult to integrate multi-source equipment information, which may face problems such as data integrity, data quality, noise, outliers or inconsistency in the data. Summary of the invention

[0005] Purpose of the invention: In order to overcome the deficiencies in the prior art, a method for identifying security risks in power communication networks based on graph neural networks is provided.

[0006] Technical solution: To achieve the above purpose, the present invention provides a method for identifying security risks of power communication network based on graph neural network, comprising the following steps:

[0007] S1: Preprocessing the collected raw data of the power communication network to obtain the optical cable feature matrix and adjacency matrix;

[0008] S2: Input the cable feature matrix and adjacency matrix into the encoder, learn the embedded representation of the node through the structure encoder, and use the attribute encoder to obtain the embedded representation of the attribute;

[0009] S3: Input the embedded representation of the node and the embedded representation of the attribute into the structure decoder and the attribute decoder respectively to reconstruct the original adjacency matrix and the feature matrix, and obtain the reconstruction error of the adjacency matrix and the reconstruction error of the feature matrix respectively;

[0010] S4: Output anomaly scores through anomaly detection models according to the reconstruction errors of feature matrices and adjacency matrices;

[0011] S5: Based on the anomaly score, the security risk identification result of the power communication network is obtained.

[0012] Furthermore, in step S1, relevant features such as overload degree, optical cable scheduling level, port occupancy, bandwidth utilization, number of optical fibers, etc. are calculated and selected, and standardized, and the influence of dimension magnitude is removed to obtain a feature matrix of the optical cable; and the adjacency matrix of the optical cable is obtained by matching the sites connected to the A end and the Z end of each optical cable.

[0013] Furthermore, the step S2 of learning the embedded representation of the node by the structure encoder includes:

[0014] In order to obtain a representative high-level node feature representation, the structure encoder transforms the observed raw node attributes X into a vector representation Z in a low-dimensional latent space. V , expressed as:

[0015] Z V =σ(XW V +b V ) (1)

[0016] Where σ(·) is the activation function ReLU, W V and b V are the weights and biases learned by the encoder.

[0017] Furthermore, obtaining the embedded representation of the attribute by using the attribute encoder in step S2 includes:

[0018] The attribute encoder uses two nonlinear feature transformation layers to map the observed attribute data into a potential attribute embedding representation, expressed as:

[0019]

[0020]

[0021] Among them, W A(1) , b A(1) , W A(2) , b A(2) are the weights and biases learned in the two layers.

[0022] Furthermore, in step S3, A represents the adjacency matrix of the network, is the adjacency matrix reconstructed by the decoder, and the reconstruction error of the adjacency matrix is ​​expressed as:

[0023]

[0024] The reconstruction error is used to evaluate the abnormality of the network structure;

[0025]

[0026]

[0027] Furthermore, in step S3, the attribute reconstruction decoder reconstructs the attribute X of the node, and the reconstruction error of the feature matrix is ​​expressed as:

[0028]

[0029] The reconstruction error is used to evaluate the anomaly of network properties;

[0030]

[0031] Furthermore, the anomaly detection model in step S4 includes a deep autoencoder. Given an input data set X, the deep autoencoder maps the data into a low-dimensional feature space through an encoding function Enc(), and then uses a decompression function Dec() to reconstruct the original data according to the representation of the feature space.

[0032] Furthermore, in step S4, the anomaly detection model is trained to minimize the reconstruction loss:

[0033] min{E[dist(X,Dec(Enc(X)))]} (9)

[0034] In order to unify the reconstruction error, the loss function is as follows:

[0035]

[0036] The final anomaly score is:

[0037]

[0038] In the anomaly detection of graphs, accurate identification cannot be achieved by using only traditional classification and clustering algorithms, and the graph model cannot be processed by using ordinary neural networks. The present invention adopts an anomaly detection algorithm based on an autoencoder. Since the position of the anomaly is on the edge, the present invention extracts the features of the edge attributes and structure of the communication network through the encoder, obtains the feature matrix and the adjacency matrix, and sends them to the decoder to reconstruct the topological structure and edge attributes. The reconstruction error of the node in the encoding and decoding stages is used as the judgment criterion for the abnormal edge in the attribute network, and then the anomaly score and anomaly ranking are obtained. The optical cable feature matrix and the adjacency matrix obtained by data preprocessing the original data of the power communication network are input into the encoder. The system learns the embedded representation of the node through the structure encoder, and uses the attribute encoder to obtain the embedded representation of the attribute. Subsequently, these embedded representations are sent to the decoder to reconstruct the original feature matrix and adjacency matrix. The structure decoder and the attribute decoder work together to accurately capture the complex interactive relationship between the network structure and the node attributes. Finally, the abnormal data is divided according to the anomaly score output by the anomaly detection model. Specifically, when the anomaly score of a sample exceeds a preset threshold, the sample is judged as abnormal data.

[0039] The present invention combines the characteristics of the power communication network itself, adopts a data-driven approach, and constructs an identification system that comprehensively considers multi-dimensional factors such as topological characteristics and performance parameters, and innovatively proposes to model the complex power communication network topology information as a graph model. The core of this method is to map each site in the network as a node in the graph, and the optical cables connecting these sites are naturally converted into edges in the graph. In particular, optical cables with security risks are defined as abnormal optical cables, thereby converting the risk identification problem into a graph anomaly detection problem. On this basis, the present invention conducts an in-depth analysis of the network topology and key attributes. These key attributes may include but are not limited to the bandwidth capacity, transmission delay, and overload of the optical cable. By comprehensively considering these factors, an identification model for power communication network security risks is constructed.

[0040] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0041] 1. Comprehensive construction of security risk identification system for power communication network

[0042] This invention innovatively proposes a power communication network security risk identification system that comprehensively considers multi-dimensional factors such as topological characteristics and performance parameters. This system not only considers the complexity of network topology, but also incorporates key attributes such as the performance parameters of optical cables, so as to more comprehensively and accurately assess the security risks of power communication networks.

[0043] 2. Graph modeling of power communication network topology information

[0044] The present invention models the complex power communication network topology information as a graph model, in which the sites in the network are mapped as nodes in the graph, and the optical cables connecting the sites are converted into edges in the graph. This modeling method can intuitively represent the structure and connection relationship of the network, which facilitates subsequent anomaly detection.

[0045] 3. Convert the risk identification problem into a graph anomaly detection problem

[0046] The present invention defines optical cables with security risks as abnormal optical cables, thus transforming the security risk identification problem of power communication networks into a graph anomaly detection problem. This transformation not only simplifies the complexity of the problem, but also enables risk identification using existing graph anomaly detection technology.

[0047] 4. Construction of risk identification model based on multi-dimensional factors

[0048] Based on an in-depth analysis of network topology and key attributes, this paper constructs a model for identifying security risks in power communication networks. This model comprehensively considers key attributes such as bandwidth capacity, transmission delay, and overload of optical cables, and can more accurately assess the security risks of power communication networks.

[0049] 5. Application of data-driven methods in security risk identification

[0050] The present invention adopts a data-driven approach to build and train a risk identification model by collecting and analyzing multi-source data in the power communication network. This approach can make full use of the information in the data and improve the accuracy and efficiency of risk identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a framework flow chart of the method of the present invention;

[0052] Figure 2 This is a schematic diagram of the structure of the autoencoder. DETAILED DESCRIPTION

[0053] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0054] like Figure 1 As shown, the present invention provides a method for identifying security risks of a power communication network based on a graph neural network, comprising the following steps:

[0055] S1: Preprocessing the collected raw data of the power communication network to obtain the optical cable feature matrix and adjacency matrix;

[0056] By calculating and selecting relevant features such as overload degree, cable scheduling level, port occupancy, bandwidth utilization, number of optical fibers, etc., they are standardized and the influence of dimension magnitude is removed to obtain the characteristic matrix of the optical cable; by matching the sites connected to the A end and the Z end of each optical cable, the adjacency matrix of the optical cable is obtained.

[0057] S2: Input the cable feature matrix and adjacency matrix into the encoder, learn the embedded representation of the node through the structure encoder, and use the attribute encoder to obtain the embedded representation of the attribute;

[0058] The embedding representation of nodes learned by the structure encoder includes:

[0059] In order to obtain a representative high-level node feature representation, the structure encoder transforms the observed raw node attributes X into a vector representation Z in a low-dimensional latent space. V , expressed as:

[0060] Z V =σ(XW V +b V ) (1)

[0061] Where σ(·) is the activation function ReLU, W V and b V are the weights and biases learned by the encoder.

[0062] Using the attribute encoder to obtain the embedded representation of the attribute includes:

[0063] The attribute encoder uses two nonlinear feature transformation layers to map the observed attribute data into a potential attribute embedding representation, expressed as:

[0064]

[0065]

[0066] Among them, W A(1) , bA(1) , W A(2) , b A(2) are the weights and biases learned in the two layers.

[0067] S3: Input the embedded representation of the node and the embedded representation of the attribute into the structure decoder and the attribute decoder respectively to reconstruct the original adjacency matrix and the feature matrix, and obtain the reconstruction error of the adjacency matrix and the reconstruction error of the feature matrix respectively;

[0068] A represents the adjacency matrix of the network, is the adjacency matrix reconstructed by the decoder, and the reconstruction error of the adjacency matrix is ​​expressed as:

[0069]

[0070] The reconstruction error is used to evaluate the abnormality of the network structure;

[0071]

[0072]

[0073] The attribute reconstruction decoder reconstructs the attribute X of the node, and the reconstruction error of the feature matrix is ​​expressed as:

[0074]

[0075] The reconstruction error is used to evaluate the anomaly of network properties;

[0076]

[0077] S4: Output anomaly scores through anomaly detection models according to the reconstruction errors of feature matrices and adjacency matrices;

[0078] The difference between the original data and the estimated data (reconstruction error) can reflect the anomalies in the data set to a certain extent. Specifically, data instances with large reconstruction errors are more likely to be considered anomalies because their patterns are significantly deviated from the majority and cannot be accurately reconstructed from the observed data;

[0079] like Figure 2 As shown, the anomaly detection model includes a deep autoencoder, which is a deep neural network that learns the potential representation of data in an unsupervised manner by stacking multiple layers of encoding and decoding functions. Given an input data set X, the deep autoencoder maps the data into a low-dimensional feature space through the encoding function Enc(), and then uses the decompression function Dec() to reconstruct the original data according to the representation of the feature space;

[0080] The anomaly detection model is trained to minimize the reconstruction loss:

[0081] min{E[dist(X,Dec(Enc(X)))]} (9)

[0082] In order to unify the reconstruction error, the loss function is as follows:

[0083]

[0084] The final anomaly score is:

[0085]

[0086] S5: Based on the anomaly score, the security risk identification result of the power communication network is obtained.

[0087] In order to verify the effectiveness and effect of the method of the present invention, this embodiment is verified by specific experiments, as follows:

[0088] 1. Dataset

[0089] The experimental data source of this embodiment is the topological data of the power communication network of a certain province, which contains a real data set of the communication network topological structure and its edge attributes (such as bandwidth utilization, overload degree, port occupancy, etc.). In order to ensure the data quality, comprehensive data preprocessing was carried out, including removing noise data to reduce errors, using appropriate methods to handle missing values ​​to maintain data integrity, and standardizing features to improve the accuracy and efficiency of subsequent analysis.

[0090] The preprocessed data set includes 2333 optical cables, each of which has 10 attributes, and there are 163 abnormal optical cables. In the data set preparation stage, this embodiment divides the sorted data into a training set, a test set, and a validation set in a ratio of 60%, 20%, and 20%.

[0091] 2. Experimental Setup

[0092] The hardware device used in the experiment is a computer with an Intel Core i7 processor and a GTX3080 graphics card. In terms of software environment, this embodiment uses the Python3 environment and the PyTorch library to conduct experiments on graph neural networks. The initialization parameters of the experiment are shown in Table 1.

[0093] Table 1 Model parameters

[0094]

[0095] 3. Method Comparison Experimental Results

[0096] The evaluation indicators are as follows:

[0097] 1. ROC-AUC

[0098] ROC-AUC is an important binary classification model evaluation indicator, which quantitatively evaluates the performance of the model by drawing the ROC curve and calculating its area under it.

[0099] 2. Recall@K

[0100] For the output anomaly sorted list, consider the first K edges as possible anomaly edges. The calculation method is as follows:

[0101] Recall = (the number of actual anomalies among the first K predicted anomalies) / (the total number of anomalies)

[0102] Table 2 AUC comparison of different algorithms

[0103] method AUC Dominant 0.7692 GCNAE 0.7713 Method of the present invention 0.8935

[0104] Table 3 Comparison of Recall@K of different algorithms

[0105] method K=300 K=500 Dominant 0.1963 0.4294 GCNAE 0.1595 0.4785 Method of the present invention 0.7423 0.8344

[0106] According to the experimental results shown in Tables 2 and 3, it can be seen that the graph anomaly detection method based on the autoencoder algorithm adopted by the present invention shows advantages in both AUC and Recall@K evaluation indicators. Specifically, compared with other comparison methods (Dominant, GCNAE), the AUC value and Recall@K of the method of the present invention are higher than those of other methods, indicating that the overall performance of the method of the present invention in identifying security risks of power communication networks is better than that of other methods. In summary, the method of the present invention has a relatively large advantage in identifying security risks of power communication networks, and can relatively accurately identify security risks in the network, providing strong technical support for the safe operation and maintenance of power communication networks.

Claims

1. A method for identifying security risks in power communication networks based on graph neural networks, characterized in that: The steps include: S1: Preprocessing the collected raw data of the power communication network to obtain the optical cable feature matrix and adjacency matrix; S2: Input the cable feature matrix and adjacency matrix into the encoder, learn the embedded representation of the node through the structure encoder, and use the attribute encoder to obtain the embedded representation of the attribute; S3: Input the embedded representation of the node and the embedded representation of the attribute into the structure decoder and the attribute decoder respectively to reconstruct the original adjacency matrix and the feature matrix, and obtain the reconstruction error of the adjacency matrix and the reconstruction error of the feature matrix respectively; S4: Output anomaly scores through anomaly detection models according to the reconstruction errors of feature matrices and adjacency matrices; S5: Based on the anomaly score, the security risk identification result of the power communication network is obtained.

2. According to claim 1, a method for identifying security risks of power communication network based on graph neural network is characterized in that: In the step S1, the characteristic matrix of the optical cable is obtained by calculating and selecting relevant features, normalizing them, and removing the influence of the dimension order; and the adjacency matrix of the optical cable is obtained by matching the sites connected to the A end and the Z end of each optical cable.

3. According to claim 1, a method for identifying security risks of power communication network based on graph neural network is characterized in that: The step S2 of learning the embedding representation of the node by the structure encoder includes: The structure encoder transforms the observed raw node attributes X into a vector representation Z in a low-dimensional latent space. V , expressed as: Z V =σ(XW V +b V ) (1) Where σ(·) is the activation function ReLU, W V and b V are the weights and biases learned by the encoder.

4. According to claim 1, a method for identifying security risks of power communication network based on graph neural network is characterized in that: The step S2 of obtaining the embedded representation of the attribute by using the attribute encoder includes: The attribute encoder uses two nonlinear feature transformation layers to map the observed attribute data into a potential attribute embedding representation, expressed as: Among them, W A(1) , b A(1) , W A(2) , b A(2) are the weights and biases learned in the two layers.

5. According to claim 1, a method for identifying security risks of power communication network based on graph neural network is characterized in that: In step S3, A represents the adjacency matrix of the network. is the adjacency matrix reconstructed by the decoder, and the reconstruction error of the adjacency matrix is ​​expressed as: The reconstruction error is used to evaluate the abnormality of the network structure; 6. The method for identifying security risks of power communication network based on graph neural network according to claim 1 is characterized in that: In step S3, the attribute reconstruction decoder reconstructs the attribute X of the node, and the reconstruction error of the feature matrix is ​​expressed as: The reconstruction error is used to evaluate the anomaly of network properties; 7. The method for identifying security risks of power communication network based on graph neural network according to claim 1 is characterized in that: The anomaly detection model in step S4 includes a deep autoencoder. Given an input data set X, the deep autoencoder maps the data into a low-dimensional feature space through an encoding function Enc(), and then uses a decompression function Dec() to reconstruct the original data according to the representation of the feature space.

8. The method for identifying security risks of power communication network based on graph neural network according to claim 7 is characterized in that: In step S4, the anomaly detection model is trained to minimize the reconstruction loss: min{E[dist(X,Dec(Enc(X)))]} (9) In order to unify the reconstruction error, the loss function is as follows: The final anomaly score is:

Citation Information

Cited By

  • UUV power system fault prediction method, device and equipment and storage medium

    CN120316620A

  • Identification system and method for intrinsic semantic difference learning

    CN121476550A