Distribution network data asset vulnerability identification method, device, system, and storage medium
By constructing a graph neural network model and combining it with a vulnerability fusion strategy, we can identify various vulnerabilities of distribution network data assets, solving the problem of the failure of existing technologies to effectively identify data asset vulnerabilities and achieving efficient and accurate vulnerability identification and risk analysis.
Patent Information
- Application Number
- CN202411812525.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-12-10
AI Technical Summary
Existing technologies fail to effectively identify and address the vulnerabilities of distribution network data assets, resulting in data leaks and attacks posing potential threats to the safe and stable operation of the power grid.
By building a graph neural network (GNN) model and combining it with vulnerability fusion strategy, we can identify the vulnerabilities of distribution network data assets, including data leakage, service denial, data integrity, data availability, and data dependency. We can also use multi-dimensional feature fusion and machine learning technology to improve recognition accuracy and efficiency.
It improves the accuracy and efficiency of identifying the vulnerability of distribution network data assets, provides strong support for the stable operation of the power system, and enhances the ability to analyze potential risks.
Smart Images

Figure CN119719670B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of distribution network data asset security technology, and in particular relates to a distribution network data asset vulnerability identification method and device, system, and storage medium. Background Art
[0002] With the development of smart grid technology, the value of data generated by smart meters, sensors, and control devices in distribution networks is becoming increasingly prominent due to its critical role in grid operation monitoring, energy optimization, and efficiency improvement. This data not only records key information such as the real-time status of the grid, user electricity usage behavior, and device operation, but also plays an important role in grid management and decision-making. Therefore, it is considered a data asset of the power system. However, the increasing value of these data assets also makes them a target for potential attackers, especially during data circulation and processing. Once attacked or leaked, it can have a serious impact on the safe and stable operation of the grid. Therefore, there is an urgent need for a vulnerability identification method for distribution network data assets, which is of great significance for maintaining grid security and promoting the sustainable development of smart grids. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a method and device, system and storage medium for identifying the vulnerability of distribution network data assets.
[0004] To achieve the above object, the present invention adopts the following technical solutions:
[0005] A method for identifying the vulnerability of distribution network data assets, comprising:
[0006] Step S1: Based on the raw data in the distribution network, a data asset set is obtained, and the data asset set is divided into a training set and a test set. The raw data includes: user electricity usage information of smart meters, grid status monitoring data of sensors, and grid control instructions of control devices.
[0007] Step S2: construct a graph based on the data asset training set, the relationships between data assets, and the feature matrix;
[0008] Step S3: The extracted vulnerability features are integrated into the graph through the graph neural network and vulnerability fusion strategy to obtain a GNN model containing vulnerability fusion features;
[0009] Step S4: training and optimizing the GNN model of vulnerability fusion features;
[0010] Step S5: Input the test set into the trained GNN model containing vulnerability fusion features to identify the vulnerability of distribution network data assets.
[0011] Preferably, step S1 includes:
[0012] According to the original data in the distribution network, a data asset set is obtained;
[0013] Perform data cleaning on the data asset set;
[0014] Divide the cleaned data asset set into a training set and a test set;
[0015] Extract key information from the training set to obtain the basic features of each type of data asset; wherein, the basic features together constitute the feature matrix of the data asset.
[0016] Preferably, the vulnerability fusion strategy is to define a feature fusion strategy for distribution network data assets for vulnerability types such as data leakage, denial of service, data integrity, data availability and data dependency.
[0017] The present invention also provides a distribution network data asset vulnerability identification device, comprising:
[0018] The preprocessing module is used to obtain a data asset set based on the raw data in the distribution network, and divide the data asset set into a training set and a test set. The raw data includes: user electricity usage information from smart meters, grid status monitoring data from sensors, and grid control instructions from control devices;
[0019] A construction module for constructing a graph based on a training set of data assets, the relationships between data assets, and a feature matrix;
[0020] The fusion module is used to fuse the extracted vulnerability features into the graph through the graph neural network and vulnerability fusion strategy to obtain a GNN model containing vulnerability fusion features;
[0021] Training module, used to train and optimize the GNN model of vulnerability fusion features;
[0022] The identification module is used to input the test set into the trained GNN model containing vulnerability fusion features to identify the vulnerability of distribution network data assets.
[0023] Preferably, the preprocessing module includes:
[0024] A processing unit, configured to obtain a data asset set based on raw data in the distribution network;
[0025] A cleaning unit, used to perform data cleaning on a data asset set;
[0026] The partitioning module is used to divide the data asset set after data cleaning into a training set and a test set;
[0027] The extraction unit is used to extract key information from the training set to obtain the basic features of each type of data asset; wherein the basic features together constitute the feature matrix of the data asset.
[0028] Preferably, the vulnerability fusion strategy is to define a feature fusion strategy for distribution network data assets for vulnerability types such as data leakage, denial of service, data integrity, data availability and data dependency.
[0029] An embodiment of the present invention also provides a distribution network data asset vulnerability identification system, comprising: a memory and a processor, wherein the memory stores a computer program run by the processor, and the computer program executes a distribution network data asset vulnerability identification method when run by the processor.
[0030] An embodiment of the present invention further provides a storage medium having a computer program stored thereon, and the computer program executes the method for identifying the vulnerability of distribution network data assets when running.
[0031] The present invention first collects raw data from the distribution network, cleans it to form a data asset set, and then extracts basic features to form a feature matrix. Next, a graph is constructed based on the data asset training set, the relationships between data assets, and the feature matrix. The extracted vulnerability features are then integrated into the graph using a graph neural network and a vulnerability fusion strategy. This invention considers the multidimensional characteristics of distribution network data assets and, through graph neural networks and machine learning techniques, improves the accuracy and efficiency of vulnerability identification, providing strong support for the stable operation of the power system. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0033] Figure 1 This is a flow chart of a method for identifying the vulnerability of distribution network data assets according to an embodiment of the present invention. DETAILED DESCRIPTION
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0035] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0036] Example 1:
[0037] like Figure 1 As shown, an embodiment of the present invention provides a method for identifying the vulnerability of distribution network data assets, including:
[0038] Step S1: Based on the original data in the distribution network, a data asset set is obtained, and the data asset set is divided into a training set and a test set.
[0039] The raw data collected from the distribution network includes user electricity consumption information from smart meters, grid status monitoring data from various sensors (such as temperature and current sensor readings), and grid control instructions from control devices (such as circuit breaker and relay status information). After preliminary classification and organization of these raw data, a data asset set D = {d1, d2, ..., d N}, where d i Represents the i-th type of data assets.
[0040] These data asset sets are cleaned to ensure data quality and accuracy. The cleaning process primarily involves handling missing values (through interpolation or deletion of missing data points) and detecting outliers (identifying and addressing abnormalities such as sudden jumps in sensor readings). This process ensures the reliability and validity of the data assets, providing a solid data foundation for subsequent vulnerability analysis and model building. This processing not only improves the usability of the data assets but also provides the necessary prerequisites for identifying vulnerabilities in distribution network data assets. The cleaned data asset set is divided into a training set and a test set.
[0041] Extract key information from the training set to form a comprehensive basic feature profile of each type of data asset. Basic features include: technical specifications, such as the model of the smart meter and the measurement range of the sensor; operating status, covering the online / offline status of the device and historical fault records; security configuration, including access control levels and encryption measures; network traffic, involving the size and direction of data traffic; and physical and environmental factors, such as the geographical location of the device and environmental conditions such as temperature and humidity. Basic features together constitute the feature matrix of the data asset. (N is the number of data assets, d is the characteristic dimension of each type of data asset):
[0042]
[0043] Among them, x ijis the value of the jth feature of the i-th data asset, i ranges from 1 to N, and j ranges from 1 to d.
[0044] Step S2: Construct a graph based on the data asset training set, the relationships between data assets, and the feature matrix
[0045] 1) Node definition: Each type of data asset is a node in the graph and contains its basic feature vector. N} represents a set of graph nodes, where v i Represents the i-th type of data asset node.
[0046] 2) Edge definition: This is to determine the interaction relationships between data assets, which are used as edges in the graph. Interaction relationships may include data flows (such as sensor data sent to the control center), control signals (such as circuit breakers receiving signals from the control center), physical connections (such as devices directly connected by cables), and logical connections (such as communication between devices achieved through network protocols). Let E∈V×V be the set of edges, then it describes the connection relationship between any two nodes in the node set V. In other words, node v i and v j If there is an edge between them, then (v i ,v j )∈E.
[0047] 3) Graph attributeization: This is the process of combining the feature matrix X with the graph structure to form a complete graph representation G. Specifically, the graph G consists of three main parts: the node set V, the edge set E, and the node feature matrix X. At this point, the graph can be represented as:
[0048] G=(V,E,X) (2)
[0049] In this process, not only are nodes and edges defined, but additional attribute information is also introduced, which directly enriches the structure and semantics of the graph G. For example, the network topology, as an attribute of the graph, describes the physical or logical connection layout between devices in the distribution network, which is reflected by the edge set E. At the same time, the dependencies between devices indicate that one device may rely on the data or services of another device to perform its function. This dependency can be represented by the features in the node feature matrix X or further refined by the attributes of the edges. This attribute information not only greatly enriches the context of the graph G but also enables graph neural networks (GNNs) to more accurately capture and analyze the interaction patterns and potential vulnerabilities between data assets. In this way, the graph G is not just a structured graph, but a complex system containing rich semantic information, providing a solid foundation for vulnerability analysis of distribution network data assets.
[0050] Step S3: The extracted vulnerability features are integrated into the graph through the graph neural network and vulnerability fusion strategy to obtain a GNN model containing vulnerability fusion features.
[0051] Fragility feature fusion integrates different types of vulnerability features into the feature representation of graph nodes. These features include not only the properties of the node itself, but also information about the edges associated with the node. The following is the process of integrating these vulnerability features into the graph G using the GNN model:
[0052] First, in the GNN model, for the lth layer (l=1,2,…,L), the node v i The feature update can be expressed as:
[0053]
[0054] in, is node v i In the feature representation of layer l, is node v i The set of neighbor nodes, W (l) is the learnable weight matrix of layer l, and τ is a nonlinear activation function such as ReLU or Sigmoid function.
[0055] Secondly, a set of feature fusion strategies are defined for distribution network data assets, targeting vulnerability types such as data leakage, denial of service, data integrity, data availability, and data dependency. These strategies effectively integrate features related to various vulnerabilities by introducing weighting factors and customized feature functions. This not only enhances the model's ability to identify different vulnerability characteristics but also improves the accuracy of its analysis of potential risks in the distribution network.
[0056] ① Data leakage vulnerability: Emphasize the data encryption characteristics and access control of the node. First, define a function to quantify the node v i The data encryption strength. This function can be calculated based on the encryption algorithm, key length and encryption protocol used:
[0057] encrypt(x i )=log2(key_length(x i ))+log2(algorithm_strength(x i )) (4)
[0058] Here, key_length(x i ) is the node v i The length of the encryption key used, algorithm_strength(x i ) is the node v iThe strength of the encryption algorithm used. This function reflects the encryption security level of the node. Secondly, the access control characteristics can be defined by quantifying the access rights and control policies of the node, such as defining a function to evaluate the access control strength of the node:
[0059]
[0060] Among them, U is the set of all users, is the set of nodes that user u is authorized to access. This function reflects the strictness of node access control. Finally, the data encryption and access control features are fused with the original node features to enhance the GNN model's ability to identify data leakage vulnerabilities. The fused features are represented as:
[0061]
[0062] x i is node v i The original eigenvectors, β and They are weight factors used to adjust the importance of encryption features and access control features, encrypt(x i ) and access_control(x i ) represent the nodes v i Data encryption and access control features.
[0063] ② Denial of Service vulnerability: Focus on the node's traffic patterns and traffic anomaly characteristics. On the one hand, the long short-term memory network (LSTM) is used to extract features from the node's traffic data patterns to capture the long-term dependencies and periodic changes in traffic. The LSTM update formula is as follows:
[0064] h t =LSTM(x t ,h t-1 ) (7)
[0065] Among them, h t is the hidden state at time step t, x t is the input feature at time step t, h t-1 is the hidden state at time step t-1. On the other hand, the attention mechanism is introduced to enhance the feature extraction capability of LSTM, thereby completing the identification of abnormal traffic behavior. The output vector z after attention weighting is t It can be expressed as:
[0066] z t =Attention(H) (8)
[0067] H is the set of hidden states at all time steps. Finally, the features extracted by LSTM and attention mechanism are fused with the original node features to enhance the GNN model’s ability to identify denial of service vulnerabilities. The fused features are expressed as:
[0068]
[0069] Among them, x i is node v i The original feature vector of traffic(x i )=z t .
[0070] ③ Data integrity vulnerability: Integrate monitoring and anomaly detection features of data processing activities. This paper uses time series analysis methods to monitor data processing activities. For example, an exponential smoothing algorithm is used to generate dynamic upper and lower bounds to detect data anomalies:
[0071] Upper Bound=μ+M·σ (10)
[0072] Lower Bound=μ-M·σ (11)
[0073] Where μ is the mean, σ is the standard deviation, and M is a constant used to determine the width of the upper and lower bounds. i If the data value of a feature exceeds these limits, it indicates that the integrity of the data asset may be threatened. In addition, an anomaly detection method based on GLR (Generalized Likelihood Ratio) is used to find anomalies. This method identifies anomalies by comparing predicted values with actual values:
[0074]
[0075] Among them, p t (y t |x 1:t-1 ) is given historical data x 1:t-1 The predicted probability under t (y t |x 1:t ,θ) is given the complete data x 1:t and the predicted probability under the model parameters θ. Finally, these features are fused with the original node features to enhance the GNN model's ability to identify data integrity vulnerabilities. The fused features are expressed as:
[0076]
[0077] Among them, x i is node v iThe original feature vector of δ is the weight factor used to adjust the importance of integrity features; integrity(x i ) represents node v i The data asset integrity monitoring characteristics can be expressed as:
[0078]
[0079] It can be seen that integrity(x i ) is a binary feature, representing the node v i Whether the characteristic data is within the normal range.
[0080] ④ Data availability vulnerability: Analyze the redundancy and backup mechanism characteristics of the nodes. Redundancy refers to the degree of replication of key components or functions in the system to improve the reliability of the system. In this invention, the redundancy characteristic R(x i )for:
[0081] R(x i )=(2N r -N) / N (15)
[0082] Among them, N r is node v i The number of redundant components, N is the total number of components. This indicator reflects the redundancy of the node in the face of a single point of failure. The backup mechanism reflects the ability to quickly recover when data assets are lost, damaged or unavailable. Here we define the backup mechanism characteristic B(x i )for:
[0083]
[0084] Among them, RTO(x i ) is the measure of node v i is an indicator of data recovery speed (the closer its value is to 1, the stronger the recovery capability), and λ is a weight parameter that adjusts the impact of data recovery time on the backup mechanism score. In other words, this binary feature indicates the impact of the backup mechanism on data availability. By integrating the redundancy and backup mechanism features into the node's feature vector, it can be expressed as:
[0085] availability(x i )=R(x i )+B(x i ) (17)
[0086] Visible availability (x i ) is a comprehensive indicator that takes into account not only the redundancy of the node but also the node backup mechanism. Finally, this comprehensive availability feature is fused with the original feature vector to obtain the fused feature vector:
[0087]
[0088] Among them, x i is node v i is the original feature vector of , and ξ is the weight factor used to adjust the importance of availability features.
[0089] ⑤Data dependency vulnerability: Evaluate the dependency of a node on an external system or component. Dependency vulnerability can be evaluated by quantifying the degree of dependency of a node on an external system or component. First, define a dependency index D(x i ), which reflects the node v i Dependence on external resources:
[0090]
[0091] in, is node v i The set of neighbor nodes, w ij is node v i With v j The edge weights between them represent the connection strength or data interaction frequency between them, dependency(x j ) is the node v j The dependency feature vector of . Secondly, the dependency index is fused with the original node features to enhance the GNN model's ability to identify dependency vulnerability. The fused features are expressed as:
[0092]
[0093] Among them, x i is node v i The original eigenvector of is a weight factor used to adjust the importance of dependent features, D(x i ) represents node v i Dependency index on external systems or components. It should be noted that in order to further evaluate the impact of dependency on data availability, an impact metric I(x i ), which reflects the node v i The impact of dependencies on its normal operation:
[0094]
[0095] It is a tuning parameter that controls the degree to which dependency affects availability.
[0096] Finally, the fusion features Substituting into the update formula of the GNN model, we can get:
[0097]
[0098] The GNN can be used to fuse different types of vulnerability features into the graph G, improving the present invention's ability to identify vulnerabilities in distribution network data assets. The feature fusion strategy for each vulnerability type adjusts the importance of features using learnable weighting factors, enabling the present invention to adaptively focus on the most relevant characteristics for each vulnerability. This approach not only improves the accuracy of vulnerability identification but also provides direct guidance for subsequent security measures.
[0099] Step S4: Training and optimizing the GNN model of vulnerability fusion features
[0100] The purpose of training the GNN model for vulnerability fusion features is to adjust the model parameters to minimize the difference between the model prediction and the true vulnerability label. The following is a detailed model training and optimization process:
[0101] First, set vulnerability category labels for the training set of the distribution network data asset dataset D. is the label matrix, where C is the total number of categories; this also means that for each type of data asset node v i Each is marked with a specific vulnerability category label.
[0102] Secondly, the cross entropy loss function is used to measure the difference between the vulnerability class predicted by the model and the true label. For multi-class classification problems, the cross entropy loss function is defined as:
[0103]
[0104] Among them, y ic is category c for node v i The true label (usually one-hot encoded), is the node v predicted by the model i The probability of belonging to category c. This loss function encourages the model to improve the prediction probability of the correct category while reducing the prediction probability of the wrong category.
[0105] Again, in each iteration, the gradient of the loss function with respect to the model parameters is calculated, and the parameters are updated to reduce the loss by the following formula:
[0106]
[0107] Among them, θ is the model parameter, η is the learning rate, is the gradient of the loss function with respect to the parameters.
[0108] Finally, to prevent overfitting, a regularization term can be added to the loss function:
[0109] ψ regularized =ψ+λ∑ l ||w (l) || 2 (25)
[0110] Where λ is the regularization coefficient, ||W (l) || 2 is the Frobenius norm of the weight matrix of the first layer. In order to make the model converge faster, it is usually necessary to make the input features distributed in a similar range, that is, to perform normalization:
[0111]
[0112] Here, μ and σ are the mean and standard deviation of the feature, respectively.
[0113] During the training process, the present invention uses graph batches to process data from multiple graphs in parallel; at the same time, the model performance is monitored on the validation set, and training is stopped when the performance no longer improves to avoid overfitting; and an independent validation set is used to adjust the model parameters and hyperparameters.
[0114] Through the above process, the GNN model can be effectively trained and optimized for the task of identifying asset vulnerabilities in distribution network data. This process involves not only minimizing the loss function but also optimizing model parameters, regularization, feature normalization, and adjusting the training strategy to ensure the model's generalization ability and prediction accuracy.
[0115] Step S5: Input the test set into the trained GNN model containing vulnerability fusion features to identify the vulnerability of distribution network data assets.
[0116] This invention constructs a comprehensive distribution network data asset feature matrix, integrating features from multiple dimensions, including technical specifications, operating status, security configuration, network traffic, and physical and environmental factors. This multi-dimensional feature fusion technology provides rich contextual information to the GNN model, significantly improving the model's ability to capture and analyze interactions between data assets and potential vulnerabilities. When constructing the graph, this invention attributes the graph by defining nodes and edges and adding attribute information such as network topology and inter-device dependencies. This not only enhances the GNN model's understanding of the complex interactions between devices in the distribution network, but also provides an effective multi-source heterogeneous data processing and fusion solution, addressing the technical challenge of extracting key information from and effectively fusing multi-source heterogeneous data. For each vulnerability type of distribution network data asset, this invention defines a feature fusion strategy for each, and adjusts the importance of features using learnable weighting factors, enabling the model to adaptively focus on the most relevant characteristics for each vulnerability. This customized feature fusion strategy directly addresses how to use graph neural networks to handle the complex interactions between devices in the distribution network, and how to enhance the model's understanding of inter-device dependencies and network topology through graph attribution. This paper employs multiple strategies during model training and optimization, including cross-entropy loss, regularization, feature normalization, and graph batch parallel processing, to ensure the model's generalization and predictive accuracy. By monitoring model performance on a validation set and adjusting model parameters and hyperparameters, and using a test set to evaluate the model's final performance, this paper effectively avoids overfitting and optimizes the model's performance in the task of identifying distribution network data asset vulnerabilities.
[0117] Example 2:
[0118] An embodiment of the present invention further provides a device for identifying the vulnerability of distribution network data assets, comprising:
[0119] The preprocessing module is used to obtain a data asset set based on the raw data in the distribution network, and divide the data asset set into a training set and a test set. The raw data includes: user electricity usage information from smart meters, grid status monitoring data from sensors, and grid control instructions from control devices;
[0120] A construction module for constructing a graph based on a training set of data assets, the relationships between data assets, and a feature matrix;
[0121] The fusion module is used to fuse the extracted vulnerability features into the graph through the graph neural network and vulnerability fusion strategy to obtain a GNN model containing vulnerability fusion features;
[0122] Training module, used to train and optimize the GNN model of vulnerability fusion features;
[0123] The identification module is used to input the test set into the trained GNN model containing vulnerability fusion features to identify the vulnerability of distribution network data assets.
[0124] As an implementation method of an embodiment of the present invention, the preprocessing module includes:
[0125] A processing unit, configured to obtain a data asset set based on raw data in the distribution network;
[0126] A cleaning unit, used to perform data cleaning on a data asset set;
[0127] The partitioning module is used to divide the data asset set after data cleaning into a training set and a test set;
[0128] The extraction unit is used to extract key information from the training set to obtain the basic features of each type of data asset; wherein the basic features together constitute the feature matrix of the data asset.
[0129] As an implementation method of an embodiment of the present invention, the vulnerability fusion strategy is: defining a feature fusion strategy for distribution network data assets for vulnerability types such as data leakage, denial of service, data integrity, data availability, and data dependency.
[0130] Example 3:
[0131] An embodiment of the present invention also provides a distribution network data asset vulnerability identification system, comprising: a memory and a processor, wherein the memory stores a computer program run by the processor, and the computer program executes a distribution network data asset vulnerability identification method when run by the processor.
[0132] Example 4:
[0133] An embodiment of the present invention further provides a storage medium having a computer program stored thereon, and the computer program executes the method for identifying the vulnerability of distribution network data assets when running.
[0134] The embodiments described above are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by persons skilled in the art should fall within the scope of protection defined by the claims of the present invention.
Claims
1. A method for identifying the vulnerability of distribution network data assets, characterized in that: include: Step S1: Based on the raw data in the distribution network, a data asset set is obtained, and the data asset set is divided into a training set and a test set. The raw data includes: user electricity usage information of smart meters, grid status monitoring data of sensors, and grid control instructions of control devices. Step S2: construct a graph based on the data asset training set, the relationships between data assets, and the feature matrix; Step S3: The extracted vulnerability features are integrated into the graph through the graph neural network and vulnerability fusion strategy to obtain a GNN model containing vulnerability fusion features; Step S4: training and optimizing the GNN model of vulnerability fusion features; Step S5: Input the test set into the trained GNN model containing vulnerability fusion features to identify the vulnerability of distribution network data assets; Step S1 includes: According to the original data in the distribution network, a data asset set is obtained; Perform data cleaning on the data asset set; Divide the cleaned data asset set into a training set and a test set; Extract key information from the training set to obtain basic features of each type of data asset; wherein the basic features together constitute a feature matrix of the data asset; The vulnerability fusion strategy is to define a feature fusion strategy for distribution network data assets based on vulnerability types such as data leakage, denial of service, data integrity, data availability and data dependency.
2. A device for identifying the vulnerability of distribution network data assets, characterized in that: include: The preprocessing module is used to obtain a data asset set based on the raw data in the distribution network, and divide the data asset set into a training set and a test set. The raw data includes: user electricity usage information from smart meters, grid status monitoring data from sensors, and grid control instructions from control devices; A construction module for constructing a graph based on a training set of data assets, the relationships between data assets, and a feature matrix; The fusion module is used to fuse the extracted vulnerability features into the graph through the graph neural network and vulnerability fusion strategy to obtain a GNN model containing vulnerability fusion features; Training module, used to train and optimize the GNN model of vulnerability fusion features; The identification module is used to input the test set into the trained GNN model containing vulnerability fusion features to identify the vulnerability of distribution network data assets; The preprocessing modules include: A processing unit, configured to obtain a data asset set based on raw data in the distribution network; A cleaning unit, used to perform data cleaning on a data asset set; The partitioning module is used to divide the data asset set after data cleaning into a training set and a test set; An extraction unit, configured to extract key information from the training set to obtain basic features of each type of data asset; wherein the basic features collectively constitute a feature matrix of the data asset; The vulnerability fusion strategy is to define a feature fusion strategy for distribution network data assets based on vulnerability types such as data leakage, denial of service, data integrity, data availability and data dependency.
3. A distribution network data asset vulnerability identification system, characterized in that: include: A memory and a processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the method for identifying the vulnerability of distribution network data assets according to claim 1 is executed.
4. A storage medium, characterized in that The storage medium stores a computer program, which executes the distribution network data asset vulnerability identification method according to claim 1 when running.
Citation Information
Patent Citations
Network node vulnerability assessment method and system based on graph attention network
CN118473960A