A Data Alignment Method for Power Business Systems Based on Graph Convolutional Neural Networks

Through graph convolution neural network aligning the power business system data, the problem of inconsistent equipment information in different business departments is solved, the reliability and consistency of data is achieved, and the information construction and resource sharing effect of power enterprises is improved.

CN115935941BActive Publication Date: 2025-07-04STATE GRID FUJIAN ELECTRIC POWER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211606845.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-14
Publication Date
2025-07-04
Estimated Expiration
2042-12-14

AI Technical Summary

Technical Problem

In the power business system, different business departments have inconsistent recording data on equipment information, resulting in a decrease in data reliability and consistency. As time goes by, the data complexity increases, affecting information construction and resource sharing.

Method used

Using a graph convolutional neural network method, the network entity alignment model is trained to associate equivalent entities of different business systems, and using graph self-attention convolutional neural network and supervised learning and perturbation perspective comparison learning, calculate node similarity and align nodes.

Benefits of technology

It improves the reliability and consistency of different data sources, realizes accurate correspondence of data pointing to the same object in multiple power business systems, and improves the effectiveness of data management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115935941B_ABST
    Figure CN115935941B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for aligning data of a power service system based on a graph convolutional neural network, including: cleaning and preprocessing device ledger information data; constructing a knowledge graph according to entities and the relationships between entities, and obtaining pre-aligned entity pairs between knowledge graphs; inputting the knowledge graph into a graph self-attention convolutional neural network for training, using the entity pairs as alignment seeds and as the supervision information of the graph self-attention convolutional neural network; obtaining the embedding vector representations of each node through the graph self-attention convolutional neural network, calculating the similarity between each node in the knowledge graphs, and taking the two most similar nodes as alignment nodes; rewriting the attribute data of the entities to be aligned in the power service system according to the alignment nodes. The present invention associates equivalent entities of different business systems through a trained network entity alignment model, predicts the correspondence of data pointing to the same object in the real world in multiple power service systems, and ensures the reliability of data from different data sources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for aligning data of a power business system based on a graph convolutional neural network, belonging to the technical field of power business data processing. Background Art

[0002] The deepening reform of the power industry requires power enterprises to further improve informatization construction, and further achieve information interconnection and resource sharing to a greater extent, so that power enterprises can effectively manage data resources, fully tap the value of data resources, achieve the goal of cost reduction and efficiency increase, and thus further expand the industry space.

[0003] However, in the process of using the power business system, there will be inconsistent recorded data of equipment information in different business departments. Moreover, as time goes by, the power network data becomes increasingly complex, and the data is frequently copied and modified when spreading on different power business systems, resulting in a decrease in the reliability of the data in different business systems. At this time, it is necessary to modify the information representing the same equipment in different business systems to ensure the reliability and consistency of the data. Summary of the Invention

[0004] To overcome the above problems, the present invention provides a method for aligning data of a power business system based on a graph convolutional neural network. This method associates equivalent entities in different business systems through a trained network entity alignment model, predicts the corresponding situation of data pointing to the same object in the real world in multiple power business systems, and ensures the reliability of data from different data sources.

[0005] The technical solution of the present invention is as follows:

[0006] A method for aligning data of a power business system based on a graph convolutional neural network, comprising:

[0007] Obtain the equipment ledger information data of two power business systems to be aligned, and clean and preprocess the equipment ledger information data;

[0008] Respectively obtain the entities in each power business system and the relationships between the entities, construct a knowledge graph according to the entities and the relationships between the entities, and obtain the pre-aligned entity pairs between the two knowledge graphs, where the entity is the node of the knowledge graph, and the relationship between the entities is the edge of the knowledge graph;

[0009] Input the knowledge graph into a graph self-attention convolutional neural network for training, and use the entity pairs as alignment seeds, which serve as supervision information during the training of the graph self-attention convolutional neural network;

[0010] Obtain the embedded vector representations of each node through the graph self-attention convolutional neural network, calculate the similarity between nodes in the knowledge graphs, and take the two most similar nodes as aligned nodes;

[0011] Rewrite the attribute data of the power business system entities to be aligned according to the aligned nodes.

[0012] Furthermore, clean and preprocess the equipment inventory information data, specifically by removing the damaged data in the equipment inventory information data and processing the remaining data into CSV format.

[0013] Furthermore, construct a knowledge graph based on the entities and the relationships between the entities, and obtain the pre-aligned entity pairs between the two knowledge graphs. Specifically:

[0014] Extract the nodes of the knowledge graph. Specifically, use the attributes for distinguishing entities in the equipment inventory information data as the nodes of the knowledge graph, and use the text information fields of the entities as the text features of the nodes;

[0015] Extract the semantic information of the text features through a language model to obtain the embedded vector representations of the nodes;

[0016] Construct the edges of the nodes according to the relationships between the entities;

[0017] Obtain two knowledge graphs G to be aligned k :

[0018] G k ={E k , R k , T k};

[0019] where k = 1 or 2, E k , R k and T k are respectively the set of nodes, the set of entity relationships, and the triple <e1, r, e2> in the knowledge graph, e1, e2 ∈ E, r ∈ R, and r is the relationship between entity e1 and entity e2;

[0020] Pre-align the entities according to the attributes with unique values of the entities to obtain entity pairs. The entity pair set S is:

[0021]

[0022] where x and y are the entity nodes of knowledge graphs G1 and G2 respectively.

[0023] Furthermore, the language model is the LaBSE model.

[0024] Further, input the knowledge graph into the graph self-attention convolutional neural network for training, and use the entity pair as the alignment seed, which serves as the supervision information during the training of the graph self-attention convolutional neural network. Specifically:

[0025] S1. Divide the entity pair set S into a training set and a validation set at a preset ratio, and the test set is the unaligned nodes.

[0026] Use a single-layer graph attention convolutional neural network to perform an aggregation operation on the neighbor information of each node in the knowledge graph. Specifically:

[0027] S2. Randomly sample 20 neighbor nodes of the target node as the neighbor set N i , where the neighbor nodes include the target node itself. Update the embedding vector representation of the target node through the embedding vector representations of each neighbor node in the neighbor set N i . The aggregation formula is as follows: where W is a trainable weight matrix that maps the embedding vector representation of the node to high-level features, and σ is the non-linear activation function Sigmoid. The calculation method is as follows:

[0028]

[0029]

[0030] where a

[0031]

[0032] is the importance of node j to node i, and the calculation method is as follows: ij

[0033]

[0034] where a is a linear layer that converts the vector into a numerical value, and LeakyReLU is a non-linear activation function, which is expressed as follows:

[0035]

[0036] where p is a coefficient;

[0037] S3. Based on the entity pair, train the graph self-attention convolutional neural network, update the network parameters according to gradient descent, and use Bayesian personalized ranking as the objective function for supervised learning. The expression is as follows:

[0038]

[0039] where (x, y, y -) constructs a training triplet for node x of the knowledge graph, y is a node pre-aligned with node x in another knowledge graph, y - is any node randomly sampled except x and y;

[0040] S4. Unsupervised training based on graph structure multi-view enhancement method, specifically:

[0041] By perturbing the encoder network parameter θ to obtain the perturbation network parameter θ′, the same knowledge graph node is input into the network and the perturbation network to obtain two view representations h and h′ of the node, which are expressed as follows:

[0042] h=f(N; θ), h′=f(N; θ′);

[0043] The way to perturb the encoder is:

[0044] θ′ l =θ l +η·Δθ l ;

[0045]

[0046] Among them, θ l and θ′ l are the parameters of the l-th layer graph self-attention convolutional neural network and the l-th layer perturbation graph self-attention convolutional neural network, η is an adjustable perturbation intensity hyperparameter, Δθ l has a mean of zero and a variance of The disturbance term of the Gaussian distribution;

[0047] InfoNCE is used as the target optimization function to bring the original representation and perturbation representation of the same node closer, and push the perturbation representation of other nodes further away, as shown below:

[0048]

[0049] Where N is the number of nodes in a training batch, sim is the cosine similarity, τ is an adjustable parameter;

[0050] S5. Repeat steps S2 to S5 until the value of the objective function converges or a preset number of training times is reached.

[0051] Furthermore, the entity pair set S is divided into a training set and a validation set in a ratio of 4:1.

[0052] Furthermore, the unaligned nodes of the knowledge graph with fewer nodes in the two knowledge graphs are used as a test set.

[0053] Further, the embedded vector representations of the nodes are obtained through the graph self-attention convolutional neural network, and the similarities between the nodes in the knowledge graphs are calculated. The two most similar nodes are used as the aligned nodes, specifically:

[0054] The embedded vector representations of the nodes in the knowledge graphs are calculated through the graph self-attention convolutional neural network;

[0055] The similarities between the representations of the nodes in the two knowledge graphs are calculated according to the Euclidean norm:

[0056]

[0057] The two most similar nodes are taken as the aligned nodes.

[0058] Further, it also includes performing performance statistics on the labeled training set and validation set nodes, and manually checking the unlabeled test set nodes to determine whether they are valid, specifically:

[0059] Calculate Hits@1 and Hits@10 for the labeled training set and validation set nodes:

[0060]

[0061] where S is the set of aligned node pairs, |S| is the number of aligned node pairs, rank i is the link prediction rank of the i-th aligned node pair, I is the judgment function, if it is true, then I = 1, otherwise, I = 0;

[0062] Manually check the unlabeled test set nodes to determine whether they are valid.

[0063] The present invention has the following beneficial effects:

[0064] Through the trained network entity alignment model, the present invention associates equivalent entities in different business systems, predicts the corresponding situations of data pointing to the same object in the real world in multiple power business systems, and ensures the reliability of data from different data sources. The present invention combines the graph self-attention convolutional neural network, supervised learning, and contrastive learning from the perspective of perturbation, and can better align the nodes in the two knowledge graphs compared with the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 It is a flowchart of the method of the present invention.

[0066] Figure 2 It is a schematic diagram of the training process of the graph self-attention convolutional neural network in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0068] refer to Figure 1-2 , a data alignment method for a power business system based on a graph convolutional neural network, comprising:

[0069] Acquire equipment ledger information data of two power business systems to be aligned, and clean and pre-process the equipment ledger information data;

[0070] Respectively obtain entities in each power business system and the relationships between the entities, construct a knowledge graph based on the entities and the relationships between the entities, and obtain pre-aligned entity pairs between the two knowledge graphs, wherein the entities are nodes of the knowledge graph and the relationships between the entities are edges of the knowledge graph;

[0071] Inputting the knowledge graph into a graph self-attention convolutional neural network for training, and using the entity pairs as alignment seeds and as supervision information during the training of the graph self-attention convolutional neural network;

[0072] Obtaining an embedded vector representation of each of the nodes through the graph self-attention convolutional neural network, calculating the similarity of each node between the knowledge graphs, and taking the two most similar nodes as alignment nodes;

[0073] The attribute data of the electric power business system entity to be aligned is rewritten according to the alignment node.

[0074] In a specific embodiment, the equipment inventory information data is cleaned and preprocessed, specifically, damaged data in the equipment inventory information data is removed, and the remaining data is processed into a CSV format.

[0075] In one embodiment of the present invention, a knowledge graph is constructed according to the entities and the relationships between the entities, and entity pairs pre-aligned between the two knowledge graphs are obtained, specifically:

[0076] Extracting nodes of the knowledge graph, specifically, using the attributes used to distinguish entities in the equipment ledger information data as nodes of the knowledge graph, and using the text information field of the entity as text features of the node;

[0077] Extracting semantic information of the text features through a language model to obtain an embedded vector representation of the node;

[0078] Constructing edges of the nodes according to the relationships between the entities;

[0079] Get two knowledge graphs G to be aligned k :

[0080] G k ={Ek , R k , T k};

[0081] Among them, k = 1 or 2, E k , R k and T k are respectively the set of nodes, the set of entity relationships, and the triple <e1, r, e2> in the knowledge graph, where e1, e2 ∈ E, r ∈ R, and r is the relationship between entity e1 and entity e2;

[0082] Pre-align the entities according to the attributes with unique values of the entities to obtain entity pairs. The entity pair set S is:

[0083]

[0084] where x and y are respectively the entity nodes of knowledge graphs G1 and G2.

[0085] In a specific embodiment, the language model is the LaBSE model.

[0086] In an implementation manner of the present invention, input the knowledge graph into a graph self-attention convolutional neural network for training, and use the entity pairs as alignment seeds, as the supervision information during the training of the graph self-attention convolutional neural network. Specifically:

[0087] S1. Divide the entity pair set S into a training set and a validation set at a preset ratio, and the test set is the unaligned nodes;

[0088] Use a single-layer graph attention convolutional neural network to perform an aggregation operation on the neighbor information of each node in the knowledge graph. Specifically:

[0089] S2. Randomly sample 20 neighbor nodes of the target node as the neighbor set N i , and the neighbor nodes include the target node itself. Update the embedding vector representation of the target node i through the embedding vector representations of each neighbor node in the neighbor set N The aggregation formula is as follows:

[0090]

[0091]

[0092] where W is a trainable weight matrix, which maps the embedding vector representation of the node to high-level features, and σ is a non-linear activation function Sigmoid, and the calculation method is as follows:

[0093]

[0094] a ij The importance of node j to node i is calculated as follows:

[0095]

[0096] where a is a linear layer that converts a vector into a numerical value, and LeakyReLU is a non-linear activation function, which is expressed as follows:

[0097]

[0098] where p is a coefficient;

[0099] S3. Train the graph self-attention convolutional neural network based on the entities, update the network parameters according to gradient descent, and use Bayesian personalized ranking as the objective function for supervised learning, and the expression is as follows:

[0100]

[0101] where (x, y, y - ) constructs a training triple for node x of the knowledge graph, y is a node pre-aligned with node x in another knowledge graph, and y - is any randomly sampled node other than x and y;

[0102] S4. Conduct unsupervised training based on the graph structure multi-view enhancement method, specifically:

[0103] Obtain the perturbed network parameter θ′ by perturbing the encoder network parameter θ, and input the same knowledge graph node into the network and the perturbed network to obtain two view representations h, h′ of the node, which are expressed as follows:

[0104] h = f(N; θ), h′ = f(N; θ′);

[0105] The way to perturb the encoder is:;

[0106] θ′ l = θ l + η·Δθ l ;

[0107]

[0108] where θ l and θ′ l are the parameters of the l-th layer of the graph self-attention convolutional neural network and the parameters of the l-th layer of the perturbed graph self-attention convolutional neural network respectively, η is an adjustable perturbation intensity hyperparameter, and Δθ l is a perturbation term of a Gaussian distribution with a mean of zero and a variance of ;

[0109] InfoNCE is used as the target optimization function to bring the original representation and perturbation representation of the same node closer, and push the perturbation representation of other nodes further away, as shown below:

[0110]

[0111] Where N is the number of nodes in a training batch, sim is the cosine similarity, τ is an adjustable parameter;

[0112] A perturbed graph self-attention convolutional neural network is obtained by perturbing the parameters of the graph self-attention convolutional neural network, wherein the node outputs the original view and the perturbed view representation of the node under the original graph self-attention convolutional neural network and the perturbed graph self-attention convolutional neural network respectively, and a more robust graph self-attention convolutional neural network can be obtained by bringing the original view representation and the perturbed view representation of the same node closer and moving the perturbed view representations with other nodes further away;

[0113] S5. Repeat steps S2 to S5 until the value of the objective function converges or a preset number of training times is reached.

[0114] In a specific embodiment, the entity pair set S is divided into a training set and a validation set in a ratio of 4:1.

[0115] In one embodiment of the present invention, unaligned nodes of the knowledge graph with fewer nodes in the two knowledge graphs are used as test sets.

[0116] Since knowledge graphs with many nodes will have a large number of residual unaligned nodes, which have no real aligned entities in knowledge graphs with few nodes, the unaligned nodes of knowledge graphs with few nodes are used as test sets.

[0117] In one embodiment of the present invention, the embedded vector representation of each node is obtained by the graph self-attention convolutional neural network, the similarity of each node between the knowledge graphs is calculated, and the two most similar nodes are used as alignment nodes, specifically:

[0118] Calculate the embedding vector representation of each node of the knowledge graph through the graph self-attention convolutional neural network;

[0119] Calculate the similarity of each node representation between two knowledge graphs based on the Euclidean norm:

[0120]

[0121] Take the two most similar nodes as the alignment nodes.

[0122] In an embodiment of the present invention, it further includes performing performance statistics on the labeled training set and validation set nodes, and manually inspecting the unlabeled test set nodes to determine whether they are valid. Specifically:

[0123] Calculate Hits@1 and Hits@10 for the labeled training set and validation set nodes:

[0124]

[0125] where S is the set of aligned node pairs, |S| is the number of aligned node pairs, and rank i is the link prediction ranking of the i-th aligned node pair, and I is a judgment function. If it is true, then I = 1; otherwise, I = 0.

[0126] Manually inspect the unlabeled test set nodes to determine whether they are valid.

[0127] Referring to Table 1, comparing the method of the present invention with the entity alignment technology based on the relational perception bi-graph convolutional network, it is proved that the method of the present invention combining the graph self-attention convolutional neural network, supervised learning, and contrastive learning from the perturbation perspective is superior to the prior art. The method of the present invention can better perform the alignment operation on the nodes in the two knowledge graphs.

[0128] Table 1

[0129] technology Hit@1 Hit@10 RDGCN 0.45 0.65 the method of the present invention 0.711 0.872

[0130] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structures made by using the specification and drawings of the present invention, directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token.

Claims

1. A data alignment method for a power service system based on a graph convolutional neural network, characterized in that include: Acquire equipment ledger information data of two power business systems to be aligned, and clean and pre-process the equipment ledger information data; Respectively obtain entities in each power business system and the relationships between the entities, construct a knowledge graph based on the entities and the relationships between the entities, and obtain pre-aligned entity pairs between the two knowledge graphs, wherein the entities are nodes of the knowledge graph and the relationships between the entities are edges of the knowledge graph; The knowledge graph is input into the graph self-attention convolutional neural network for training, and the entity pair is used as an alignment seed and as supervision information during the training of the graph self-attention convolutional neural network, specifically: S1, dividing the entity pair set S into a training set and a validation set according to a preset ratio, and the test set is the unaligned nodes; S2, use a single-layer graph attention convolutional neural network to aggregate the neighbor information of each node in the knowledge graph, S3. Based on the entity, the graph self-attention convolutional neural network is trained, the network parameters are updated according to the gradient descent, and the Bayesian personalized ranking is used as the objective function of supervised learning. The expression is as follows: Among them, (x, y, y - ) is a triple constructed for training the node x of the knowledge graph, y is a node pre-aligned with the node x in another knowledge graph, and y - is any randomly sampled node other than x and y; S4. Unsupervised training based on graph structure multi-view enhancement method, specifically: By perturbing the encoder network parameter θ to obtain the perturbation network parameter θ′, we can then obtain the node representations h, h′ of the same node x in the knowledge graph under the network and the perturbation network, as shown below: h=f(x; θ), h′=f(x; θ′); The way to perturb the encoder is: θ l ' = θ l + η · Δθ l ; where, θ l and θ l ′ are the parameters of the l-th layer graph self-attention convolutional neural network and the parameters of the l-th layer perturbed graph self-attention convolutional neural network respectively, η is an adjustable perturbation intensity hyperparameter, and Δθ l is a Gaussian distribution with mean zero and variance and is the perturbation term; InfoNCE is used as the target optimization function to bring the original representation h of the same node n closer. n and the perturbation h′ n , push away the disturbance representation with other nodes, expressed as follows: where N is the number of nodes in a training batch, n′ is a node other than node n, sim is the cosine similarity, τ is an adjustable parameter; S5, repeating steps S2 to S5 until the value of the objective function converges or reaches a preset number of training times; Obtaining an embedded vector representation of each of the nodes through the graph self-attention convolutional neural network, calculating the similarity of each node between the knowledge graphs, and taking the two most similar nodes as alignment nodes; The attribute data of the electric power business system entity to be aligned is rewritten according to the alignment node.

2. The data alignment method for the power service system based on the graph convolutional neural network according to claim 1, wherein The equipment ledger information data is cleaned and preprocessed, specifically, damaged data in the equipment ledger information data is removed, and the remaining data is processed into CSV format.

3. The method for aligning power service system data based on a graph convolutional neural network according to claim 1, wherein A knowledge graph is constructed based on the entities and the relationships between the entities, and a pre-aligned entity pair between the two knowledge graphs is obtained, specifically: Extracting nodes of the knowledge graph, specifically, using the attributes used to distinguish entities in the equipment ledger information data as nodes of the knowledge graph, and using the text information field of the entity as text features of the node; Extracting semantic information of the text features through a language model to obtain an embedded vector representation of the node; Constructing edges of the nodes according to the relationships between the entities; Obtain two knowledge graphs G to be aligned k : G k = {E k , R k , T k}; where k = 1 or 2, E k , R k and T k are respectively the set of nodes, the set of entity relationships, and the triple <e1, r, e2> in the knowledge graph, where e1, e2 ∈ E, r ∈ R, and r is the relationship between entity e1 and entity e2; The entities are pre-aligned according to the attributes of the entities with unique values ​​to obtain entity pairs. The entity pair set S is: Among them, x and y are the entity nodes of the knowledge graphs G1 and G2 respectively.

4. The data alignment method for the power service system based on the graph convolutional neural network according to claim 3, wherein The language model is a LaBSE model.

5. The method for aligning power service system data based on a graph convolutional neural network according to claim 4, characterized in that A single-layer graph attention convolutional neural network is used to aggregate the neighbor information of each node in the knowledge graph, specifically: Randomly sample 20 neighbor nodes of the target node as the neighbor set N i , where the neighbor nodes include the target node itself. Through the embedding vectors of each neighbor node in the neighbor set N i to update the embedding vector representation of the target node . The aggregation formula is as follows: The aggregation formula is as follows: Among them, W is a trainable weight matrix that maps the embedding vector representation of the node to high-level features, and σ is the non-linear activation function Sigmoid. The calculation method is as follows: a ij The importance of node j to node i is calculated as follows: Among them, a is a linear layer that converts the vector into a numerical value, and LeakyReLU is a non-linear activation function, which is expressed as follows: Among them, p is an adjustable coefficient.

6. The method for aligning power service system data based on a graph convolutional neural network according to claim 5, characterized in that, The entity pair set S is divided into a training set and a validation set in a ratio of 4:

1.

7. The method for aligning power service system data based on a graph convolutional neural network according to claim 5, characterized in that The unaligned nodes of the knowledge graph with fewer nodes in the two knowledge graphs are used as the test set.

8. The method for aligning power service system data based on a graph convolutional neural network according to claim 5, characterized in that, The embedding vector representation of each node is obtained through the graph self-attention convolutional neural network, the similarity between the nodes of the knowledge graphs is calculated, and the two most similar nodes are used as the aligned nodes. Specifically: The embedding vector representation of each node of the knowledge graph is calculated through the graph self-attention convolutional neural network; The similarity between the representations of the nodes between the two knowledge graphs is calculated according to the Euclidean norm: The two most similar nodes are taken as the aligned nodes.

9. The method for aligning power service system data based on a graph convolutional neural network according to claim 5, characterized in that It also includes performing performance statistics on the labeled training set and validation set nodes, and manually checking the unlabeled test set nodes to determine whether they are valid. Specifically: Calculate Hits@1 and Hits@10 for the labeled training set and validation set nodes: where S is the set of aligned node pairs, |S| is the number of aligned node pairs, rank i is the link prediction rank of the i-th aligned node pair, I is a judgment function. If it is true, then I = 1; otherwise, I = 0. Manually check the unlabeled test set nodes to determine whether they are valid.

Citation Information

Patent Citations

  • A text relationship extraction method and system based on a hierarchical knowledge graph attention model

    CN109902171A

  • Training method of knowledge graph alignment model based on graph neural network

    CN113807520A