Transformer substation anti-error entity alignment method and system based on multi-modal structure fusion

Through the multimodal structure fusion method, the multimodal information of the power grid is integrated, combined with Wasserstein and Gromov-Wasserstein optimization, efficient, precise alignment and real-time intelligent detection of power grid equipment are achieved, and the problems of insufficient modal utilization and computational scalability in the existing technology are solved.

CN120408306APending Publication Date: 2025-08-01WUXI POWER SUPPLY BRANCH OF STATE GRID JIANGSU ELECTRIC POWER CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510490082.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the power grid scenario, the existing technology has modal utilization, rough structural alignment, weak anti-interference ability and insufficient computational scalability, making it difficult to achieve efficient and accurate cross-system device alignment and real-time intelligent detection of anti-error operations.

Method used

The multimodal structure fusion method is adopted, and the device text, images, attributes and topological structure information are integrated through the cross-modal dynamic fusion mechanism, combined with Wasserstein semantic alignment and Gromov-Wasserstein structure alignment, to build a hybrid optimal transmission optimization target, introduce a multimodal cross-verification strategy, and the block acceleration strategy reduces the computational complexity.

Benefits of technology

It realizes efficient and accurate cross-system equipment alignment, supports grid equipment portrait construction, rapid fault traceability and preventive maintenance, provides reliable data support for real-time intelligent detection, and improves the accuracy and robustness of grid knowledge graph entity alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408306A_ABST
    Figure CN120408306A_ABST
Patent Text Reader

Abstract

The invention discloses a transformer substation anti-error entity alignment method and system based on multi-modal structure fusion, and belongs to the technical field of intelligent power grids, and the method comprises the steps: carrying out the cross-system multi-modal data collection of a power grid, carrying out the preprocessing and feature extraction, and generating multi-modal features; multi-modal features are fused to generate joint embedding, a multi-modal similarity matrix is calculated, an initial coupling matrix is solved in combination with a Sinkhorn algorithm, and an initial anchor point set is screened; iteratively updating the initial anchor point set in combination with multi-modal consistency verification to obtain a verification anchor point set; and taking the iterated verification anchor point set as priori knowledge, injecting the priori knowledge into a Gromov-Wasserstein structure optimization model, and carrying out block calculation on the structure similarity of an adjacent matrix to realize global topology alignment. According to the method, cross-system entity alignment is realized through three-stage progressive optimization, the alignment result is fused with equipment entities, attributes, relationships and dynamic states, and the method plays an important role in scenes such as real-time intelligent detection of substation anti-misoperation and the like in combination with a logical reasoning engine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of smart grids, relates to the technology of multimodal entity alignment in knowledge graphs, and more specifically, relates to a substation anti-misoperation entity alignment method and system based on multimodal structure fusion. Background Art

[0002] With the rapid development of smart grids, the power system has gradually shifted from a traditional isolated data management mode to an intelligent and integrated knowledge-driven mode. As a technical means capable of effectively integrating multi-source heterogeneous data, knowledge graphs have shown great potential in fields such as power grid equipment management, fault diagnosis, and load forecasting. By constructing a unified power grid knowledge graph, cross-system data fusion and intelligent reasoning can be achieved. However, the core challenge in this process lies in how to accurately align entities scattered in different systems, which may exhibit high heterogeneity due to differences in naming rules, data modalities, and storage formats. This heterogeneity not only leads to the phenomenon of data islands but also makes cross-system collaborative analysis and decision-making extremely difficult. Currently, entity alignment technologies are mainly divided into three categories: rule-based methods, embedding learning methods, and optimization alignment methods. However, these methods all have significant limitations in the power grid scenario.

[0003] Rule-based methods rely on manually predefined matching rules, such as identifying variants of device names through string similarity or regular expressions, or screening candidate entities through numerical attributes. Although such methods have certain effects in simple scenarios, their generalization ability is poor and it is difficult to handle the complex diversity of power grid device naming rules. Embedding learning methods map entities to a low-dimensional vector space and use embedding similarity to achieve alignment, which alleviates the rigidity problem of rule-based methods to a certain extent. In addition, high-quality annotations in the power grid field are scarce and the acquisition cost is high, which limits the practicality and scalability of the model. Optimization alignment methods (such as optimal transport and graph matching) directly solve the entity correspondence relationship through mathematical programming, avoiding the representation bias problem in embedding learning. The computational complexity of traditional GW alignment is relatively high and cannot support the real-time processing requirements of millions of entities in the power grid. The errors of low-quality anchor points will gradually accumulate during the iterative optimization process, ultimately destroying the stability of the global alignment.

[0004] Meanwhile, the complexity and real-time requirements of the anti-misoperation logic are also increasing day by day. Traditional anti-misoperation systems mainly rely on manually predefined rule libraries (such as the "five-prevention" logic) and static topology checks, and there are the following pain points: First, the rule coverage is incomplete: it is difficult for manual rules to cover the complex interlock relationships of multiple devices and scenarios (such as operations across voltage levels and topological changes after the access of new energy); Second, the dynamic adaptability is poor: when the device status (such as the opening and closing of circuit breakers and the position of grounding switches) and the power grid operation mode change, the static rules cannot adjust the logic constraints in real time; Third, the multi-source data is fragmented: SCADA real-time signals, GIS topologies, device ledgers, and sensor data are scattered and isolated, lacking unified knowledge expression and reasoning capabilities; Fourth, the computational scalability is insufficient, and the algorithm complexity limits its application in large-scale power grids.

[0005] These challenges have given rise to the technical necessity of the present invention. There is an urgent need for an entity alignment method that can deeply integrate multi-modal data, accurately capture topological constraints, be highly noise-resistant and scalable. Summary of the Invention

[0006] Object of the Invention: The existing technologies have one-sided utilization of modalities, rough structure alignment, weak anti-interference ability, and insufficient computational scalability in the power grid scenario. The present invention aims to provide a substation anti-misoperation logic verification entity alignment method based on the fusion of multi-modal and structural information. Through a cross-modal dynamic fusion mechanism, it adaptively integrates multi-modal information such as device text, images, attributes, and topological structures to generate a robust joint embedding representation; combines Wasserstein semantic alignment and Gromov-Wasserstein structural alignment to construct a hybrid optimal transport optimization objective to ensure global topological consistency; introduces a multi-modal cross-validation strategy to iteratively screen high-confidence anchor points to reduce noise interference; adopts a block acceleration strategy to divide the large-scale power grid topology by voltage level, significantly improving the computational efficiency. The present invention can efficiently and accurately achieve cross-system device alignment, support the construction of power grid device portraits, rapid fault tracing, and preventive maintenance optimization, and provide reliable data support and knowledge services for the realization of real-time intelligent detection of anti-misoperation operations.

[0007] To solve the deficiencies in the existing technologies, the present invention provides a substation anti-misoperation entity alignment method and system based on multi-modal structure fusion.

[0008] The present invention adopts the following technical solutions.

[0009] In the first aspect of the present invention, a substation anti-misoperation entity alignment method based on multi-modal structure fusion is provided, including the following steps:

[0010] Perform cross-system multi-modal data collection of the power grid, including collecting text data and attribute data from the SCADA system, topological data from the GIS system, and device image data from the inspection system;

[0011] Preprocess and extract features from the collected multi-modal power grid data to generate multi-modal features, including text semantic features, image features, attribute features, and topological structure features;

[0012] Fuse the multi-modal features to generate a joint embedding, calculate the multi-modal similarity matrix, and combine with the Sinkhorn algorithm to solve the initial coupling matrix and screen the initial anchor point set;

[0013] Iteratively update the initial anchor point set by combining multi-modal consistency verification to obtain the verified anchor point set;

[0014] Inject the iterated verified anchor point set into the Gromov-Wasserstein structure optimization model as prior knowledge, calculate the structural similarity of the adjacency matrix in blocks, and achieve global topological alignment.

[0015] Optionally, the text data includes device names, specification descriptions, technical parameters, and operation and maintenance record data;

[0016] The attribute data includes static attributes, dynamic attributes, and status indicators. The static attributes include commissioning time and installation location, the dynamic attributes include real-time voltage, load rate, and temperature, and the status indicators include cumulative operation times and remaining life prediction values;

[0017] The topological data includes an adjacency matrix, hierarchical relationships, and dynamic topologies. The dynamic topologies include the real-time status of circuit breakers and the flow direction of line loads;

[0018] The image data includes device appearance images, infrared thermal imaging maps, and historical defect images.

[0019] Optionally, the preprocessing of the multi-modal data includes:

[0020] Clean and standardize the text data, fill in missing values and normalize the attribute data, construct a matrix and enhance self-connection for the topological data, and generate and normalize missing virtual images for the image data.

[0021] Optionally, the feature extraction of the preprocessed multi-modal data includes:

[0022] Extract the features of the preprocessed text data, encode the device name and specification description through the SimCSE model, and extract the device text semantic feature F text , and output a 256-dimensional semantic vector F text ∈R 256 ;

[0023] Extract the features of the preprocessed image data, and extract the original image feature F′ through the pre-trained ResNet-50 image ∈R2048 , and reduce the dimension to 256 dimensions to obtain the image embedding F image ∈R 256 ;

[0024] Extract the feature of the preprocessed attribute data, and generate the attribute embedding F attr ∈R 256 ;

[0025] Extract the feature of the preprocessed topological data, encode the power grid adjacency matrix through GCN, and output the 256-dimensional structure embedding F struc ∈R 256 .

[0026] Optionally, the fusion of multi-modal features to generate the joint embedding includes:

[0027] Calculate the weights of each modality through the attention mechanism and generate the joint embedding F fused ;

[0028] Perform component extraction of the joint embedding F fused to obtain the text component image component attribute component topological component

[0029] Optionally, the calculation of the multi-modal similarity matrix includes:

[0030] According to the text embedding F text , image embedding F image , attribute embedding F attr and structure embedding F struc , calculate the text similarity matrix C text , image similarity matrix C image , attribute similarity matrix C attr and structure similarity matrix C struc ;

[0031] According to the text similarity matrix C text , image similarity matrix C image , attribute similarity matrix C attr and structure similarity matrix C struc , calculate the multi-modal similarity matrix C sum .

[0032] Optionally, the combination of the Sinkhorn algorithm to solve the initial coupling matrix and screen the initial anchor point set includes:

[0033] According to the multi-modal similarity matrix C sum , use the Sinkhorn algorithm to solve the initial coupling matrix π 0, and filter the initial anchor set

[0034]

[0035] c = 1 / max(m,n)

[0036]

[0037] where π is the coupling matrix of the alignment probabilities between entities in two knowledge graphs; KL(π‖π prior ) is the Kullback-Leibler divergence between the coupling matrix π and the prior distribution π prior ; π ij is the alignment probability between entity e i in the source graph and entity e' j in the target graph; π prior,ij represents the alignment probability between entity e i in the source graph and entity e' j in the target graph in the prior distribution; c represents the maximum possible value of each element in the coupling matrix π; ∈ represents the threshold parameter for filtering the initial anchors; m and n are the numbers of entities in the source graph and the target graph respectively; e i represents the i-th entity in the source knowledge graph; e' j represents the j-th entity in the target knowledge graph.

[0038] Optionally, the iterative update of the initial anchor set by combining multi-modal consistency verification to obtain the verified anchor set includes:

[0039] Update the anchor set Retain the candidate pairs with text similarity and image similarity higher than the thresholds:

[0040]

[0041] In the formula, θ text and θ image represent the text similarity threshold and the image similarity threshold respectively.

[0042] Optionally, the injection of the iterated verified anchor set as prior knowledge into the Gromov-Wasserstein structure optimization model includes:

[0043] The objective function f FGW of the Gromov-Wasserstein structure optimization model is:

[0044]

[0045] where α is the balance parameter between semantic alignment and structural alignment; is the adjacency matrix value of entities i and j in the source knowledge graph; is the adjacency matrix value of entities k and l in the target knowledge graph; π ik is the coupling probability between source entity i and target entity k; π jl is the coupling probability between source entity j and target entity l; λ is the weight parameter; is the set of entity pairs;

[0046] Update π using the BPG algorithm k+1 , and use the Sinkhorn iteration to solve the sub-problem until convergence:

[0047]

[0048] where f GWD is the Gromov-Wasserstein distance; A s and A t are the adjacency matrices of the source knowledge graph and the target knowledge graph respectively; β is the step size parameter.

[0049] The second aspect of the present invention provides a substation anti-misoperation entity alignment system based on multi-modal structure fusion. Based on the method for substation anti-misoperation entity alignment based on multi-modal structure fusion described in the first aspect of the present invention, the system includes:

[0050] A multi-modal data acquisition module for performing cross-system multi-modal data acquisition of the power grid;

[0051] A multi-modal feature processing module for preprocessing and feature extraction of the acquired power grid multi-modal data to generate multi-modal features;

[0052] A dynamic fusion and embedding generation module for fusing multi-modal features to generate joint embeddings and calculating a multi-modal similarity matrix;

[0053] An initial coupling solving module for solving the initial coupling matrix based on the Sinkhorn algorithm and screening the initial anchor point set;

[0054] An anchor point iterative verification module for iteratively updating the initial anchor point set by combining multi-modal consistency verification to obtain a verified anchor point set;

[0055] A structure optimization and alignment module injects the iteratively verified anchor point set as prior knowledge into the Gromov-Wasserstein structure optimization model, and calculates the structural similarity of the adjacency matrix in blocks to achieve global topological alignment.

[0056] Compared with the prior art, the beneficial effects of the present invention at least include:

[0057] First of all, in the present invention, a multi-modal dynamic fusion mechanism is adopted. Through the cross-modal attention mechanism, multi-modal information such as device text, images, attributes, and topological structures is adaptively integrated to generate a robust joint embedding representation. This mechanism can not only make full use of the complementarity of multi-modal data, but also reduce the number of model parameters and computational complexity through weight sharing, while avoiding the information loss problem caused by modal fragmentation in traditional methods. In addition, the multi-modal dynamic fusion mechanism can effectively highlight the differences of entities in different data sources without any prior alignment seed data, thereby improving the alignment accuracy.

[0058] Secondly, the present invention introduces an improved Gromov-Wasserstein distance optimization method, which combines Wasserstein semantic alignment and Gromov-Wasserstein structural alignment to construct a hybrid optimal transport optimization objective. This method not only considers the semantic similarity between entities, but also ensures that the alignment result conforms to the physical connection rules of the power grid through global topological consistency constraints. The improved optimization algorithm significantly reduces the computational complexity of traditional GW alignment, enabling it to support the real-time processing requirements of large-scale power grids.

[0059] Finally, the present invention combines the self-attention mechanism and the residual structure to design a multi-modal cross-validation strategy. By iteratively screening high-confidence anchor points, the impact of noise interference on the alignment result is reduced. The self-attention mechanism can capture the interaction relationships between global modalities, while the residual structure enhances the stability of feature extraction, ensuring that the model can still maintain a high alignment accuracy in the presence of noise or missing multi-modal data.

[0060] In summary, compared with the prior art, the present invention, by means of the multi-modal dynamic fusion mechanism and the improved Gromov-Wasserstein optimization method, ingeniously integrates multi-modal information such as text, images, attributes, and topological structures, and precisely regulates the information interaction between different modalities through the cross-modal self-attention mechanism and global topological consistency constraints, ensuring that the synergistic advantages of multi-modal data are fully exerted. This method significantly improves the accuracy and robustness of entity alignment in the power grid knowledge graph, and at the same time reduces the computational complexity through the block acceleration strategy, enabling it to efficiently support the real-time application of large-scale power grids and having important applications in the intelligent detection technology of anti-misoperation logic in automated substations. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 is a flowchart of the method provided by an embodiment of the present invention;

[0062] Figure 2 is a schematic diagram of the entity alignment application in a specific power grid scenario provided according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0064] To more clearly introduce the prominent substantive features of the present invention and the significant progress brought to the prior art, the following introduces an application example of implementing the present invention.

[0065] The following will detail the embodiments of the present invention in conjunction with the accompanying drawings. The application example specifically includes:

[0066] The present invention provides a method for entity alignment of anti-maloperation logic verification in a substation based on the fusion of multi-modal and structural information.

[0067] Based on the above method, as Figure 1 shown, in this embodiment, the above method is applied to power grid equipment management and fault diagnosis. The specific process is as follows:

[0068] S1. Perform cross-system multi-modal data collection of the power grid, including collecting text data and attribute data from the SCADA system, topological data from the GIS system, and device image data from the inspection system.

[0069] Preferably, the S1 specifically includes:

[0070] A1: Collect text data from the SCADA system, including device name, specification description, technical parameters, and operation and maintenance records.

[0071] Exemplarily, the device name includes "Main Transformer #1", "Circuit Breaker CB-01"; the specification description includes signal and manufacturer, model (such as "SFZ-220MVA"), manufacturer (such as "ABC"); the technical parameters include rated voltage and capacity, rated voltage (such as "220 kV"), capacity (such as "300 MVA"); the operation and maintenance records include maintenance time and fault description, maintenance time (such as "2025-01-10"), fault description (such as "Winding overheating on 2025-01-10");

[0072] A2: Collect attribute data from the SCADA system simultaneously, including static attributes: commissioning time (such as "2025-01"), installation location (such as "longitude 118.78°, latitude 32.04°"); dynamic attributes: real-time voltage (such as "220kV±5%"), load rate (such as "85%"), temperature (such as "65°C"); status indicators: cumulative operation times (such as "circuit breaker opening and closing times"), predicted remaining life value.

[0073] A3: Collect topological data from the GIS system, including adjacency matrix: regarding the connection relationship of devices (A ij = 1 indicates that device i is connected to device j); hierarchical relationship: the tree structure of substation → busbar → feeder → terminal equipment; dynamic topology: real-time status of circuit breaker (closed / open), line load flow direction.

[0074] A4: Collect image data from the inspection system, including device appearance images: the overall view of the device and key components; infrared thermal imaging diagrams: surface temperature distribution of the device, image resolution (such as "1024×768, storage format is JPEG"); historical defect images: archiving of abnormal states such as rust and cracks.

[0075] S2. Preprocess and extract features from the collected power grid multi-modal data to generate multi-modal features, including text semantic features, image features, attribute features, and topological structure features.

[0076] Preferably, in the above S2, the preprocessing of multi-modal data includes:

[0077] Clean and standardize the text data, fill in missing values and normalize the attribute data, construct matrices and enhance self-connections for the topological data, and generate and normalize missing virtual images for the image data, specifically including:

[0078] B1. Clean the text data collected by the SCADA system: remove special characters (such as "T-01@Sub_A" → "T-01"); standardize: unify the naming rules (such as "Main Transformer #1" → "Transformer-01");

[0079] B2. Simultaneously fill in the missing values for the attribute data collected by the SCADA system: fill with the mean value of similar devices (such as taking the mean value of the same model when the capacity is missing); normalize: standardize the numerical attributes to the interval [0,1] to obtain x norm .

[0080]

[0081] where μ hist is the average value of historical data; σ histis the standard deviation of historical data; x is the original sensor data value; x norm is the normalized sensor data value.

[0082] B3, construct a matrix for the topological data collected by the GIS system: generate an adjacency matrix A ∈ R from the data N×N ; and use enhanced self-connection to generate Increase I N self-loops;

[0083]

[0084] Then perform feature initialization: construct a node attribute matrix X ∈ R N×n (n includes attributes such as voltage u, capacity v, and commissioning time t);

[0085] where N is the number of nodes; n is the number of node attributes; I N is the N×N identity matrix.

[0086] B4, use the conditional generative adversarial network CGAN to generate virtual images I for the image data collected by the inspection system gen ;

[0087] I gen = CGAN(concat(enc text (T), enc attr (A)))

[0088] where T is the device name text, A is the attribute vector, (enc text (T)) is the SimCSE text encoder, and (enc attr (A)) is the attribute encoding MLP;

[0089] Then perform normalization processing to adjust the image size to a unified resolution (such as "512×512").

[0090] Preferably, in S2, the feature extraction of the preprocessed multi-modal data includes: extracting the device text semantic feature F through SimCSE text , extracting the image feature F through ResNet-50 image , extracting the attribute feature F through MLP attr , extracting the topological structure feature F through GCN struc , and then generating a joint embedding representation F through a cross-modal attention mechanism fused , specifically including:

[0091] C1, text feature extraction: use the SimCSE model to encode the device name and specification description, and output a 256-dimensional semantic vector F text ∈R256 ;

[0092] C2, Image feature extraction: Extract the original image feature F' through pre-trained ResNet-50 image ∈R 2048 , and reduce the dimension to 256 dimensions to obtain the image embedding F image ∈R 256 ;

[0093] F image = ReLU(W I ·F' image + b I )

[0094] where W I is the weight matrix used to perform a linear transformation on F' image to achieve the dimension reduction operation; b I is the bias term.

[0095] C3, Attribute feature extraction: Input the numerical attributes (here, "voltage, capacity, commissioning time") into a 3-layer MLP (input layer: 3 dimensions, hidden layer: 64 dimensions, output layer: 256 dimensions) to generate the attribute embedding F attr ∈R 256 ;

[0096] F attr = MLP([u, v, t])

[0097] C4, Topological feature extraction: Use a 2-layer GCN to encode the power grid adjacency matrix and output a 256-dimensional structural embedding F struc ∈R 256 ;

[0098]

[0099] F struc = H (2)

[0100] where H (1) and H (2) respectively represent the output feature matrices of the first and second layers of GCN; W (0) and W (1) are the weight matrices in the first and second layers of GCN respectively; is the degree matrix after adding self-connections.

[0101] S3, Fuse multi-modal features to generate a joint embedding, calculate the multi-modal similarity matrix, and combine the Sinkhorn algorithm to solve the initial coupling matrix and screen the initial anchor set.

[0102] Preferably, in the above S3, fusing multi-modal features to generate a joint embedding includes:

[0103] Cross-modal dynamic fusion: Calculate the weights of each modality through the attention mechanism and generate the joint embedding F fused , F fused is the high-level representation of the entity's multi-modal features. Its core role is to provide unified semantic and structural perception features for subsequent entity alignment, ensuring the complementarity and consistency of different modality information.

[0104]

[0105] F fused = ∑ m∈{text,image,attr,struc} α m F m where W m is the learnable weight matrix of modality m; F m is the feature vector (embedding representation) of modality m; α m is the attention weight of modality m, and its range is [0, 1].

[0106] To further improve the alignment effect, the component extraction of the joint embedding F fused needs to be carried out separately, including the text component image component attribute component topology component The calculation formula is as follows:

[0107]

[0108]

[0109] Preferably, in the step S3, calculating the multi-modal similarity matrix includes:

[0110] Calculating the corresponding similarity matrix C text , C image , C attr , C struc , and calculating the multi-modal similarity matrix C sum , then using the Sinkhorn algorithm to solve the initial coupling matrix π 0 , then screening the initial high-confidence anchor point set and gradually iteratively updating the anchor point set Finally, optimizing the global structural alignment in combination with the Gromov-Wasserstein distance, specifically including:

[0111] D1: According to the text embedding F text , image embedding F image , attribute embedding F attr and structural embedding F struc, calculate the text similarity matrix C text , the image similarity matrix C image , the attribute similarity matrix C attr and the structure similarity matrix C struc ;

[0112]

[0113] Among them, the superscripts s and t respectively represent the entity feature vectors in the source knowledge graph and the entity feature vectors in the target knowledge graph;

[0114] D2: According to the text similarity matrix C text , the image similarity matrix C image , the attribute similarity matrix C attr and the structure similarity matrix C struc , calculate the multimodal similarity matrix C sum ;

[0115] C sum = C text + C image + C attr + C struc

[0116] Preferably, in the S3, combining the Sinkhorn algorithm to solve the initial coupling matrix and screening the initial anchor point set includes:

[0117] According to the multimodal similarity matrix C sum , use the Sinkhorn algorithm to solve the initial coupling matrix π 0 , and screen the initial high-confidence anchor point set

[0118]

[0119] c = 1 / max(m,n)

[0120]

[0121] Among them, π represents the coupling matrix of the alignment probabilities between entities in the two knowledge graphs; KL(π‖π prior ) represents the Kullback-Leibler divergence between the coupling matrix π and the prior distribution π prior ; π ij represents the alignment probability between the entity e i in the source graph and the entity e' j in the target graph; π prior,ij represents the alignment probability between the entity e i in the source graph and the entity e' j in the prior distribution and the target graphAlignment probability; c represents the maximum possible value of each element in the coupling matrix π; ∈ represents the threshold parameter for screening high-confidence anchor points, usually taking m and n are the number of entities in the source graph and the target graph respectively; e i represents the i-th entity in the source knowledge graph; e′ j represents the j-th entity in the target knowledge graph.

[0122] S4. Combine multi-modal consistency verification to iteratively update the initial anchor point set screened in S3 to obtain a high-confidence anchor point set.

[0123] Preferably, the S4 includes updating the anchor point set Only retain the candidate pairs that are higher than the threshold in text similarity and image similarity;

[0124]

[0125] In the formula, θ text and θ image respectively represent the text similarity threshold (such as 0.85) and the image similarity threshold (such as 0.8); and are respectively the text feature vectors of e i and e′ j ; and are respectively the image feature vectors of e i and e′ j .

[0126] It should be noted that, aiming at the problem of single-modal noise sensitivity existing in the entity alignment method in the prior art, the present invention reduces the noise interference through multi-modal consistency anchor point screening, and the cross-verification means also reduces the mis-matching rate of entity alignment.

[0127] S5. Inject the iterated high-confidence anchor point set into the Gromov-Wasserstein structure optimization model as prior knowledge, calculate the structural similarity of the adjacency matrix in blocks, and achieve global topological alignment.

[0128] Preferably, in the S5, the objective function f FGW of the Gromov-Wasserstein structure optimization model is:

[0129]

[0130] Among them, α is the balance parameter between semantic alignment and structural alignment, and its value range is [0,1]; is the adjacency matrix value of entities i and j in the source knowledge graph; is the adjacency matrix value of entities k and l in the target knowledge graph; πik is the coupling probability between the source entity i and the target entity k; π jl is the coupling probability between the source entity j and the target entity l; λ is the weight parameter, which is adjusted according to task requirements, model convergence, etc.; is the set of entity pairs;

[0131] Update π using the BPG algorithm k+1 , and use the Sinkhorn iteration to solve the sub-problem until convergence:

[0132]

[0133] where f GWD is the Gromov-Wasserstein distance; A s and A t are the adjacency matrices of the source knowledge graph and the target knowledge graph respectively; β is the step size parameter.

[0134] It should be noted that, aiming at the problem of poor dynamic adaptability in preventing misoperations in the prior art, the present invention can elastically adapt to topological changes by introducing the Gromov-Wasserstein optimization model, and solves the problem that it is difficult for traditional methods to adapt to the dynamic differences of adjacency relationships due to the time-varying power grid topology.

[0135] It is worth noting that the present invention realizes cross-system entity alignment through three-stage progressive optimization. The three stages include initial semantic matching, iterative anchor point screening, and global structure alignment. The alignment results integrate device entities, attributes, relationships, and dynamic states, and combine with a logical reasoning engine, which plays an important role in scenarios such as real-time intelligent detection of substation misoperation prevention.

[0136] In Embodiment 2 of the present invention, a substation misoperation prevention entity alignment system based on multi-modal structure fusion is provided. Based on the substation misoperation prevention entity alignment based on multi-modal structure fusion described in Embodiment 1, the system includes:

[0137] A multi-modal data acquisition module, which is used to perform cross-system multi-modal data acquisition of the power grid, synthesize multi-scale features, and integrate multi-level information;

[0138] A multi-modal feature processing module, which is used to preprocess and extract features from the acquired power grid multi-modal data, generate multi-modal features, and mine the deep semantic information of the multi-modal data;

[0139] A dynamic fusion and embedding generation module, which is used to fuse multi-modal features to generate joint embeddings, calculate a multi-modal similarity matrix, and regulate the information attention of text, images, attributes, and topological features through a self-attention mechanism to enhance different modal features;

[0140] An initial coupling solving module for solving an initial coupling matrix based on the Sinkhorn algorithm and screening an initial anchor point set;

[0141] An anchor point iterative verification module for iteratively updating the initial anchor point set by combining multimodal consistency verification to obtain a verified anchor point set;

[0142] A structure optimization and alignment module injects the iterated verified anchor point set as prior knowledge into the Gromov-Wasserstein structure optimization model, calculates the structural similarity of the adjacency matrix in blocks, and realizes global topological alignment;

[0143] It should be noted that by passing the separately extracted multimodal features through the Gromov-Wasserstein optimization module, on the one hand, combining Wasserstein semantic alignment and Gromov-Wasserstein structural alignment to optimize the global topological consistency, and on the other hand, reducing the computational complexity through a block acceleration strategy.

[0144] Preferably, the system further includes an anti-error detection application module for mapping the entity alignment result to an anti-error operation rule library, real-time detecting abnormal topological connections and generating a locking control instruction to realize real-time intelligent detection of anti-error operations.

[0145] It should be noted that in view of the technical defects such as insufficient semantic structure collaboration and cross-modal representation deviation caused by the lack of multimodal supervision and / or in-depth optimization of topological consistency in the existing frameworks in the prior art, the present invention aims to make the deep integration of multimodal supervision signals and the optimization of topological consistency support each other functionally through cross-modal dynamic fusion, hybrid optimal transport optimization, and block acceleration strategy, and has achieved the technical effects of significantly improving the accuracy and robustness of entity alignment in the power grid knowledge graph and being able to efficiently support the real-time applications of large-scale power grids, breaking through the technical bottleneck of isolated processing of multimodal information and insufficient adaptation to the topological characteristics of the power grid in the prior art.

[0146] To further illustrate the method for preventing misoperation entity alignment in a substation based on multimodal structure fusion described in Embodiment 1 of the present invention, in Embodiment 3, it is combined with Figure 2 for specific implementation and application description.

[0147] Figure 2 Exemplarily shows the entity alignment process in the anti-error logic verification of the method of the present invention, covering the corresponding relationship between method steps and system modules, including:

[0148] Taking the entity alignment of a circuit breaker as an example, cross-system entity alignment is realized through three-stage progressive optimization of initial semantic matching, iterative anchor point screening, and global structure alignment. Specifically, it includes:

[0149] In the initial semantic matching stage, corresponding to the multi-modal data acquisition and feature extraction module in the figure, it is executed to extract the text attributes of the circuit breaker "CB-01" (such as the name "CB-01" and rated current "50kA") from the SCADA system, obtain the topological connection information of "Breaker_A1" (such as connected bus "Bus-A") from the GIS system, and collect the appearance image of "CB-01" (such as nameplate model "VD4-12kV") through the inspection system. Using SimCSE to encode text similarity (such as 0.75), ResNet-50 to compare image features (such as similarity 0.90), and attribute normalization calculation, an initial coupling matrix is generated. (such as 0.82);

[0150] In the iterative anchor point screening stage, corresponding to the verification process of the optimization alignment module in the figure, multi-modal cross-verification is executed - only candidate pairs with text similarity (such as >0.7) and image similarity (such as >0.85) are retained (such as "CB-01" and "Breaker_A1"), and their topological verification is carried out (such as verifying whether both are connected to the bus "Bus-A"), and after the confidence is improved (such as 0.95), they are added to the high-confidence anchor point set.

[0151] In the global structure alignment stage, corresponding to the Gromov-Wasserstein optimization module in the figure, the topological adjacency matrix comparison is executed (such as the structural difference between "CB-01→Bus-A" in SCADA and "Breaker_A1→Bus-A" in GIS is 0), and the objective function f is optimized through the BPG algorithm. FGW (such as 0.175), and finally the coupling matrix is updated to π. CB01,BreakerA1 (such as 0.98);

[0152] Finally, the alignment of the two circuit breaker entities is completed, and unified attributes are output (such as name "Circuit Breaker A1", rated current "50kA", connected bus "Bus-A") for the anti-misoperation logic system to verify operation conflicts (such as triggering an alarm when the disconnect instruction conflicts with the topological state).

[0153] This disclosure can be a system, method, and / or computer program product. The computer program product can include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of this disclosure.

[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent substitutions can still be made to the specific implementation manners of the present invention, and any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A substation anti-error entity alignment method based on multi-modal structure fusion, characterized in that Including the following steps: Perform cross-system multimodal data collection for the power grid, including collecting text data and attribute data from the SCADA system, topological data from the GIS system, and device image data from the inspection system; Preprocess and extract features from the collected power grid multimodal data to generate multimodal features, including text semantic features, image features, attribute features, and topological structure features; Fuse multimodal features to generate a joint embedding, calculate a multimodal similarity matrix, and combine the Sinkhorn algorithm to solve the initial coupling matrix and screen the initial anchor point set; Iteratively update the initial anchor point set by combining multimodal consistency verification to obtain a verified anchor point set; Inject the iterated verified anchor point set as prior knowledge into the Gromov-Wasserstein structure optimization model, calculate the structural similarity of the adjacency matrix in blocks, and achieve global topological alignment.

2. A substation anti-error entity alignment method based on multimodal structure fusion according to claim 1, characterized in that: The text data includes device name, specification description, technical parameters, and operation and maintenance record data; The attribute data includes static attributes, dynamic attributes, and status indicators. The static attributes include commissioning time and installation location. The dynamic attributes include real-time voltage, load rate, and temperature. The status indicators include cumulative operation times and remaining life prediction values; The topological data includes an adjacency matrix, hierarchical relationship, and dynamic topology. The dynamic topology includes the real-time status of circuit breakers and the load flow direction of lines; The image data includes device appearance images, infrared thermal imaging maps, and historical defect images.

3. A substation anti-error entity alignment method based on multimodal structure fusion according to claim 2, characterized in that: The preprocessing of the multimodal data includes: Clean and standardize the text data, fill in missing values and normalize the attribute data, construct a matrix and enhance self-connection for the topological data, and generate and normalize missing virtual images for the image data.

4. A substation anti-error entity alignment method based on multimodal structure fusion according to claim 3, characterized in that: The feature extraction from the preprocessed multimodal data includes: Extract the feature of the preprocessed text data, encode the device name and specification description through the SimCSE model, and extract the device text semantic feature F text , and output a 256-dimensional semantic vector F text ∈R 256 ; Extract the features of the preprocessed image data, and extract the original image features F' through the pre-trained ResNet-50 image ∈R 2048 , and reduce the dimension to 256 dimensions to obtain the image embedding F image ∈R 256 ; Extract the preprocessed attribute data features and generate the attribute embedding F through the MLP attr ∈R 256 ; Extract the topological data features after preprocessing, encode the power grid adjacency matrix through GCN, and output a 256-dimensional structural embedding F struc ∈R 256 .

5. A substation anti-error entity alignment method based on multimodal structure fusion according to claim 4, characterized in that: The fusion of multimodal features to generate a joint embedding includes: Calculate the weights of each modality through the attention mechanism and generate the joint embedding F fused ; Perform joint embedding F fused for component extraction to obtain text components image components attribute components topological components 6. A substation anti-error entity alignment method based on multimodal structure fusion according to claim 5, characterized in that: The calculation of the multimodal similarity matrix includes: According to the text embedding F text 、the image embedding F image 、the attribute embedding f attr and the structure embedding F struc , calculate the text similarity matrix C text 、the image similarity matrix C image 、the attribute similarity matrix C attr and the structure similarity matrix C struc ; According to the text similarity matrix C text , the image similarity matrix C image , the attribute similarity matrix C attr and the structure similarity matrix C struc , calculate the multimodal similarity matrix C sum .

7. A substation anti-error entity alignment method based on multimodal structure fusion according to claim 6, characterized in that: The combination of the Sinkhorn algorithm to solve the initial coupling matrix and screen the initial anchor point set includes: According to the multimodal similarity matrix C sum , the initial coupling matrix π is solved using the Sinkhorn algorithm 0 , and the initial anchor point set is screened c = 1 / max(m, n) where π is the coupling matrix of the alignment probabilities between entities in two knowledge graphs; KL(π||π prior ) is the Kullback-Leibler divergence between the coupling matrix π and the prior distribution π prior ; π ij is the alignment probability between entity e i in the source graph and entity e' j in the target graph; π prior,ij represents the alignment probability between entity e i in the source graph and entity e' j in the target graph in the prior distribution; c represents the maximum possible value of each element in the coupling matrix π; ∈ represents the threshold parameter for screening initial anchor points; m and n are the numbers of entities in the source graph and the target graph respectively; e i represents the i-th entity in the source knowledge graph; e' j represents the j-th entity in the target knowledge graph.

8. A substation anti-error entity alignment method based on multimodal structure fusion according to claim 7, characterized in that: The iterative update of the initial anchor point set by combining multimodal consistency verification to obtain a verified anchor point set includes: Update the anchor set Retain candidate pairs with text similarity and image similarity higher than the threshold: where θ text and θ image represent the text similarity threshold and the image similarity threshold respectively.

9. A substation anti-error entity alignment method based on multi-modal structure fusion according to claim 8, characterized in that: Injecting the iterated verified anchor point set into the Gromov-Wasserstein structure optimization model as prior knowledge includes: The objective function f of the Gromov-Wasserstein structure optimization model FGW is as follows: Among them, α is the balance parameter between semantic alignment and structural alignment; is the adjacency matrix value of entities i and j in the source knowledge graph; is the adjacency matrix value of entities k and l in the target knowledge graph; π ik is the coupling probability of source entity i and target entity k; π jl is the coupling probability of source entity j and target entity l; λ is the weight parameter; is the set of entity pairs; Update π using the BPG algorithm k+1 , and use the Sinkhorn iteration to solve the subproblem until convergence: where f GWD is the Gromov-Wasserstein distance; A s and A t are the adjacency matrices of the source knowledge graph and the target knowledge graph respectively; β is the step size parameter.

10. A substation anti-error entity alignment system based on multi-modal structure fusion, based on the method for aligning anti-error entities of a substation based on multi-modal structure fusion according to any one of claims 1-9, characterized in that, The system includes: A multi-modal data acquisition module for performing cross-system multi-modal data acquisition of the power grid; A multi-modal feature processing module for preprocessing and feature extraction of the acquired power grid multi-modal data to generate multi-modal features; A dynamic fusion and embedding generation module for fusing multi-modal features to generate a joint embedding and calculating a multi-modal similarity matrix; An initial coupling solving module for solving an initial coupling matrix based on the Sinkhorn algorithm and screening an initial anchor point set; An anchor point iterative verification module for iteratively updating the initial anchor point set by combining multi-modal consistency verification to obtain a verified anchor point set; A structure optimization and alignment module injects the iterated verified anchor point set into the Gromov-Wasserstein structure optimization model as prior knowledge, calculates the structural similarity of the adjacency matrix in blocks, and realizes global topological alignment.

Citation Information

Cited By

  • GW distance-based graph alignment method, apparatus and system, and storage medium

    CN122388699A