A data governance method and system for power spatial data
Through deep learning model and semantic mapping technology, the problem of insufficient logical association and contextual relationship extraction in power space data governance is solved, efficient and accurate data repair and governance is achieved, and the data processing capabilities of the power system are improved.
Patent Information
- Application Number
- CN202510368858.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-03-27
AI Technical Summary
In the governance of power space data, it is difficult for the prior art to effectively extract the logical relationship and contextual relationship between data, and the detection and repair of abnormal data, redundant data and conflict data are insufficient, making it difficult to meet the needs of efficient and accurate governance.
Deep learning model is used to jointly model the spatial dimensions, time dimensions and business logic dimensions of power spatial data, generate semantic models of spatiotemporal characteristics and business logic constraints, extract multi-level semantic features through embedding and attention mechanisms, build semantic maps and global knowledge bases, and combine data cleaning algorithms to detect and repair abnormal, redundant and conflicting data.
It significantly improves the processing capability of complex multi-source heterogeneous data, realizes efficient detection and precise repair, improves the relevance and availability of data governance, and supports grid operation optimization and precise management.
Smart Images

Figure CN119884610B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power spatial data processing, and a data governance method and system for power spatial data. Background Art
[0002] Power spatial data is an important basic data for power grid operation and management, covering multi-dimensional information such as equipment status, geographical location, time series, and business logic. In the prior art, the processing methods for power spatial data mainly focus on traditional data cleaning and rule matching methods, and data governance is achieved through statistical analysis or some artificial intelligence-based algorithms. These technologies have played a certain role in the efficiency and accuracy of data governance and have been widely applied in power system management.
[0003] However, with the complexity of the power system and the rapid growth of data volume, the prior art faces many limitations in governing multi-source heterogeneous data and in-depth semantic analysis. Traditional methods often lack an in-depth understanding of data semantics and are difficult to effectively extract the logical associations and context relationships between data. In addition, the detection and repair of abnormal data, redundant data, and conflicting data rely on static rules, with insufficient flexibility, and the processing results are also difficult to meet the requirements of efficient and accurate governance.
[0004] In view of the above problems, the present invention proposes an innovative method to improve the governance effect of power spatial data. Summary of the Invention
[0005] The present application provides a data governance method and system for power spatial data to improve the governance effect of power spatial data.
[0006] The present application provides a data governance method for power spatial data, including:
[0007] Based on the spatio-temporal characteristics of power spatial data and power business logic, a deep learning model is used to jointly model the spatial dimension, time dimension, and business logic dimension of power spatial data to generate a semantic model including spatio-temporal characteristics and business logic constraints;
[0008] Using the constructed semantic model, feature extraction is performed on power spatial data through an embedding and attention mechanism to generate multi-level semantic features, where the multi-level semantic features include the spatial correlation, time series relationship, and business logic constraints of power spatial data;
[0009] Using the extracted multi-level semantic features, a semantic graph is constructed through graph database technology; based on the semantic graph, a global power spatial data knowledge base is generated, and the logical associations and context relationships between power spatial data are dynamically annotated in the knowledge base;
[0010] Based on the global power spatial data knowledge base, combined with data cleaning algorithms, detect abnormal data, redundant data, and conflicting data in the power spatial data, and generate preliminary repair suggestions;
[0011] According to the preliminary repair suggestions, based on the logical rules generated by the semantic model, perform semantic consistency verification on the repair contents of abnormal data, redundant data, and conflicting data;
[0012] Based on the verification results of the semantic consistency verification, repair the abnormal data, redundant data, and conflicting data to generate a processed power spatial data set.
[0013] Furthermore, based on the spatio-temporal characteristics of the power spatial data and power business logic, use a deep learning model to jointly model the spatial dimension, time dimension, and business logic dimension of the power spatial data, and generate a semantic model with spatio-temporal characteristics and business logic constraints, including:
[0014] Obtain multi-source heterogeneous input data of the power spatial data, where the multi-source heterogeneous input data includes the geographical location, operating status, historical time series data of power equipment, and corresponding business logic rules; perform hierarchical preprocessing on the multi-source heterogeneous input data to form spatial feature vectors, time series feature vectors, and logical rule feature vectors;
[0015] Based on a graph neural network, model the spatial feature vectors to generate spatial semantic features including the physical connection relationship and spatial correlation between devices; based on a long short-term memory network, model the time series feature vectors to generate time semantic features reflecting the historical operating status and time series relationship of devices; based on an attention mechanism combined with rule reasoning, model the logical rule feature vectors to generate business semantic features describing device operation constraints and event logical dependencies;
[0016] Fuse the spatial semantic features, time semantic features, and business semantic features through a multi-layer fully connected network to generate a semantic model including spatio-temporal characteristics and business logic constraints.
[0017] Furthermore, the hierarchical preprocessing of the multi-source heterogeneous input data includes:
[0018] According to the following formula (1), calculate the weighted distance:
[0019] ;
[0020] Where represents the weighted distance between device and device ; is device and device The geographical distance between; is the operating load of the device ; is the operating load of the device ; is the power transmission capacity between the device and the device ; are respectively the health states of the device and the device , where the value range of the health state is , where 1 represents completely healthy; is the weight coefficient; is the smoothing parameter to avoid division by zero;
[0021] Based on the dynamic weighted distance , the weighted mean clustering algorithm is used to group power devices, where the weighted mean clustering algorithm's optimization objective function adopts the following formula 2:
[0022] ;
[0023] Among them, is the target number of clusters; is the th cluster; represents the number of devices in the th cluster; is the total number of power devices; is the cluster balance penalty coefficient;
[0024] According to the following formula (3), the spatial feature vector is generated for the th device:
[0025] ;
[0026] Among them, is the geometric center coordinate of the th device's cluster ; is the weighted distance calculated according to formula 1 between the th device and the center of its belonging cluster ; represents the number of devices in the th cluster; is the th device's weighted distance from all devices within its cluster .
[0027] Furthermore, modeling the spatial feature vectors based on a graph neural network to generate spatial semantic features including the physical connection relationships and spatial correlations between devices includes:
[0028] Obtain a dynamic graph representing the topological structure between power devices, where the nodes of the dynamic graph represent power devices, the edges represent the physical connection relationships between devices, and the weights of the edges are dynamically calculated from the physical distance, transmission capacity, and connection status between devices. The weights of the edges are calculated using the following formula (4):
[0029] ;
[0030] Wherein, represents the weight of the edge between device and device at time instant ; is the physical distance between nodes and ; represents the connection status of device and device at time , with a range of , where 1 represents full connection and 0 represents no connection; is the transmission capacity between nodes and ; , and are weight coefficients; is a smoothing coefficient;
[0031] Normalize the weights of the edges according to the following formula (5);
[0032] ;
[0033] Wherein, represents the normalized edge weight between device and device at time instant ; represents the set of all nodes connected to node ;
[0034] Based on the normalized edge weights, perform spatial feature modeling using a dynamic graph neural network, where the update formula for the nodes represented in each layer of the graph neural network is expressed by the following formula (6):
[0035] ;
[0036] Among them, represents the feature representation of node at the layer; is the trainable weight matrix of the layer; represents the feature representation of node at the layer; is the activation function; represents the set of all nodes connected to node ;
[0037] The spatial features are updated layer by layer through a dynamic graph neural network to generate spatial semantic features reflecting the spatio-temporal correlation between power equipment.
[0038] Furthermore, the logical rule feature vector is modeled based on the attention mechanism combined with rule reasoning to generate business semantic features describing equipment operation constraints and event logical dependencies, including:
[0039] Extract the logical rule set related to the operation and scheduling of power equipment, and generate its feature representation for each rule in the rule set. The feature representation includes the constraint conditions, applicable scope of the rule, and the dependency relationship with other rules;
[0040] Use a rule-based inference engine to analyze the logical rule set and construct a dependency graph between rules. The nodes of the dependency graph represent rules, and the edges represent the logical dependency strength between rules;
[0041] Combine the self-attention mechanism to model the rule nodes in the rule dependency graph, and calculate the priority weight of each rule according to the dependency strength and constraint scope of the rule;
[0042] According to the priority weights, perform weighted aggregation on the logical rule feature vector to generate business semantic features describing equipment operation constraints and event logical dependencies.
[0043] The beneficial effects of the technical solution provided by this application include:
[0044] (1) Through joint modeling and semantic model construction based on deep learning models, the spatio-temporal characteristics and business logic constraints of power spatial data can be deeply analyzed, multi-level semantic features can be extracted, and the processing ability for complex, multi-source heterogeneous data can be significantly improved. (2) By using the semantic model and semantic consistency verification mechanism, abnormal data, redundant data, and conflict data can be efficiently detected and accurately repaired, ensuring the logical consistency and semantic integrity of the repaired content. (3) Through the construction of semantic graphs and the generation of global knowledge bases, dynamic annotation of the logical associations and contextual relationships between power spatial data can be realized, comprehensively enhancing the relevance and usability of power data governance. (4) Through the construction of semantic graphs and global knowledge bases, semantic support can be provided for subsequent power grid operation optimization, precise management, and user queries, greatly improving the efficiency and intelligence of data query and analysis. Description of the Drawings
[0045] Figure 1 is a flowchart of a data governance method for power spatial data provided by the first embodiment of the present application.
[0046] Figure 2 is a schematic diagram of a data governance system for power spatial data provided by the second embodiment of the present application. Detailed Embodiments
[0047] Many specific details are set forth in the following description in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.
[0048] The first embodiment of the present application provides a data governance method for power spatial data. Please refer to Figure 1 , which is a schematic diagram of the first embodiment of the present application. The following will describe in detail a data governance method for power spatial data provided by the first embodiment of the present application in combination with Figure 1 .
[0049] Step S101: Based on the spatio-temporal characteristics of power spatial data and power business logic, a deep learning model is used to jointly model the spatial dimension, time dimension, and business logic dimension of power spatial data, and a semantic model of spatio-temporal characteristics and business logic constraints is generated.
[0050] Step S101 involves jointly modeling based on the spatio-temporal characteristics of power spatial data and power business logic through a deep learning model to generate a semantic model that can describe data characteristics and business constraints. Specifically, this step includes the following implementation manners:
[0051] First, preprocess the original power spatial data to meet the input requirements of the deep learning model. The preprocessing includes but is not limited to data cleaning, format standardization, and feature encoding. The cleaning operation is used to remove obvious noisy data or incomplete data; format standardization can transform spatial positions, time series, and related power business events into a unified input structure; feature encoding can convert unstructured data (such as power equipment types or business rules) into vector forms that can be input into the model through embedding methods.
[0052] After the data preprocessing is completed, analyze each dimension of the power spatial data. The analysis of the spatial dimension data includes converting geographical location information into a coordinate system or topological structure to represent the spatial relationship between power equipment. The analysis of the time dimension data includes sampling and segmenting time series data. For example, organize data on power load or equipment operating status by time window. The data in the business logic dimension needs to extract key events or operation sequences in combination with specific power business rules and equipment operation logics.
[0053] Next, design and train a deep learning model. The model consists of three parts: a feature extraction module in the spatial dimension, a feature extraction module in the time dimension, and a feature extraction module in the business logic dimension. In the spatial dimension, a graph neural network can be used to capture the spatial correlation between power equipment; in the time dimension, a long short-term memory network (LSTM) or a Transformer structure can be used to capture the time series pattern of the data; in the business logic dimension, a model based on the attention mechanism can be used to model the logical constraints between events. The output features of the above modules are comprehensively modeled through a fully connected network or other fusion mechanisms to form a joint representation.
[0054] The training of the model is based on the labeled training dataset and is carried out in a supervised learning, semi-supervised learning, or unsupervised learning manner. During the training process, multiple loss functions can be used, such as the mean squared error (MSE) loss for time prediction, the categorical cross-entropy loss for logical constraints, etc. After training, the model can generate a semantic model that describes the spatio-temporal characteristics and business logic constraints of the power spatial data.
[0055] The generation of the semantic model mainly includes the fixation of the model's parameters and function encapsulation. The generated semantic model can be directly used for subsequent tasks such as feature extraction. The model structure and parameters can be stored in a file format that can be called (such as HDF5 or ONNX). As the output, the semantic model can be used to describe and quantify the multi-dimensional characteristics of the power spatial data and provide a basis for subsequent steps. Through this step, the complex characteristics of the power spatial data are systematized and modeled, facilitating subsequent processing and application.
[0056] Furthermore, based on the spatio-temporal characteristics and power business logic of power spatial data, a deep learning model is used to jointly model the spatial dimension, time dimension, and business logic dimension of power spatial data, generating a semantic model with spatio-temporal characteristics and business logic constraints, including:
[0057] Obtain multi-source heterogeneous input data of power spatial data, where the multi-source heterogeneous input data includes the geographical location, operating status, historical time series data of power equipment, and corresponding business logic rules; perform hierarchical preprocessing on the multi-source heterogeneous input data to form spatial feature vectors, time series feature vectors, and logical rule feature vectors;
[0058] Based on a graph neural network, model the spatial feature vectors to generate spatial semantic features containing the physical connection relationships and spatial correlations between devices; based on a long short-term memory network, model the time series feature vectors to generate time semantic features reflecting the historical operating status and time series relationships of devices; based on an attention mechanism combined with rule reasoning, model the logical rule feature vectors to generate business semantic features describing device operation constraints and event logical dependencies;
[0059] Fuse the spatial semantic features, time semantic features, and business semantic features through a multi-layer fully connected network to generate a semantic model including spatio-temporal characteristics and business logic constraints.
[0060] First, it is necessary to obtain multi-source heterogeneous input data. These data may include the geographical location information of power equipment, the real-time operating status of the equipment, historical time series records, and business logic rules related to the operation of the power grid. The geographical location information can be obtained through the geographical coordinates or topological connection data of the equipment. The operating status may include real-time parameters such as voltage, current, and load. The time series data records the changes in the operating status of the equipment over time, while the business logic rules reflect the operating specifications and dispatching strategies of the power system.
[0061] After obtaining the data, it is necessary to perform hierarchical preprocessing on the multi-source heterogeneous data to form a feature representation suitable for model processing. The spatial dimension data is transformed into spatial feature vectors after encoding and normalization. For example, the geographical location information can be mapped into coordinate points or topological matrices. The time dimension data is processed through time window segmentation and serialization to generate time series feature vectors, which capture the regularity of the changes in the operating status of the equipment over time. The business logic data is transformed into logical rule feature vectors through rule parsing and embedding techniques, retaining the causal relationships and operation constraints between events.
[0062] In the modeling stage, the spatial feature vectors are processed by a graph neural network. The graph neural network models based on the physical connection relationships and topological structures of the devices, extracts spatial semantic features from them, and these features reflect the spatial correlations between power grid devices, such as direct connection or adjacency relationships between devices. The time series feature vectors are modeled by a long short-term memory network (LSTM). LSTM is good at processing data with strong time dependencies and can capture the historical patterns and future trends of the device operating states. For the business logic feature vectors, the attention mechanism is combined with rule reasoning. The attention mechanism dynamically assigns importance weights to the rules, while rule reasoning ensures the derivation and constraint consistency of the logical features, thereby generating business semantic features.
[0063] After completing the modeling, the spatial semantic features, time semantic features, and business semantic features are fused through a multi-layer fully connected network. In the fully connected network, each type of feature is weighted to obtain a comprehensive representation, and the fusion process highlights the interaction relationships and global significance among the features. The fused features are further subjected to non-linear activation and dimensional compression, and finally a semantic model is generated. This model simultaneously contains the spatio-temporal characteristics of the power data and business logic constraints.
[0064] This method fully captures the complex characteristics of power spatial data through the joint modeling of deep learning models, providing a reliable semantic basis for subsequent data governance, anomaly detection, and decision support.
[0065] Furthermore, the hierarchical preprocessing of the multi-source heterogeneous input data includes:
[0066] Calculate the weighted distance according to the following formula (1):
[0067] ;
[0068] where represents the weighted distance between device and device ; is the geographical distance between device and device , which can be calculated from the geographical coordinates of the devices.
[0069] is the operating load of device ; is the operating load of device ; is the power transmission capacity between device and device ;
[0070] are respectively the and the device of the health status, where the health status value range is , where 1 represents completely healthy; the health status of the device is determined by monitoring the key performance indicators of the device, and these indicators usually include the operation history of the device, maintenance records, the number of failures, electrical parameters (such as the stability of voltage and current), etc., and are obtained through comprehensive evaluation. The value of the health status is between 0 and 1, where:
[0071] 1 indicates that the device is completely healthy: the device is in good operating condition, there are no obvious abnormalities in the historical data, and all key performance indicators are within the normal range. For example, the device has never had a failure or damage, the electrical parameters remain stable, and the regular maintenance is completed on time.
[0072] 0 indicates that the device is completely unhealthy: the device has serious problems or cannot operate, such as continuous failures, abnormal electrical parameters, or damage to key components.
[0073] Values between 0 and 1: reflect the health degree of the device. For example, if the device occasionally shows fluctuations during operation, or there is slight aging but it does not affect the function, a health status of 0.8 or 0.9 can be given; if the device has experienced several failure repairs but is still operating, it can be evaluated as 0.5 or lower.
[0074] is the weight coefficient, obtained through experiments or set according to expert knowledge;
[0075] is the smoothing parameter to avoid division by zero; based on the dynamic weighted distance , the weighted mean clustering algorithm is used to group power equipment, where the weighted mean clustering algorithm's optimization objective function uses the following formula 2:
[0076] ;
[0077] where, is the target number of clusters; is the th cluster; represents the th device in the th cluster; is the total number of power equipment;
[0078] According to the following formula (3), the spatial feature vector of the th device is generated :
[0079] ;
[0080] wherein, is the geometric center coordinate of the cluster where the th device is located; is the weighted distance between the th device and the center of the cluster to which it belongs, calculated according to Formula 1; represents the number of devices in the th cluster; is the average weighted distance between the th device and all devices within the cluster where it is located.
[0081] First, by obtaining the basic attributes of power devices, such as geographical location, operating load, power transmission capacity, and health status, these attributes are used to calculate the weighted distance between devices. The weighted distance not only considers physical attributes (such as geographical distance and transmission capacity) but also integrates operating characteristics (such as load differences and health status), providing a comprehensive measure of device correlation.
[0082] On this basis, the weighted k - means clustering algorithm is used to group the devices. The optimization objective of the algorithm ensures that the grouping result can reflect the close association of devices while avoiding the situation where a single cluster is too sparse or dense by balancing the compactness within the cluster and the balance between clusters.
[0083] Finally, a spatial feature vector is generated for each device. The spatial feature vector consists of geometric center coordinates, weighted distance, and average weighted distance, comprehensively describing the spatial position of the device and its relationship with the cluster center and other devices. The design of this feature vector provides high - quality input data for subsequent modeling and analysis.
[0084] Furthermore, the spatial feature vector is modeled based on a graph neural network to generate spatial semantic features including the physical connection relationship and spatial correlation between devices, including:
[0085] Obtain a dynamic graph representing the topological structure between power devices, where the nodes of the dynamic graph represent power devices, the edges represent the physical connection relationships between devices, and the weights of the edges are dynamically calculated from the physical distance, transmission capacity, and connection status between devices. The weights of the edges are calculated using the following formula (4):
[0086] ;
[0087] wherein, represents device and device at time instance the weight of the edge; this value is calculated by Formula 4 and is a dynamically updated quantity used to measure the real-time association degree between two devices. A larger weight indicates a more important relationship between the two devices.
[0088] is the physical distance between nodes and ;
[0089] represents the connection status of device at time instance and device , with a range of , where 1 indicates full connection and 0 indicates no connection; the connection status is a value between 0 and 1 used to represent the connectivity degree between devices. It can be obtained by monitoring the operating status of the devices. For example, if two devices are fully connected (such as a circuit breaker being closed), then = 1; if they are completely disconnected (such as a circuit breaker being open), then = 0. When the connection quality is between the two, a reasonable value can be given according to the operating conditions of the devices, such as the current transmission performance or stability. For example, 0.5 indicates partial connection or medium connection quality.
[0090] is the transmission capacity between nodes and ;
[0091] , and are weight coefficients, which can be determined through experiments or optimization methods. For example, if the physical distance has a small impact on the device association, can be set to a lower value, while when more dependent on the transmission capacity, can be increased.
[0092] is the smoothing coefficient, which is used to avoid the case of a zero denominator and is usually set to a very small positive number. Its role is to improve the stability of the calculation.
[0093] Normalize the weight of the edge according to the following formula (5);
[0094] ;
[0095] where, represents device and device The normalized edge weight at time instant ; denotes the set of all nodes connected to node ;
[0096] Equation (5) is used to normalize the edge weights and calculate the normalized edge weight . The normalization is achieved by dividing the edge weight between device and device by the sum of all adjacent edge weights of node . This normalization operation ensures that the sum of the weights of all adjacent edges is 1, thus distributing reasonable weight ratios in the graph neural network calculation. For example, if node is connected to 3 nodes with edge weights respectively, after normalization, they still retain the relative ratio but the sum is 1.
[0097] Based on the normalized edge weights, a dynamic graph neural network is used for spatial feature modeling, where the update formula for the node representation in each layer of the graph neural network is expressed by the following Equation (6):
[0098] ;
[0099] where, denotes the feature representation of node at the -th layer; is the trainable weight matrix at the -th layer; denotes the feature representation of node at the -th layer; is the activation function; denotes the set of all nodes connected to node ;
[0100] Equation (6) is the node update formula of the dynamic graph neural network, which is used to update the node features layer by layer. At the -th layer, the feature representation of node is the weighted sum of the features of all its adjacent nodes, weighted by the normalized edge weight . The weight matrix is a trainable parameter used to learn the linear transformation of node features. The activation function (such as ReLU or Leaky ReLU) introduces non-linearity to enhance the expressive power of the network.
[0101] The spatial features are updated layer by layer through a dynamic graph neural network to generate spatial semantic features reflecting the spatio-temporal correlation between power equipment.
[0102] During the entire modeling process, the dynamic graph neural network gradually captures the spatial correlation and dynamic spatio-temporal characteristics between power equipment through multi-layer node updates. For example, in a transmission network, if transformer A is physically close to circuit breaker B and is currently in the closed state, its edge weight will be higher, and the features of A and B will have a greater mutual influence during the update process.
[0103] Step S102: Using the constructed semantic model, extract features from power spatial data through an embedding and attention mechanism to generate multi-level semantic features, where the multi-level semantic features include the spatial correlation, time series relationship, and business logic constraints of the power spatial data.
[0104] The implementation of step S102 involves using the semantic model generated in step S101 to extract features from power spatial data through an embedding and attention mechanism to generate multi-level semantic features including spatial correlation, time series relationship, and business logic constraints. The specific implementation is as follows.
[0105] First, based on the generated semantic model, initialize the input data for feature extraction. The input of power spatial data includes the feature vectors encoded in its spatial dimension, time dimension, and business logic dimension, and these feature vectors are further embedded based on the semantic model. The data embedding in the spatial dimension can adopt a spatial embedding algorithm based on the position relationship. For example, the geographical location is mapped into a vector form to quantify the spatial correlation between devices; the data embedding in the time dimension is processed through a time series feature processing module, such as time window encoding or timestamp embedding method, to map the data into the time feature space; the data embedding in the business logic dimension utilizes business rules and event sequences to transform them into structured features that can be input into the deep learning model.
[0106] After the preliminary embedding is completed, feature extraction relies on the attention mechanism to capture the multi-dimensional correlation of the embedded vectors. Specifically, the attention module of the semantic model is used to weight the input embedded data, thereby extracting the important interaction characteristics between the data. The attention mechanism in the spatial dimension can dynamically allocate weights between each power equipment, highlighting the pairs of equipment with strong correlation. The attention mechanism in the time dimension analyzes the correlation between different time points in the time series data to capture the key time-dependent features. The attention mechanism in the business logic dimension dynamically allocates weights according to the logical relationship of events, thereby extracting the important patterns in the business logic constraints.
[0107] Through the embedding and attention mechanisms, the obtained feature vectors are input into the feature fusion layer of the deep learning model. The fusion layer combines features of different dimensions into a unified multi-level semantic feature representation through non-linear activation functions (such as ReLU or Leaky ReLU) and fully connected layers. This multi-level semantic feature contains spatial correlation, time series relationship, and business logic constraints, and can comprehensively reflect the inherent characteristics and complex relationships of power spatial data.
[0108] After feature extraction is completed, the multi-level semantic features are further formatted for subsequent steps. The formatting process includes normalization and dimension compression of the feature vectors to optimize storage and computational efficiency. These feature data can be stored in the form of matrices or tensors as the basis for the next step of semantic graph construction. Through this series of operations, step S102 lays a solid semantic foundation for the deep governance of power spatial data and provides high-quality input for subsequent data analysis and processing.
[0109] Furthermore, the use of the constructed semantic model to extract features from power spatial data through the embedding and attention mechanisms to generate multi-level semantic features includes:
[0110] Perform embedding processing on the spatial dimension of power spatial data, and convert the geographical location information, connection relationship, and physical topology characteristics of devices into high-dimensional feature representations that can be processed by the model;
[0111] Perform sequential embedding on the time dimension of power spatial data, and map the historical records and time-dependent relationships of device operating states into continuous time feature representations to capture the impact of key time points;
[0112] Perform rule embedding processing on the business logic dimension of power spatial data, and convert the power business rules and logical dependencies of events into logical feature vectors to retain the causal relationships between events;
[0113] Perform weighted processing on the above embedding features through the attention mechanism, dynamically allocate the weights of each feature, and extract key features according to the importance of spatial correlation, time series patterns, and business rules, generating multi-level semantic features including spatial correlation, time series relationship, and business logic constraints.
[0114] First, it is necessary to perform embedding processing on each dimension of the power spatial data to generate a high-dimensional feature representation suitable for model calculation. For the data in the spatial dimension, the key point of the embedding processing is to encode the geographical location information of the equipment, the physical connection relationship between power equipment, and the grid topology characteristics into high-dimensional features available to the model. Specifically, the geographical location of the equipment can be transformed into a coordinate vector, and the connection relationship between the equipment can be represented by an adjacency matrix or a topological graph. Physical topology characteristics, such as the degree and connectivity of nodes, can be further transformed into embedding features using specific encoding techniques, so as to provide the model with the association information between equipment.
[0115] The embedding processing in the time dimension focuses on the time series information of the equipment operation status. The historical record data of the equipment, such as load, voltage, or fault status, can be serialized into time series features and segmented using time window techniques. These segmented data are then mapped to continuous time features to capture the dependencies and trends between time points. For example, through timestamp embedding or position encoding methods, the model can be enabled to recognize the impact of key time points on data patterns.
[0116] For the data in the business logic dimension, the core of the embedding processing is to encode the power business rules and the logical dependencies of events into logical feature vectors. Business rules may include scheduling processes, fault handling rules, or maintenance plans, which are transformed into logical expressions through semantic parsing techniques and further mapped into feature vectors. The logical dependencies of events, such as causal chains or priority orders, are integrated into the logical feature vectors through rule reasoning to ensure that the model can recognize the causal relationships and constraint conditions between events.
[0117] After the above embedding processing, the generated features are weighted through the attention mechanism. The role of the attention mechanism is to dynamically adjust the weights of each feature, allocating resources according to its relative importance to the current task or goal. For example, for some tasks, spatial features may be more important than time features, while for other tasks, logical rules may play a key role. The attention mechanism assigns appropriate weights to each feature by learning the correlations between features, thereby highlighting the contributions of key features to the model.
[0118] Finally, multi-level semantic features are generated through the embedding and attention mechanisms, including spatial correlation, time series relationship, and business logic constraints. These features comprehensively represent the multi-dimensional information of the power spatial data, providing high-quality input support for subsequent semantic modeling, analysis, and data governance. This method ensures that the complex characteristics of the power spatial data in different dimensions can be fully captured and effectively utilized.
[0119] Step S103: Construct a semantic graph using the extracted multi-level semantic features through graph database technology; based on the semantic graph, generate a global power spatial data knowledge base, and dynamically annotate the logical associations and contextual relationships between power spatial data in the knowledge base.
[0120] Step S103 involves using the multi-level semantic features extracted in Step S102 to construct a semantic graph through graph database technology, generating a global power spatial data knowledge base based on the semantic graph, and simultaneously dynamically annotating the logical associations and contextual relationships between power spatial data in the knowledge base. The specific implementation is as follows:
[0121] First, the extracted multi-level semantic features need to be formatted to meet the input requirements of the graph database. This includes hierarchically organizing the feature vectors, clarifying the node types of the data (such as power equipment, geographical location, time nodes, etc.) and the types of edges between them (such as physical connection relationships between devices, time series associations of events, business logic constraints, etc.). These nodes and edges are labeled with unique identifiers to ensure their uniqueness and identifiability when constructing the graph database.
[0122] Next, use graph database technology to construct a semantic graph. Graph database technology can be implemented using existing graph database platforms (such as Neo4j, JanusGraph, etc.). The construction steps include the definition of nodes and edges, the attachment of attributes, and the creation of relationships. In the node definition, specific attributes are assigned to each node type. For example, power equipment nodes contain attributes such as device ID, device type, geographical location, etc., and event nodes contain attributes such as event type, timestamp, related devices, etc. The definition of edges is based on the relationship types between nodes, such as "connected to", "occurred at", or "controlled by", etc. Each edge type can be attached with weight or directionality information to reflect the strong or weak associations or causal relationships between data.
[0123] During the construction process of the semantic graph, it is necessary to dynamically map the multi-level semantic features into specific elements in the graph structure. For example, spatial correlation features are used to define the spatial adjacency relationships between device nodes, time series relationships are used to connect time nodes with related events, and business logic constraints are used to establish logical connections between business rule nodes and device nodes or event nodes. This mapping process can be achieved through scripts or automated tools to ensure that all data elements are correctly reflected in the graph structure.
[0124] After the preliminary construction of the semantic graph, a global power spatial data knowledge base is generated based on the graph. The generation of the knowledge base is based on the semantic graph, and further integrates external data sources (such as power grid topology data, historical operation data, business rule documents, etc.) to supplement and improve the existing graph information. During the process of generating the knowledge base, logical rules and context relationship rules can be defined to dynamically annotate the logical associations and context information between data. For example, the rule engine is used to dynamically annotate the current flow relationship according to the electrical characteristics between devices, or to dynamically generate the context description of the event chain according to the time series of event occurrences.
[0125] The dynamic annotation mechanism of the knowledge base is realized through real-time update and the inference engine. For example, when new data is input or existing data changes, the knowledge base can automatically update the attributes or relationships of relevant nodes and edges, so as to ensure that the logical associations and context relationships of the data are always kept up-to-date. This dynamic annotation mechanism relies on the real-time computing ability of the graph database and the logical rule support provided by the semantic model.
[0126] Through the above steps, the construction of the semantic graph and the generation of the global power spatial data knowledge base provide comprehensive, accurate and efficient support for subsequent data analysis and processing. The method makes full use of the advantages of multi-level semantic features and graph database technology, and realizes the in-depth semantic parsing and structured organization of power spatial data.
[0127] Furthermore, constructing the semantic graph by using the extracted multi-level semantic features through graph database technology includes:
[0128] Dividing the extracted multi-level semantic features into feature sets of nodes and edges, where the node features include the spatial location, time state and business logic attributes of power equipment, and the edge features include the physical connection relationship between devices, the time-dependent relationship between events, and the logical rule constraint relationship;
[0129] Defining the graph structure based on the node features and edge features, where the nodes represent power equipment or key events, and the edges represent the association relationships between devices or the logical dependencies between events;
[0130] Creating a storage structure of nodes and edges in the graph database, and attaching attributes to each node and edge, including spatial characteristics, time characteristics and context information of business rules;
[0131] Dynamically updating the nodes and edges in the graph structure according to the spatio-temporal changes of the power spatial data, and ensuring that the semantic graph can reflect the changes of device states and the adjustment of business logic in real time.
[0132] First, it is necessary to systematically partition the extracted multi-level semantic features. The multi-level semantic features include three main dimensions: space, time, and logic. The node feature set mainly covers the core attributes related to power equipment, such as the geographical location of the equipment, electrical parameters, real-time status, and business logic constraints. These pieces of information together describe the spatial location, temporal state, and operational attributes of the equipment. The edge feature set includes the physical connection relationships between devices, the temporal dependency relationships between events, and the logical rule constraints. These edge features are used to describe the correlations and logical interactions between nodes.
[0133] Based on the partitioned node and edge features, define the graph structure for the semantic graph. The nodes of the graph structure can represent power equipment or key events, such as device nodes like transformers, switches, circuit breakers, or event nodes like system scheduling and abnormal alarms. The edges represent the relationships between the nodes, such as the power transmission paths in the physical topology, the temporal sequential dependencies in the event trigger chain, and the logical associations in the business rule constraints. In this way, the graph structure can comprehensively express the multi-dimensional correlations of power spatial data.
[0134] In the graph database, create the storage forms of nodes and edges according to the defined graph structure. Each node stores its attribute information, such as geographical coordinates, device category, and real-time operating status; each edge stores its associated context information, such as the physical distance between devices, transmission capacity, and the temporal trigger interval and logical dependency rules between events. These attributes and context information are attached to the nodes and edges, making each element have a complete semantic description that can be traced and queried.
[0135] To ensure that the semantic graph can adapt to the dynamic changes of power spatial data, it is necessary to update the nodes and edges in the graph structure in real time. For example, when the status of a device changes (such as from an operating state to a faulty state), the corresponding node attributes need to be updated immediately; when the physical topology is adjusted or new event chains are added, the corresponding edge attributes or connection relationships need to be modified synchronously. Through these update operations, the semantic graph can remain consistent with the actual power spatial data and reflect the latest information in a timely manner when the data changes dynamically.
[0136] Finally, through this semantic graph constructed based on graph database technology, it can provide an efficient structured representation for power spatial data. The semantic graph not only contains multi-dimensional information about power equipment and events but also can reveal complex logical associations and context dependencies through the relationships between nodes and edges, providing a powerful support tool for power data governance, analysis, and decision-making.
[0137] Furthermore, based on the semantic graph, generate a global knowledge base for power spatial data, including:
[0138] Extract the semantic information of nodes and edges from the semantic graph, where the semantic information of nodes includes the geographical location, power operation parameters, and real-time status of devices, and the semantic information of edges includes the physical connections between devices, the chronological relationship between events, and the association strength;
[0139] Based on the extracted node and edge information, define the core structure of the knowledge base, map each node to an entity in the knowledge base, map the edges to the association relationships between entities, and attach context attributes related to device operation, time logic, and business rules;
[0140] Store the data in the semantic graph into the knowledge base, adopt an extensible storage form to support large-scale data operations, and dynamically maintain the update mechanism of nodes and edges to ensure that the knowledge base can synchronize the changes in the semantic graph in real time.
[0141] First, it is necessary to extract the complete semantic information of nodes and edges from the constructed semantic graph. The semantic information of nodes should cover the geographical location, power operation parameters, and real-time status of power devices. For example, each node can represent a specific power device, including its location in the power grid topology, current load, voltage level, and operation health status, etc. The semantic information of edges is used to describe the relationships between nodes, including the physical connection characteristics between devices, such as the transmission capacity and length of power lines, and the chronological relationship and logical association strength between events, such as the triggering dependency of abnormal events.
[0142] Based on the extracted node and edge information, define the core structure of the global power space data knowledge base. The entities in the knowledge base correspond to the nodes in the semantic graph, such as transformers, switches, circuit breakers, or specific scheduling events. The association relationships between entities are mapped from the edges in the semantic graph, such as the power transmission path between devices or the triggering logic between events. Each entity and association relationship is attached with rich context attributes, including spatial characteristics (such as the geographical coordinates of devices), time logic (such as the chronological order of event occurrence), and business rules (such as scheduling dependencies or fault impact ranges). These context attributes ensure that the knowledge base can comprehensively represent the multi-dimensional correlations of power space data.
[0143] When the data in the semantic graph is stored in the knowledge base, an extensible storage form needs to be adopted to support large-scale power data operations. For example, the knowledge base can be implemented based on a relational database or a graph database to ensure that, on the basis of storing nodes and edges, complex association relationships can also be quickly queried. To adapt to the dynamic changes of actual power spatial data, the information of nodes and edges in the knowledge base needs to have a dynamic maintenance mechanism. For example, when the operating state of a certain device changes, the attributes of the corresponding entity in the knowledge base should be updated in real time; when the power grid topology is adjusted, the newly added or removed devices and their connection relationships should be synchronously reflected in the knowledge base. This dynamic update mechanism ensures that the content of the knowledge base is always consistent with the semantic graph.
[0144] The finally generated global power spatial data knowledge base provides comprehensive semantic data support and can be queried, inferred, and analyzed through standardized interfaces. The knowledge base not only reflects the current state of power devices and events but also reveals the complex logic and associations of power spatial data through its structured semantic representation, providing a reliable basis for subsequent power data governance, dispatching optimization, and decision-making support. Through this method, the global power spatial data knowledge base can manage and utilize power spatial data in an efficient, dynamic, and extensible manner.
[0145] Step S104: Based on the global power spatial data knowledge base, combined with a data cleaning algorithm, detect abnormal data, redundant data, and conflicting data in the power spatial data, and generate preliminary repair suggestions.
[0146] The implementation of step S104 involves detecting abnormal data, redundant data, and conflicting data in the power spatial data based on the global power spatial data knowledge base and combining a data cleaning algorithm, and generating preliminary repair suggestions. Specifically, in this step, the target data set to be cleaned is first extracted from the global power spatial data knowledge base, which usually includes all power data related to the spatial dimension, time dimension, and business logic dimension. These data will be classified and sorted according to their node types (such as device nodes, event nodes) and edge types (such as association relationships, time series relationships) for subsequent processing.
[0147] For the sorted data, the target data set is processed in combination with data cleaning algorithms. Anomaly data detection is mainly achieved by setting context constraints and logical rules. For example, for power equipment operation data, outliers can be identified by monitoring whether the equipment parameters exceed the normal working range; for time series data, missing data or time logic conflict data can be identified by analyzing the continuity of time nodes; for business logic, operations or states that violate the rules can be detected using the business rules defined in the knowledge base. In specific implementations, anomaly data detection can adopt rule-based methods or combine statistical analysis models or machine learning algorithms to improve the robustness and flexibility of detection.
[0148] Redundant data detection focuses on identifying records that are stored repeatedly or have the same content but different sources in the data. For example, by comparing the topological structures between power equipment nodes, it is detected whether there are repeatedly defined nodes; or through the annotation information of nodes and edges in the semantic graph, it is confirmed whether there are duplicate logical relationships. The detected redundant data will be marked as candidate data for further confirmation during subsequent repair.
[0149] Conflict data detection aims to identify contradictions existing between data. For example, data from multiple sources have different values under the same context conditions. When detecting conflict data, reasoning can be carried out based on the context relationship rules in the knowledge base. For example, whether the logical association between equipment conforms to the topological characteristics of the power system, or whether the sequence of time series events conforms to the power dispatching rules. In complex conflict situations, the reasoning ability of the semantic model can also be used to analyze the reasons for the conflicts.
[0150] After the detection is completed, preliminary repair suggestions are generated in combination with the detection results. The repair suggestions include specific operation guidelines. For example, the proposed correction values for anomaly data, the suggestions for merging or deleting redundant data, and the preferred processing solutions for conflict data. When generating repair suggestions, weight values can be attached to each repair suggestion according to the importance of the data or the credibility of the source for reference in subsequent processing links. This process is usually achieved through a rule engine or a decision tree model to ensure that the generated repair suggestions have logical consistency and operability.
[0151] Through the above steps, the quality problems in the power space data are effectively detected and marked in this stage, and preliminary repair suggestions are generated, laying a solid foundation for the semantic consistency verification and final repair in the subsequent steps.
[0152] Step S105: According to the preliminary repair suggestions and based on the logical rules generated by the semantic model, semantic consistency verification is performed on the repair contents of anomaly data, redundant data, and conflict data.
[0153] The implementation of step S105 involves performing semantic consistency verification on the repair contents of abnormal data, redundant data, and conflicting data according to the logical rules generated by the semantic model based on the preliminary repair suggestions, so as to ensure that the repair process is consistent with the internal logic and business semantics of the power spatial data.
[0154] First, the preliminary repair suggestions generated in step S104 are input into the semantic consistency verification module, and the core of this module depends on the logical rule library generated by the semantic model. The logical rule library includes multi-level rule sets in the spatial dimension, time dimension, and business logic dimension. For example, the physical connection constraint rules between power equipment, the time sequence relationship rules of power dispatching events, and the causal association rules between business operations and states. These rules are generated by the semantic model in the joint modeling and feature extraction stages and can accurately describe the logical associations and context relationships of the power spatial data.
[0155] When performing semantic consistency verification, for each repair suggestion, first match the repair content with the existing data in the knowledge base to confirm whether the repair content conforms to the existing context. For example, for abnormal data repair suggestions, the verification module will verify whether the repair value falls within the normal working range marked in the knowledge base; for the merging suggestions of redundant data, it will confirm whether the merged data meets the requirements of the device topology or business logic; for the preferred solutions of conflicting data, it will evaluate the credibility or consistency of the selected data through logical rule reasoning.
[0156] The verification process adopts a layer-by-layer verification method. First, verify the rules in the spatial dimension, such as whether the physical connection relationship between power equipment is complete and whether the repaired spatial data introduces new conflicts or inconsistencies. Then verify the time dimension rules to confirm whether the events in the time series conform to the power dispatching or operation rules, such as whether the order of the repaired events conforms to the time dependence. Finally, verify the rules in the business logic dimension to ensure that the repair content conforms to the causal relationship and operation specifications of the power business. For example, if the repair value of an event causes inconsistencies with the pre- or post-operations in the business process, the repair content will be marked as inconsistent.
[0157] After completing the semantic consistency verification, the verification module will output the verification results for each repair suggestion, including three statuses: "passed", "needs adjustment", or "suggested rejection". If the verification passes, the repair content can directly enter the next execution stage; if adjustment is needed, the system will feedback the rules or data attributes that need to be improved according to the specific reasons for the verification failure; if rejection is suggested, it indicates that the repair content seriously conflicts with the semantic rules and does not meet the execution conditions.
[0158] Through this process, step S105 ensures that the repair suggestions are not only effective for the data itself, but also consistent with the global logic and business requirements of the power spatial data in terms of overall semantics, thus laying a foundation for the high quality and high reliability of the final data governance results.
[0159] Furthermore, according to the logical rules generated based on the semantic model from the preliminary repair suggestions, semantic consistency verification is performed on the repair contents of abnormal data, redundant data, and conflicting data, including:
[0160] Extract the target data and its context information involved in the preliminary repair suggestions, including the source of the repair data, the applicable logical rules, and the context association conditions;
[0161] Based on the logical rule library generated by the semantic model, perform rule verification on the repair contents one by one to verify whether the repaired data meets the physical constraints of the device, the time series logic, and the business operation specifications;
[0162] For the repair contents involving multiple data sources or depending on context conditions, verify whether the repaired data is consistent with the associated data through the logical reasoning module, including the logical integrity of the event chain and the logical continuity of the device state.
[0163] First, it is necessary to perform a detailed analysis of the preliminary repair suggestions to extract the target data and its context information. This includes clarifying the specific source of the repair content, such as the acquisition device or data source system of the original data record, and at the same time extracting the logical rules and context conditions related to the repair data. For example, if the repair suggestion involves the correction of abnormal device operation status, it is necessary to include the device type, the associated physical topology structure, the relevant time series records, and the applicable business rules.
[0164] After extracting the target data, the verification process depends on the logical rule library generated by the semantic model. This rule library contains a series of rule sets used to describe the logical constraints of power spatial data, including but not limited to the physical constraint conditions of the device, the logical continuity of the time series, and the process specifications of business operations. Each rule is associated with specific context conditions. For example, the physical connection rule between devices can be defined as part of the topology structure, while the time logic rule depends on the time sequence of event occurrence.
[0165] Next, match the repair contents with the logical rules in the rule library one by one to verify whether the logical constraints are met. For example, for the repair of device operation parameters, verify whether it conforms to the normal operation range of the device and whether it is consistent with the status of adjacent devices; for the repair of missing or abnormal values in the time series, verify whether the repaired value is consistent with the trend and dependency relationship of the time series, such as whether it conforms to the numerical change law between time points or the time sequence of event triggering.
[0166] When dealing with repair content involving multiple data sources or complex context conditions, the verification process further invokes the logical reasoning module. Through semantic reasoning, this module compares the repaired data with associated data to verify its consistency with the context information. For example, in the verification of the event chain, the reasoning module needs to verify whether the repaired content maintains logical integrity, such as whether the repaired event order conforms to the causal relationship; in the verification of the device status, it is necessary to ensure that there are no conflicts between the repaired status and the physical topology and operating rules.
[0167] If the verification finds that the repaired content is inconsistent with the logical rules or context information, the repaired content will be marked as verification failed, and the specific failure reasons and conflict points will be recorded, such as violating specific physical constraints or logical rules. This information will be passed as feedback to the repair module for further adjustment of the repair suggestions.
[0168] The entire verification process ensures that the repaired data is logically consistent with the global power space data, conforms to physical and business constraints, and supports subsequent data analysis and governance operations.
[0169] Step S106: Repair the abnormal data, redundant data, and conflicting data according to the verification result of the semantic consistency verification, and generate a governed power space dataset.
[0170] The implementation of step S106 involves performing the final repair operation on the abnormal data, redundant data, and conflicting data according to the verification result of the semantic consistency verification in step S105, thereby generating a governed power space dataset. Specifically, this step strictly combines the repair operation with the semantic logic, context relationship, and business rules of the power space data to ensure the accuracy and reliability of the repair result.
[0171] First, according to the verification result of the semantic consistency verification, classify and process the repair content. For the repair suggestions that pass the verification, directly perform the repair operation and update the repair value or the corrected data to the target dataset. For example, for the repair of abnormal data, directly replace the original abnormal value with the corrected value that passes the verification; for the repair of redundant data, retain the most credible records through merging or deduplication operations; for the repair of conflicting data, adopt the preferred solution to replace the conflict item and remove other conflicting values.
[0172] For the repair suggestions that need to be adjusted, the system will improve the repair content according to the adjustment feedback provided by the verification module. This usually includes adjusting the range of repair values, optimizing the strategy of data merging, or re-evaluating the preference rules for conflicting data. For example, if a time series repair suggestion is marked as needing adjustment because it does not conform to the time continuity rule, the repair value can be recalculated to meet the dependency of the time logic. The improved repair suggestions need to pass the consistency verification again to ensure that the adjusted repair content conforms to all semantic rules.
[0173] For the repair content that is recommended to be rejected, since it has irreconcilable conflicts with the semantic rules, the system will mark it as unrepaired and record the conflict situation for subsequent manual intervention or further analysis. Such data usually needs to be judged by domain experts in combination with the actual business scenario.
[0174] After all repair operations are completed, the system will perform unified formatting and standardization processing on the repaired data set. The formatting process includes field verification, type conversion, and unified encoding of the repaired data to ensure the structural consistency and usability of the data set. At the same time, the system will generate a repair log, which details the original data, repair content, verification results, and final repair values of each repair operation, providing a basis for subsequent auditing and verification.
[0175] Finally, the repaired data set is stored as a governed power spatial data set. This data set not only eliminates abnormal, redundant, and conflicting data, but also ensures the logic and context integrity of the data through semantic consistency verification. The governed data set can be directly used for power grid operation optimization, business analysis, and other subsequent applications, thus significantly improving the quality and utilization value of the data. Through this series of strict repair operations, step S106 provides a reliable basis for the efficient governance and in-depth application of power spatial data.
[0176] In the above embodiments, a data governance method for power spatial data is provided. Correspondingly, the present application also provides a data governance system for power spatial data. Please refer to Figure 2 , which is a schematic diagram of an embodiment of a data governance system for power spatial data of the present application. Since this embodiment, that is, the second embodiment, is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiment. The system embodiments described below are only illustrative.
[0177] A data governance system for power spatial data provided by the second embodiment of the present application includes:
[0178] A modeling unit 201, which is used to jointly model the spatial dimension, time dimension, and business logic dimension of power spatial data by using a deep learning model based on the spatio-temporal characteristics of power spatial data and power business logic, and generate a semantic model with spatio-temporal characteristics and business logic constraints.
[0179] A generating unit 202, which is used to extract features from power spatial data by using the constructed semantic model through an embedding and attention mechanism, and generate multi-level semantic features, where the multi-level semantic features include the spatial correlation, time series relationship, and business logic constraints of power spatial data.
[0180] A constructing unit 203, which is used to construct a semantic graph by using the extracted multi-level semantic features through graph database technology; based on the semantic graph, generate a global power spatial data knowledge base, and dynamically annotate the logical association and context relationship between power spatial data in the knowledge base.
[0181] A detecting unit 204, which is used to detect abnormal data, redundant data, and conflict data in power spatial data based on the global power spatial data knowledge base in combination with a data cleaning algorithm, and generate preliminary repair suggestions.
[0182] A verifying unit 205, which is used to perform semantic consistency verification on the repair content of abnormal data, redundant data, and conflict data according to the preliminary repair suggestions based on the logical rules generated by the semantic model.
[0183] A repairing unit 206, which is used to repair abnormal data, redundant data, and conflict data according to the verification result of the semantic consistency verification, and generate a processed power spatial data set.
[0184] The third embodiment of the present application provides an electronic device, and the electronic device includes:
[0185] A processor;
[0186] A memory, which is used to store a program, and when the program is read and executed by the processor, it executes a data governance method for power spatial data provided in the first embodiment of the present application.
[0187] The fourth embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it executes a data governance method for power spatial data provided in the first embodiment of the present application.
[0188] Although the present application is disclosed above with preferred embodiments, it is not used to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be defined by the scope defined in the claims of the present application.
Claims
1. A data governance method for power spatial data, characterized in that, Including: Based on the spatio-temporal characteristics of power spatial data and power business logic, a deep learning model is used to jointly model the spatial dimension, time dimension, and business logic dimension of power spatial data, generating a semantic model including spatio-temporal characteristics and business logic constraints; Using the constructed semantic model, feature extraction is performed on power spatial data through embedding and attention mechanisms, generating multi-level semantic features, where the multi-level semantic features include the spatial correlation, time series relationship, and business logic constraints of power spatial data; Using the extracted multi-level semantic features, a semantic graph is constructed through graph database technology; based on the semantic graph, a global power spatial data knowledge base is generated, and the logical associations and context relationships between power spatial data are dynamically annotated in the knowledge base; Based on the global power spatial data knowledge base, combined with data cleaning algorithms, abnormal data, redundant data, and conflicting data in power spatial data are detected, and preliminary repair suggestions are generated; According to the preliminary repair suggestions, based on the logical rules generated by the semantic model, semantic consistency verification is performed on the repair contents of abnormal data, redundant data, and conflicting data; Based on the verification results of the semantic consistency verification, abnormal data, redundant data, and conflicting data are repaired, generating a processed power spatial data set; Among them, the method of using a deep learning model to jointly model the spatial dimension, time dimension, and business logic dimension of power spatial data based on the spatio-temporal characteristics of power spatial data and power business logic, generating a semantic model of spatio-temporal characteristics and business logic constraints, includes: Obtaining multi-source heterogeneous input data of power spatial data, where the multi-source heterogeneous input data includes the geographical location, operating status, historical time series data of power equipment, and corresponding business logic rules; performing hierarchical preprocessing on the multi-source heterogeneous input data to form a spatial feature vector, a time series feature vector, and a logical rule feature vector; Based on a graph neural network, the spatial feature vector is modeled to generate a spatial semantic feature including the physical connection relationship and spatial correlation between devices; based on a long short-term memory network, the time series feature vector is modeled to generate a time semantic feature reflecting the historical operating status and time series relationship of devices; based on an attention mechanism combined with rule reasoning, the logical rule feature vector is modeled to generate a business semantic feature describing device operation constraints and event logical dependencies; The spatial semantic feature, time semantic feature, and business semantic feature are fused through a multi-layer fully connected network to generate a semantic model including spatio-temporal characteristics and business logic constraints.
2. The data governance method for power spatial data according to claim 1, wherein The hierarchical preprocessing of the multi-source heterogeneous input data includes: Calculating the weighted distance according to the following formula (1): ; Among them, represents the weighted distance between device and device ; is the geographical distance between device and device ; is the operating load of device ; is the operating load of device ; is the power transmission capacity between device and device ; are the health states of device and device respectively, where the value range of the health state is , where 1 represents completely healthy; is the weight coefficient; is a smoothing parameter to avoid division by zero; Based on dynamic weighted distance , weighted mean clustering algorithm is used to group power equipment, where the weighted optimization objective function of the mean clustering algorithm adopts the following formula (2): ; Among them, is the target number of clusters; is the th cluster; represents the number of devices in the th cluster; is the total number of power devices; is the penalty coefficient for cluster balance; According to the following formula (3), generate a spatial feature vector for the th device : ; Among them, is the geometric center coordinate of the cluster where the th device is located; is the weighted distance between the th device and the center of the cluster to which it belongs, calculated according to Formula 1; represents the number of devices in the th cluster; is the average weighted distance between the th device and all devices within the cluster where it is located.
3. The data governance method for power spatial data according to claim 1, wherein The method of using a graph neural network to model the spatial feature vector to generate a spatial semantic feature including the physical connection relationship and spatial correlation between devices includes: Obtain a dynamic graph representing the topological structure between power equipment, where the nodes of the dynamic graph represent power equipment, the edges represent the physical connection relationships between the equipment, and the weights of the edges are dynamically calculated based on the physical distance, transmission capacity, and connection status between the equipment. The weights of the edges are calculated using the following formula (4): ; Among them, represents the device and the device at the time instant of the edge weight; is the physical distance between nodes and ; represents the time instant when the devices and the devices are in a connection state, with the range being , where 1 represents full connection and 0 represents no connection; is the transmission capacity between nodes and ; , and are weight coefficients; is the smoothing coefficient; Normalize the weights of the edges according to the following formula (5); ; Among them, represents the device and the device at the time instant of the normalized edge weight; represents all the node sets connected to the node represents the device and the device at the moment of the edge weight; Based on the normalized edge weights, use a dynamic graph neural network to perform spatial feature modeling, where the update formula for the nodes represented in each layer of the graph neural network is expressed using the following formula (6): ; Among them, represents the feature representation of node at the layer; is the trainable weight matrix of the layer; represents the feature representation of node at the layer; is the activation function; represents the set of all nodes connected to node ; Perform multi-layer updates on the spatial features through a dynamic graph neural network to generate spatial semantic features reflecting the spatio-temporal correlation between power equipment.
4. The data governance method for power spatial data according to claim 1, wherein Model the logical rule feature vector based on the attention mechanism combined with rule reasoning to generate business semantic features describing equipment operation constraints and event logical dependencies, including: Extract a set of logical rules related to the operation and scheduling of power equipment, and generate a feature representation for each rule in the rule set. The feature representation includes the constraint conditions, applicable scope, and dependencies with other rules of the rule; Use a rule-based inference engine to analyze the set of logical rules and construct a dependency graph between the rules. The nodes of the dependency graph represent rules, and the edges represent the logical dependency strength between the rules; Combine the self-attention mechanism to model the rule nodes in the rule dependency graph, and calculate the priority weight of each rule according to the dependency strength and constraint scope of the rule; According to the priority weights, perform weighted aggregation on the logical rule feature vector to generate business semantic features describing equipment operation constraints and event logical dependencies.
5. The data governance method for power spatial data according to claim 1, characterized in that, The spatial semantic features, temporal semantic features, and business semantic features are fused through a multi-layer fully connected network to generate a semantic model including spatio-temporal characteristics and business logic constraints, including: Normalize the spatial semantic features, temporal semantic features, and business semantic features respectively to ensure that the range of feature values is consistent and avoid the impact of feature scale differences on model training; Construct a multi-layer fully connected network, where the input of each layer of the network is the fused features output by the previous layer of the network, and the network weights adjust the contributions of different feature types through trainable parameters; Add regularization constraints between features in each layer of the fully connected network to limit the excessive influence of specific types of features and ensure the balance of the contributions of spatial, temporal, and business features to the final semantic model; In the last layer of the network for feature fusion, perform a non-linear mapping on the fused features through an activation function to generate a semantic model comprehensively describing spatial correlation, time series relationship, and business logic constraints.
6. The data governance method for power spatial data according to claim 1, wherein Use the constructed semantic model to extract features from power spatial data through embedding and attention mechanisms to generate multi-level semantic features, including: Perform embedding processing on the spatial dimension of power spatial data, and convert the geographical location information, connection relationship, and physical topology characteristics of the equipment into high-dimensional feature representations that can be processed by the model; Serialize and embed the time dimension of power spatial data, map the historical records and time dependencies of device operating states into continuous time feature representations, and capture the impacts of key time points; Perform rule embedding processing on the business logic dimension of power spatial data, transform the logical dependencies of power business rules and events into logical feature vectors, and retain the causal relationships between events; Perform weighted processing on the high-dimensional feature representations, continuous time feature representations, and logical feature vectors through an attention mechanism, dynamically allocate the weights of each feature, extract key features according to the importance of spatial relevance, time series patterns, and business rules, and generate multi-level semantic features including spatial relevance, time series relationships, and business logic constraints.
7. The data governance method for power spatial data according to claim 1, wherein Using the extracted multi-level semantic features, construct a semantic graph through graph database technology, including: Divide the extracted multi-level semantic features into feature sets of nodes and edges, where node features include the spatial location, time state, and business logic attributes of power devices, and edge features include the physical connection relationships between devices, the time dependencies between events, and the logical rule constraint relationships; Define the graph structure based on node features and edge features, where nodes represent power devices or key events, and edges represent the association relationships between devices or the logical dependencies between events; Create storage structures for nodes and edges in the graph database, and attach attributes to each node and edge, including spatial characteristics, time characteristics, and context information of business rules; Dynamically update the nodes and edges in the graph structure according to the spatio-temporal changes of power spatial data to ensure that the semantic graph can reflect the changes in device states and the adjustments of business logic in real time.
8. The data governance method for power spatial data according to claim 1, wherein According to the preliminary repair suggestions, perform semantic consistency verification on the repair contents of abnormal data, redundant data, and conflict data based on the logical rules generated by the semantic model, including: Extract the target data and its context information involved in the preliminary repair suggestions, including the source of the repair data, the applicable logical rules, and the context association conditions; Based on the logical rule library generated by the semantic model, perform rule verification on the repair contents one by one to verify whether the repair data meets the physical constraints of the device, the time series logic, and the business operation specifications; For the repair contents involving multiple data sources or depending on context conditions, verify whether the repaired data is consistent with the associated data through a logical reasoning module, including the logical integrity of the event chain and the logical continuity of the device state.
9. A data governance system for power spatial data, characterized in that, Include: A modeling unit for jointly modeling the spatial dimension, time dimension, and business logic dimension of power spatial data using a deep learning model based on the spatio-temporal characteristics of power spatial data and power business logic, and generating a semantic model of spatio-temporal characteristics and business logic constraints; A generation unit for using the constructed semantic model to extract features from power spatial data through embedding and attention mechanisms, and generating multi-level semantic features, where the multi-level semantic features include the spatial relevance, time series relationships, and business logic constraints of power spatial data. A construction unit, which is used to construct a semantic map by using the extracted multi-level semantic features through graph database technology; based on the semantic map, generate a global power space data knowledge base, and dynamically annotate the logical associations and context relationships between power space data in the knowledge base; A detection unit, which is used to detect abnormal data, redundant data and conflict data in power space data based on the global power space data knowledge base and in combination with a data cleaning algorithm, and generate preliminary repair suggestions; A verification unit, which is used to perform semantic consistency verification on the repair contents of abnormal data, redundant data and conflict data according to the preliminary repair suggestions and based on the logical rules generated by the semantic model; A repair unit, which is used to repair abnormal data, redundant data and conflict data according to the verification results of the semantic consistency verification and generate a processed power space data set; wherein, the modeling unit is specifically used for: Obtain multi-source heterogeneous input data of power space data, where the multi-source heterogeneous input data includes the geographical location, operating status, historical time series data of power equipment, and corresponding business logic rules; perform hierarchical preprocessing on the multi-source heterogeneous input data to form a spatial feature vector, a time series feature vector and a logical rule feature vector; Based on a graph neural network, model the spatial feature vector to generate a spatial semantic feature including the physical connection relationship and spatial correlation between devices; based on a long short-term memory network, model the time series feature vector to generate a time semantic feature reflecting the historical operating status and time series relationship of the device; based on an attention mechanism combined with rule reasoning, model the logical rule feature vector to generate a business semantic feature describing device operation constraints and event logical dependencies; Fuse the spatial semantic feature, the time semantic feature and the business semantic feature through a multi-layer fully connected network to generate a semantic model including spatio-temporal characteristics and business logic constraints.
Citation Information
Patent Citations
Business process arrangement method and system based on power grid operation knowledge
CN113962549A
Electric power data semantic analysis method and device, storage medium and program product
CN118861969A