Data management and intelligent analysis method oriented to power multi-source heterogeneity
Through ontology model and space-time joint indexing technology, the multi-source heterogeneous data of power are integrated, and the safety threshold is dynamically adjusted, which solves the problem of poor adaptability of data islands and fixed thresholds, and realizes efficient data integration and accurate anomaly detection.
Patent Information
- Application Number
- CN202510581432.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-15
AI Technical Summary
The existing multi-source heterogeneous data governance methods of power are difficult to effectively integrate structured, semi-structured and unstructured data, resulting in serious data silos, and fixed security thresholds cannot dynamically adapt to environmental changes, resulting in false positives or missed reports.
Ontology model alignment, temporary ontology generation and multimodal semantic embedding technologies are used to achieve semantic unification across data types, establish a spatio-temporal joint indexing mechanism, dynamically adjust the operating data security threshold, and perform abnormal detection in combination with device parameters and environmental data.
It improves data integration efficiency and semantic consistency, improves the accuracy and credibility of abnormal detection, reduces false alarm rates, and quickly locates equipment failures and abnormal causes.
Smart Images

Figure CN120492678A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-source heterogeneous data quality and analysis technology, and relates to a data governance and intelligent analysis method for multi-source heterogeneous power. Background Art
[0002] Multi-source heterogeneous data refers to data sets within the power system that originate from different devices, platforms, and business processes, with varying formats, structures, protocols, and spatiotemporal characteristics. Currently, multi-source heterogeneous data in the power system faces challenges such as cross-system data fusion and dynamic data governance. Therefore, methods for the governance and analysis of multi-source heterogeneous data in the power system are of paramount importance.
[0003] The existing governance and analysis methods for multi-source heterogeneous power data can already meet basic needs, but they still have certain shortcomings: on the one hand, traditional governance methods for multi-source heterogeneous power data are difficult to effectively integrate structured, semi-structured and unstructured data, and cannot achieve data integration efficiency and semantic consistency, resulting in serious data silos.
[0004] On the other hand, existing technologies for detecting data anomalies mostly rely on fixed safety thresholds and cannot dynamically adapt to environmental changes such as temperature fluctuations and equipment aging. This can easily lead to false alarms or missed alarms due to equipment aging, affecting operation and maintenance efficiency. Summary of the Invention
[0005] In view of this, in order to solve the problems raised in the above background technology, a data governance and intelligent analysis method for multi-source heterogeneous power is proposed.
[0006] The purpose of the present invention can be achieved through the following technical solutions: The present invention provides a data governance and intelligent analysis method for multi-source heterogeneous power, including: S1, multi-source heterogeneous data acquisition: multi-source heterogeneous power data is collected in real time according to the multi-source heterogeneous equipment deployed in the target power grid area, and the data type is determined, including structured data, semi-structured data or unstructured data.
[0007] S2. Semantic mapping of multi-source heterogeneous data: triggering corresponding semantic mapping logic according to the data type of the multi-source heterogeneous power data, integrating and storing it in the database, and establishing a spatiotemporal joint indexing mechanism to realize multi-dimensional linkage query of data.
[0008] S3. Data anomaly analysis: Based on the environmental data in the multi-source heterogeneous power data, the dynamic safety threshold of the power equipment operation data is analyzed through the built-in dynamic threshold generation unit, and the operation data in the current multi-source heterogeneous power data is dynamically compared with its corresponding safety threshold to determine whether it is abnormal data and mark it.
[0009] S4. Confirm the cause of the anomaly: Confirm the cause of the anomaly according to the data anomaly confirmation logic, and generate an anomaly confidence assessment index. At the same time, associate the cause of the anomaly with the confidence assessment index to generate an abnormal event record, push it to the operation and maintenance system, and trigger an alarm.
[0010] Compared with the existing technology, the beneficial effects of the present invention are as follows: 1. The present invention classifies and identifies structured, semi-structured and unstructured data, and adopts ontology model alignment, temporary ontology generation and multimodal semantic embedding technology respectively to achieve semantic unification across data types. At the same time, it establishes a spatiotemporal joint indexing mechanism to bind and store geographic coordinates and timestamps, supports efficient multi-dimensional queries, and improves data integration efficiency and semantic consistency.
[0011] 2. The present invention solves the problem of poor adaptability of fixed thresholds by dynamically adjusting the operating data safety threshold and anomaly confidence assessment based on the deviation coefficient between environmental monitoring data and historical benchmarks, improves the detection accuracy and credibility of anomalies in multi-source heterogeneous power data, improves the accuracy of anomaly detection, and reduces the false alarm rate.
[0012] 3. The present invention constructs multi-level abnormality confirmation logic, combines equipment parameters, environmental data and communication status, and accurately locates the causes of equipment failure, environmental interference and communication abnormalities, which helps to quickly locate the root cause and shorten the time, and push differentiated alarm strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0014] Figure 1 Schematic diagram of the method steps of the present invention. DETAILED DESCRIPTION
[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0016] See also Figure 1As shown, the present invention provides a data governance and intelligent analysis method for multi-source heterogeneous power, and the specific steps are as follows: S1, multi-source heterogeneous data acquisition: multi-source heterogeneous power data is collected in real time according to the multi-source heterogeneous equipment deployed in the target power grid area, and the data type is determined, including structured data, semi-structured data or unstructured data.
[0017] As a preferred feasible embodiment, the specific process of determining the data type of the multi-source heterogeneous power data includes: A1, data source identification and classification: matching the multi-source heterogeneous devices corresponding to the multi-source heterogeneous power data in the target power grid area with the physical devices in the database (such as sensor time series data (SCADA), smart meter readings (CSV / database)) and non-physical devices (such as work order texts, inspection images, voice records, log files, etc.), and determining that if the multi-source heterogeneous power data corresponds to a physical device, the multi-source heterogeneous power data is recorded as structured data; otherwise, the multi-source heterogeneous power data is recorded as semi-structured data or unstructured data.
[0018] A2. Extension identification and classification: Match the file extensions of the target power grid area's multi-source heterogeneous data with the corresponding extensions of each data type in the database to preliminarily determine the data type of the multi-source heterogeneous data.
[0019] A specific example is structured data: .csv (smart meter readings), .db (database files), .sql (relational database scripts), etc.
[0020] Semi-structured data: .json (configuration file), .xml (device parameter log), .log (system log), etc.
[0021] Unstructured data: .jpg / .png (inspection images), .wav (voice recordings), .pdf (work order documents), etc.
[0022] A3. Data type determination: Obtain the file byte count, entropy value, and keyword density of the target power grid area multi-source heterogeneous data, and determine the data type of the multi-source heterogeneous data based on the feature-data type mapping table.
[0023] It needs to be further explained that the specific method of obtaining the file byte number, entropy value and keyword density of the target power grid area multi-source heterogeneous data is: directly read the number of binary bytes stored in the file and record it as the file byte number of the target power grid area multi-source heterogeneous data.
[0024] Entropy is a measure of the uncertainty or complexity of data. The higher the entropy, the more dispersed the data distribution (such as unstructured images and texts), and the lower the entropy, the more regular the data format (such as structured tables). Count the occurrence frequency P of each byte (or character) in the target power grid area's multi-source heterogeneous data. m , where m = 1, 2, ..., n, m is the number of each byte, n is the number of bytes, according to the application information entropy formula The entropy value H of the target power grid area's multi-source heterogeneous data is obtained. Among them, structured data (such as .csv): field separators (commas, tabs) appear repeatedly, and the entropy value is low, while unstructured data (such as .jpg): pixel values are randomly distributed, and the entropy value is high.
[0025] Keyword density is the frequency of occurrence of pre-set keywords (such as business-related terms and data field names) in the target power grid region's multi-source heterogeneous data, reflecting the semantic concentration of the data content. The number of occurrences of pre-set keywords in the target power grid region's multi-source heterogeneous data is counted and compared with the total number of words in the target power grid region's multi-source heterogeneous data. The result is recorded as the keyword density of the target power grid region's multi-source heterogeneous data.
[0026] Specifically, the feature-data type mapping is shown in Table 1 below.
[0027] Table 1 Example of feature-data type mapping
[0028] Data Type File Bytes Characteristics Entropy range Keyword density Structured data Small to medium size (1KB-10MB) H<5 High (>15%) Semi-structured data Medium to large scale (10MB-1GB) 5≤H<7 Moderate (5%-15%) Unstructured data Wide range (KB-GB+) 7≤H Low (<5%)
[0029] If the number of bytes is moderate, the format is regular (such as a table), the field names (keywords) appear repeatedly, and the entropy value is low (data distribution is concentrated), then it is structured data.
[0030] If the number of bytes is large (nested structure), the keywords (such as tags and key names) are regular but not repeated, and the entropy value is medium (the format is semi-regular), then it is semi-structured data.
[0031] If the number of bytes varies greatly (such as image resolution), there is no fixed format, and the keywords are scattered (such as image content without text semantics), then it is unstructured data.
[0032] S2. Semantic mapping of multi-source heterogeneous data: triggering corresponding semantic mapping logic according to the data type of the multi-source heterogeneous power data, integrating and storing it in the database, and establishing a spatiotemporal joint indexing mechanism to realize multi-dimensional linkage query of data.
[0033] As a preferred feasible embodiment, the specific content of the data type corresponding semantic mapping logic includes: (1) structured data corresponding semantic mapping logic: using the ontology model to directly match the database table structure and field semantics.
[0034] It's important to further clarify that the ontology model shown here is a formalized knowledge representation method used to describe concepts, entities, and their relationships within a specific domain, aiming to build a shared framework with unified semantics. It's more than just a definition of data structure; it's an abstraction and logical representation of real-world knowledge, enabling computers to understand the meaning behind the data and supporting automated reasoning and cross-system collaboration.
[0035] For example, the database table PowerMeter corresponds to the ontology class ElectricMeter, and the field voltage is mapped to the ontology property hasVoltage. The relationship between classes and properties is defined using OWL (Web Ontology Language), supporting automated reasoning (e.g., voltage anomaly detection rules).
[0036] (2) Semantic mapping logic for semi-structured data: Parse the nested structure of JSON / XML, generate a temporary ontology model (e.g., JSON Schema → OWL class), align the temporary ontology with the global ontology to match the database table structure and field semantics. Resolve naming conflicts (e.g., "load" in the log refers to "user load", but is defined as ConsumerLoad in the ontology).
[0037] (3) Semantic mapping logic for semi-structured data: Semantic vectors are extracted through models (BERT, RoBERTa, etc.), context-related word embeddings are generated, and visual feature vectors are extracted (using neural networks such as ResNet and CNN) to capture equipment failure modes (such as cracks in transformer oil chromatogram images). Multimodal Transformer (such as VideoBERT) is used for joint encoding, and the feature vectors are input into the ontology embedding model (such as TransE, RotatE) and mapped to the ontology concept space.
[0038] For example, the BERT vector corresponding to the text "line overload" is mapped to the ontology class OverloadEvent. The feature vector of a device crack in an image is mapped to the ontology attribute hasPhysicalDamage.
[0039] As a preferred feasible embodiment, the specific establishment process of the spatiotemporal joint indexing mechanism includes: B1, data integration storage: converting the geographic coordinates of the target power grid area multi-source heterogeneous power data according to the established two-dimensional coordinate system, and combining them with the corresponding timestamps to obtain the standard storage format of the multi-source heterogeneous power data, and storing it in the database.
[0040] Specifically, a two-dimensional coordinate system is established with the center point of the target power grid area as the origin. The geographic coordinates of the target power grid area's multi-source heterogeneous data are mapped to the corresponding positions in the two-dimensional coordinate system using the longitude and latitude of the geographic coordinates as the horizontal and vertical axes, respectively. The converted geographic coordinates (e.g., coordinates (500, 300)) and timestamp (e.g., 1682847600123 milliseconds) are then "bound together" to form a "space-time tag."
[0041] When each multi-source heterogeneous power data is stored in the database, in addition to its own numerical value (such as voltage and current), it must contain a "time-space label" - that is, the converted two-dimensional coordinates (X, Y) and timestamp (T).
[0042] B2. Index structure selection: Construct query spatiotemporal index keys according to query requirements, and each query spatiotemporal index key is unique, thereby performing spatiotemporal joint query of the multi-source heterogeneous power data.
[0043] Specifically, each index key is composed of a timestamp (T) and geographic coordinates (X, Y) and is unique (one index key corresponds to only one data record).
[0044] For example, the spatiotemporal index key of a piece of data may be: T=1682841600(2025-04-29 08:00:00)+(X=850, Y=750), which corresponds to the voltage data of a substation at that moment.
[0045] Construct an index key range of "time interval + space region" (such as time [T1, T2], space [(X1, Y1), (X2, Y2)]).
[0046] Nearest neighbor query: Build an index key of "target coordinates + distance sorting" (for example, with the coordinates of the fault point as the center, generate an index key sequence by distance).
[0047] The present invention classifies and identifies structured, semi-structured and unstructured data, and adopts ontology model alignment, temporary ontology generation and multimodal semantic embedding technology to achieve semantic unification across data types. At the same time, it establishes a spatiotemporal joint indexing mechanism to bind and store geographic coordinates and timestamps, supports efficient multi-dimensional queries, and improves data integration efficiency and semantic consistency.
[0048] S3. Data anomaly analysis: Based on the environmental data in the multi-source heterogeneous power data, the dynamic safety threshold of the power equipment operation data is analyzed through the built-in dynamic threshold generation unit, and the operation data in the current multi-source heterogeneous power data is dynamically compared with its corresponding safety threshold to determine whether it is abnormal data and mark it.
[0049] As a preferred feasible embodiment, the specific analysis process of the dynamic safety threshold of the power equipment operation data includes: analyzing the environmental monitoring data Evn of each power equipment in the environmental data of the power multi-source heterogeneous data ig , and the basic environmental monitoring data Evn′ corresponding to the basic safety threshold when the operating data of the corresponding power equipment corresponds to the basic environmental monitoring data Evn′ ig Perform deviation amplitude fusion calculation to obtain the deviation coefficient of the environmental monitoring data of each power device, where i = 1, 2, ..., j, i is the number of each power device, j is the number of power devices, g = 1, 2, ..., d, g is the number of each environmental monitoring data, and d is the number of environmental monitoring data.
[0050] In a specific example, the source of each basic environmental monitoring data corresponding to the basic safety threshold of the power equipment's operating data can be based on the normal state statistics of the power equipment's historical operating data. For example, the historical stable range of environmental data (such as the average temperature over the past three months) corresponding to the power equipment's operating data during a fault-free period (such as the average voltage fluctuation over seven consecutive days) can be used.
[0051] Specifically, the specific method of calculating the deviation amplitude fusion is: according to the formula Perform deviation amplitude fusion calculation to obtain the deviation coefficient of environmental monitoring data of each power equipment By calculating the relative deviation between the environmental monitoring data and the operating data of the corresponding power equipment corresponding to the basic safety threshold, the degree of influence of environmental factors on the operating data of the power equipment is quantified. When the deviation coefficient is larger, it indicates that the potential interference of environmental changes on the operating data of the power equipment is greater.
[0052] According to the environmental monitoring data deviation-safety threshold mapping table, the safety threshold corresponding to the deviation coefficient of the environmental monitoring data of each power device is obtained, and the safety threshold corresponding to the deviation coefficient of the environmental monitoring data of each power device is further multiplied by the basic safety threshold corresponding to the corresponding operating data to obtain the operating data safety threshold of each power device.
[0053] As a specific example, assume that the current of an electric device corresponds to a basic safety threshold of 100 amperes, and the environmental monitoring data is temperature. Under these conditions, a mapping table of environmental monitoring data deviations to safety thresholds is shown in Table 1 below, and an example calculation of the current corresponding to the basic safety threshold of an electric device is shown in Table 2.
[0054] Table 1 Example of environmental monitoring data deviation-safety threshold mapping table
[0055] Coefficient of variation interval Operational data security threshold ≤0.1 1.00 0.1-0.2 (including 0.2) 0.95 0.2-0.3 (including 0.3) 0.90 >0.30 0.85
[0056] Table 2 Calculation example of the basic safety threshold corresponding to the current of power equipment
[0057]
[0058] The deviation coefficient of environmental monitoring data (reflecting the magnitude of environmental fluctuations) is converted into an operational data safety threshold, reflecting the dynamic impact of the environment on the operational safety threshold of power equipment. For example, when the temperature deviation coefficient is 0.2 (20% beyond the normal range), the operational data safety threshold is 0.95, indicating that the safety threshold of the power equipment under the current environment needs to be reduced by 5% from the base value to reserve a safety margin.
[0059] By integrating environmental data with equipment operating data, safety thresholds are more closely aligned with the equipment's actual operating conditions. For example, in high-temperature environments, equipment insulation performance degrades, lowering the safety current threshold. If the actual current approaches the adjusted threshold, the system will issue an early warning, preventing missed faults caused by fixed thresholds.
[0060] As a preferred feasible embodiment, the specific judgment process of whether the operating data in the multi-source heterogeneous power data within the current preset detection time period is abnormal data includes: extracting the operating monitoring data of each current power device from the operating data in the current multi-source heterogeneous power data, and comparing them with the corresponding safety thresholds respectively. If a certain operating monitoring data of a certain power device is less than or equal to the corresponding safety threshold, the operating monitoring data of the current power device is recorded as abnormal data.
[0061] The present invention solves the problem of poor adaptability of fixed thresholds by dynamically adjusting the operating data safety threshold and anomaly confidence assessment based on the deviation coefficient between environmental monitoring data and historical benchmarks, improves the detection accuracy and credibility of anomalies in multi-source heterogeneous power data, improves the accuracy of anomaly detection, and reduces the false alarm rate.
[0062] S4. Confirm the cause of the anomaly: Confirm the cause of the anomaly according to the data anomaly confirmation logic, and generate an anomaly confidence assessment index. At the same time, associate the cause of the anomaly with the confidence assessment index to generate an abnormal event record, push it to the operation and maintenance system, and trigger an alarm.
[0063] As a preferred feasible embodiment, the specific content of the data anomaly confirmation logic includes: if the abnormal data is accompanied by equipment parameters (such as insulation resistance, vibration amplitude) that continuously deviate from the baseline value and conforms to the preset fault mode library (such as short circuit, overload characteristics), it is determined to be an equipment failure.
[0064] If the abnormal data coincides with the period of sudden change in data from external environmental sensors (such as temperature, humidity, and wind speed), and there is no persistent abnormality in the equipment parameters, it is determined to be environmental interference.
[0065] If multiple related parameters of the same device (such as current and power) simultaneously experience non-physical regular jumps, or the packet loss rate suddenly increases, it is determined to be a communication anomaly.
[0066] As a preferred feasible embodiment, the specific generation method of the anomaly confidence evaluation index includes: obtaining the anomaly duration of the anomaly data and the number of associated device alarms, and normalizing them to obtain the relative proportion of the anomaly duration of the anomaly data and the relative proportion of the number of associated device alarms.
[0067] It should be further explained that the specific method of obtaining the abnormal duration of the abnormal data, the number of associated device alarms and the model output probability value is: record the duration of the abnormal data from the first trigger to the current moment, and use it as the abnormal duration of the abnormal data.
[0068] The number of concurrent alarms of devices associated with the power equipment described in the abnormal data (such as adjacent transformers and electric meters on the same feeder) in the target power grid area is counted and recorded as the number of alarms of associated devices of the abnormal data.
[0069] It should be further explained that the normalization process to obtain the relative proportion of abnormal duration and the relative proportion of the number of associated device alarms of the abnormal data specifically includes: comparing the abnormal duration with the maximum permitted duration stored in the database (e.g., 30 minutes) to obtain the relative proportion of the abnormal duration of the abnormal data. Comparing the number of associated alarms (N) with the maximum permitted number of associated devices stored in the database (e.g., 10) to obtain the relative proportion of the number of associated device alarms of the abnormal data (N% = N / 10 × 100%).
[0070] The relative proportion of the abnormal duration of the abnormal data and the relative proportion of the number of associated device alarms are weighted and summed according to the preset corresponding weights to obtain the abnormal confidence evaluation index of the abnormal data.
[0071] In a specific example, the preset corresponding weights of the relative proportion of the abnormal duration and the relative proportion of the number of associated device alarms may be 0.6 and 0.4, respectively.
[0072] Specifically, power equipment faults (such as aging insulation and poor contact) usually do not disappear instantly. The longer the duration of abnormal data, the more likely it is to be a real fault rather than a temporary interference. For example, a transformer overload may last for several minutes or even longer, while anomalies caused by environmental interference (such as instantaneous lightning strikes) are often short-lived. Giving the abnormality duration a high weight of 0.6 reflects the judgment logic of "the longer it lasts, the more serious the problem", avoiding the misjudgment of instantaneous fluctuations (such as voltage spikes) as serious faults. If only relying on the number of associated device alarms (the weight is too high), it may cause misjudgment due to local communication interference or common-mode interference of multiple devices (such as multiple devices in the same area triggering alarms at the same time due to sudden weather changes). The duration indicator can filter out short-term noise. For example, if a device abnormality lasts for 5 minutes (with a high duration ratio), even if there is only one associated alarm, it may be more worthy of attention than a situation that lasts for 30 seconds but has three associated alarms.
[0073] Power system equipment is physically interconnected (e.g., multiple meters or adjacent transformers on the same feeder). If abnormal data is accompanied by alarms from multiple connected devices, this indicates a regional or systemic problem (e.g., a busbar failure causing synchronization anomalies in surrounding equipment). A weight of 0.4 is assigned to the number of connected device alarms because it provides important clues about the spread of an anomaly. However, this information must be considered in conjunction with its duration: a short, sudden, and potentially multi-device alarm could indicate external interference (e.g., lightning causing widespread communication anomalies), while a prolonged single-device alarm is more likely to indicate a fault within the device itself.
[0074] As a preferred feasibility embodiment, the specific generation process of the abnormal event record includes: if the cause of the abnormality is equipment failure and the confidence assessment index is greater than or equal to 60%, it is determined to be a valid fault event; if the cause of the abnormality is equipment failure and the confidence assessment index is less than the set threshold but there is an alarm of related equipment (such as an abnormality of an adjacent transformer), it is upgraded to a suspected fault event.
[0075] If the cause of the anomaly is environmental interference and the confidence assessment index is greater than or equal to 80% and the environmental sensor data synchronization is abnormal, it is determined to be a confirmed environmental event, otherwise it is marked as an event to be verified.
[0076] If the abnormality is caused by communication abnormality and the confidence assessment index is greater than or equal to 50% and the data packet loss rate exceeds the threshold, a communication failure event is directly generated.
[0077] As a preferred feasible embodiment, the method uses a database during the execution process to store multi-source heterogeneous power data in the target power grid area, store each physical device and each non-physical device, store the extension name corresponding to each data type, store the maximum permitted duration and the maximum permitted number of associated devices.
[0078] The present invention constructs multi-level anomaly confirmation logic, combines equipment parameters, environmental data and communication status, and accurately locates causes such as equipment failures, environmental interference and communication anomalies, which helps to quickly locate the root cause and shorten the time, and push differentiated alarm strategies.
[0079] The above contents are merely examples and explanations of the concept of the present invention. Those skilled in the art may make various modifications or additions to the described specific embodiments or replace them in a similar manner. As long as they do not deviate from the concept of the invention or exceed the scope defined by the present invention, they should all fall within the scope of protection of the present invention.
Claims
1. A data governance and intelligent analysis method for multi-source heterogeneous power systems, characterized by: include: S1. Multi-source heterogeneous data acquisition: Real-time collection of multi-source heterogeneous power data from multi-source heterogeneous devices deployed in the target power grid area, and determination of data types, including structured data, semi-structured data, or unstructured data; S2. Semantic mapping of multi-source heterogeneous data: triggering corresponding semantic mapping logic based on the data type of the multi-source heterogeneous power data, integrating and storing it in the database, and establishing a spatiotemporal joint indexing mechanism to realize multi-dimensional data linkage query; S3. Data anomaly analysis: Based on the environmental data in the multi-source heterogeneous power data, the built-in dynamic threshold generation unit analyzes the dynamic safety threshold of the power equipment operation data. The current operation data in the multi-source heterogeneous power data is dynamically compared with the corresponding safety threshold to determine whether it is abnormal data and mark it; S4. Confirm the cause of the anomaly: Confirm the cause of the anomaly according to the data anomaly confirmation logic, and generate an anomaly confidence assessment index. At the same time, associate the cause of the anomaly with the confidence assessment index to generate an abnormal event record, push it to the operation and maintenance system, and trigger an alarm.
2. The method for data governance and intelligent analysis of multi-source heterogeneous power systems according to claim 1 is characterized by: The specific process of determining the data type of the multi-source heterogeneous power data includes: S1. Multi-source heterogeneous data acquisition: Real-time collection of multi-source heterogeneous power data from multi-source heterogeneous devices deployed in the target power grid area, and determination of data types, including structured data, semi-structured data, or unstructured data; A1. Data Source Identification and Classification: Match the target power grid area's multi-source heterogeneous data corresponding to the multi-source heterogeneous devices with the physical devices and non-physical devices in the database. If the multi-source heterogeneous data corresponds to a physical device, the data is recorded as structured data. Otherwise, the data is recorded as semi-structured data or unstructured data. A2. Extension identification and classification: Match the file extensions of the target power grid area's multi-source heterogeneous data with the corresponding extensions of each data type in the database to preliminarily determine the data type of the multi-source heterogeneous data; A3. Data type determination: Obtain the file byte count, entropy value, and keyword density of the target power grid area multi-source heterogeneous data, and determine the data type of the multi-source heterogeneous data based on the feature-data type mapping table.
3. The method for data governance and intelligent analysis of multi-source heterogeneous power systems according to claim 1 is characterized by: The specific contents of the semantic mapping logic corresponding to the data type include: (1) Structured data corresponding semantic mapping logic: using the ontology model to directly match the database table structure and field semantics; (2) Semantic mapping logic for semi-structured data: Parse the nested structure of JSON / XML, generate a temporary ontology model (e.g., JSON Schema → OWL class), and align the temporary ontology with the global ontology to match the database table structure and field semantics; (3) Semantic mapping logic for semi-structured data: Semantic vectors are extracted through models (BERT, RoBERTa, etc.), context-related word embeddings are generated, and visual feature vectors are extracted to capture equipment failure modes. Multimodal Transformer joint encoding is used to input the feature vectors into the ontology embedding model and map them to the ontology concept space.
4. The method for data governance and intelligent analysis of multi-source heterogeneous power systems according to claim 1 is characterized by: The specific process of establishing the spatiotemporal joint indexing mechanism includes: B1. Data integration and storage: The geographic coordinates of the target power grid area's multi-source heterogeneous power data are converted according to the established two-dimensional coordinate system, and combined with the corresponding timestamps to obtain a standard storage format for the multi-source heterogeneous power data, which is then stored in the database; B2. Index structure selection: Construct query spatiotemporal index keys according to query requirements, and each query spatiotemporal index key is unique, thereby performing spatiotemporal joint query of the multi-source heterogeneous power data.
5. The method for data governance and intelligent analysis of multi-source heterogeneous power systems according to claim 1 is characterized by: The specific analysis process of the dynamic safety threshold of the power equipment operation data includes: From the environmental monitoring data of each power device in the environmental data of the multi-source heterogeneous power data, the deviation amplitude of each basic environmental monitoring data corresponding to the basic safety threshold of the operating data of the corresponding power device is fused and calculated to obtain the deviation coefficient of the environmental monitoring data of each power device; According to the environmental monitoring data deviation-safety threshold mapping table, the safety threshold corresponding to the deviation coefficient of the environmental monitoring data of each power device is obtained, and the safety threshold corresponding to the deviation coefficient of the environmental monitoring data of each power device is further multiplied by the basic safety threshold corresponding to the corresponding operating data to obtain the operating data safety threshold of each power device.
6. The method for data governance and intelligent analysis of multi-source heterogeneous power systems according to claim 1 is characterized by: The specific process of determining whether the operating data in the electric power multi-source heterogeneous data within the current preset detection time period is abnormal data includes: The operation data of each current power equipment is extracted from the current multi-source heterogeneous power data, and compared with the corresponding safety thresholds respectively. If a certain operation monitoring data of a current power equipment is less than or equal to the corresponding safety threshold, the operation monitoring data of the current power equipment will be recorded as abnormal data.
7. The method for data governance and intelligent analysis of multi-source heterogeneous power systems according to claim 1 is characterized by: The specific contents of the data anomaly confirmation logic include: If the abnormal data is accompanied by the continuous deviation of the equipment parameters from the reference value and meets the preset fault mode library, it is determined to be an equipment failure; If the abnormal data coincides with the period of sudden change in the external environmental sensor data, and there is no persistent abnormality in the device parameters, it is determined to be environmental interference; If multiple parameters associated with the same device change simultaneously without physical regularity, or if the packet loss rate suddenly increases, it is considered a communication anomaly.
8. The method for data governance and intelligent analysis of multi-source heterogeneous power systems according to claim 7 is characterized by: The specific generation method of the abnormal confidence evaluation index includes: Obtain the abnormal duration of abnormal data and the number of associated device alarms, and normalize them to obtain the relative proportion of the abnormal duration of abnormal data and the relative proportion of the number of associated device alarms; The relative proportion of the abnormal duration of the abnormal data and the relative proportion of the number of associated device alarms are weighted and summed according to the preset corresponding weights to obtain the abnormal confidence evaluation index of the abnormal data.
9. The method for data governance and intelligent analysis of multi-source heterogeneous power systems according to claim 1 is characterized by: The specific generation process of the abnormal event record includes: If the cause of the abnormality is a device failure and the confidence level is greater than or equal to 60%, it is considered a valid fault event. If the cause of the abnormality is a device failure and the confidence level is less than the set threshold but there are related device alarms, it is upgraded to a suspected fault event. If the cause of the anomaly is environmental interference and the confidence assessment index is greater than or equal to 80%, and the environmental sensor data synchronization is abnormal, it is determined to be a confirmed environmental event; otherwise, it is marked as an event to be verified; If the abnormality is caused by communication abnormality and the confidence assessment index is greater than or equal to 50% and the data packet loss rate exceeds the threshold, a communication failure event is directly generated.
10. The method for data governance and intelligent analysis of multi-source heterogeneous power systems according to claim 1 is characterized by: During the execution of this method, a database is used to store multi-source heterogeneous power data in the target power grid area, store each physical device and each non-physical device, store the extension name corresponding to each data type, and store the maximum permitted duration and the maximum permitted number of associated devices.
Citation Information
Patent Citations
Multi-source heterogeneous data semantic integration model constructed based on domain ontology and method
CN104182454A
Power equipment image data warehouse and power equipment defect detection method
CN111078912A
Data query method and device for multi-source heterogeneous data, storage medium and equipment
CN112699141A
Method, device and equipment for automatically generating alarm information and storage medium
CN117376093A
Target system data intelligent monitoring method and system based on large model
CN119004367A
Cited By
Real-time adaptive constant value checking method and system for relay protection device
CN120873637A
A real-time adaptive setting value checking method and system for a relay protection device
CN120873637B
Power sensitive database lightweight supervision method and system based on metadata analysis
CN121029706A
Multi-source access and abnormal retry switchgear waveform acquisition method and alarm method
CN121164884A
Multi-source heterogeneous data fusion method and device
CN121166672A