Quasi-real-time lake and reservoir fusion calculation component construction method based on data platform

By constructing a near real-time lake warehouse fusion computing component for the data middle platform, the problems of semantic inconsistency and insufficient real-time performance of multi-source data in the power grid system were solved. This enabled unified modeling of multi-source data and efficient and stable cross-system collaborative computing, thereby improving the real-time performance and adaptive capabilities of power grid data processing.

CN121743409APending Publication Date: 2026-03-27STATE GRID LIAONING ELECTRIC POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing power grid systems suffer from semantic inconsistencies, insufficient real-time performance, weak collaborative capabilities, and poor stability in multi-source data processing, making it difficult to meet the needs of new power systems for multi-source fusion, real-time interaction, and intelligent decision-making.

Method used

We construct a near real-time lake warehouse converged computing component based on a data middle platform. By unifying metadata, table structure, and lineage information to form a data semantic layer, we achieve near real-time access, format normalization, and time-series alignment of multi-source data. We also design a converged computing engine that supports cross-storage computing in the lake warehouse, with execution plan splitting, incremental processing, and heterogeneous scheduling functions to ensure computing consistency and high concurrency.

Benefits of technology

It has achieved unified modeling of multi-source power grid measurement data in lake areas, warehouse areas and log systems, improved the real-time performance, stability and adaptability of data processing, significantly improved the throughput and result consistency of the power grid online monitoring and analysis link, and can cope with adaptive optimization in complex power grid scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121743409A_ABST
    Figure CN121743409A_ABST
Patent Text Reader

Abstract

The invention discloses a quasi-real-time lake-cabin fusion calculation component construction method based on a data platform, and the method comprises the following specific steps: collecting the measurement data of each business system of a power grid, and forming a data source system of a lake-cabin integrated structure; the metadata, the table structure and the blood relationship information of the lake-warehouse integrated architecture are integrated through a data platform to form a unified data semantic layer; on the basis of a unified data semantic layer, performing quasi-real-time access, format regulation and time sequence alignment on the data to generate standardized intermediate data; and designing a fusion calculation engine supporting lake and reservoir cross-storage calculation by utilizing the intermediate data. According to the method, unified modeling, quasi-real-time access and cross-storage cooperative calculation of multi-source data of power grid multi-source measurement data in a lake and warehouse integrated scene are realized, and the timeliness of a data processing link, the stability of a calculation task and the self-adaptive capability of a component in a complex business scene are also remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power data engineering and lake warehouse integrated computing, and particularly relates to a quasi-real-time lake warehouse fusion computing component construction method based on a data middle platform. BACKGROUND

[0002] With the advancement of new power system construction, power grid business has gradually evolved from the traditional mode mainly relying on manual inspection and offline statistics into a digital operation mode driven by massive real-time data. The current power system generally has the characteristics of "high concurrency measurement, high frequency state reporting, cross-system collaborative computing, and large-scale multi-source data storage". Multi-source data such as voltage, current, load, power quality, oil temperature, temperature and humidity, wind speed, and action events from SCADA, DMS, OMS, metering systems, online monitoring devices, distribution network terminals, power consumption sensors, and weather stations continue to be generated, making the structure type, collection period, and business semantics of power grid operation data more complex.

[0003] In the traditional architecture, although the data warehouse can support structured modeling and offline analysis, its data update period is usually minutes or hours, which is difficult to meet the requirements of "quasi-real-time processing" in scenarios such as voltage limit detection, transformer temperature rise monitoring, line load trend judgment, and wind and light power prediction. Although the data lake has massive storage capacity, it lacks field-level governance, consistency management, and high-performance computing capabilities for power grid multi-source measurement data, making it difficult to undertake cross-system fusion computing tasks.

[0004] Lakehouse technology bridges the gap between the capabilities of lakes and warehouses to some extent, making data lakes have features such as tabular management, ACID transactions, columnar acceleration, and unified metadata. However, it still faces obvious bottlenecks in the landing of power grid business: first, the semantics of multi-source data are not unified. The data caliber, field naming, physical quantity unit, and sampling frequency of different power grid systems differ, making it difficult for the lake-warehouse system to form a unified interpretation layer, and cross-lake-warehouse computing is prone to field conflicts or inconsistent meanings. Second, real-time access to multi-source data is difficult to standardize. There are various forms of data in power grid systems, such as streaming data (such as voltage, current, oil temperature), batch data (such as topology, archives), and log data (such as protection actions, fault events). The existing access links lack unified management of format, timing, and error codes, which can easily cause out-of-order, sudden delay, or data drift, affecting the accuracy and real-time performance of computing. Third, the computing and scheduling capabilities of cross-lake-warehouse are insufficient. Power grid analysis involves data lake engines (such as Spark, Flink), warehouse engines (such as ClickHouse, MPP), and log analysis engines. Existing solutions lack mechanisms for execution plan splitting, operator pushing, and incremental computing based on voltage load characteristics, and computing is prone to high latency, high resource consumption, and even inconsistent results. Fourth, there is a lack of adaptive optimization mechanisms. Factors such as power load fluctuations, wind and solar power output changes, device data delays, and collection link congestion frequently change the computing environment. Existing lake-warehouse platforms use static configurations and cannot dynamically adjust strategies based on latency, resource bottlenecks, or data bias, which can lead to unstable or failed tasks such as voltage monitoring, temperature rise analysis, and power evaluation.

[0005] As a data governance and sharing platform for power grid enterprises, the data middle platform has the ability of metadata management, field governance, and service orchestration, but it still lacks the ability of real-time data access, multi-engine fusion computing, and feedback-driven optimization in the power grid scenario, making it difficult to support the business model of "multi-source fusion, real-time interaction, and intelligent decision-making" required by the new power system.

[0006] Therefore, to fundamentally break through the technical bottlenecks of existing systems in terms of semantic inconsistency, real-time performance, weak collaboration, and poor stability, it is a technical problem that has been long desired to be solved in the field. SUMMARY

[0007] The purpose of the present application is to provide a method for building a quasi-real-time lake-warehouse fusion computing component based on a data middle platform, which effectively solves the problem of inconsistent physical quantity meanings of multi-source measurement data in traditional lake-warehouse systems and the difficulty of cross-system collaboration. It also significantly improves the timeliness, stability of data processing links, and the adaptive ability of components in complex business scenarios.

[0008] The technical solution of the present application to solve the above technical problems is as follows:

[0009] This invention provides a method for constructing a near real-time lake warehouse fusion computing component based on a data middle platform, comprising the following steps:

[0010] Collect measurement data from various power grid business systems to form a data source system with an integrated lake-warehouse architecture;

[0011] By integrating the metadata, table structure, and lineage information of the lake warehouse integrated architecture through the data middle platform, a unified data semantic layer is formed;

[0012] Based on the unified data semantic layer, streaming data, batch data and logs are accessed in near real-time, formatted and time-series aligned, and standardized intermediate data is generated.

[0013] Using the aforementioned intermediate data, a converged computing engine supporting cross-storage computing in lake warehouses is designed. The converged computing engine has the functions of execution plan splitting, incremental processing, and heterogeneous scheduling. The converged computing engine realizes high-concurrency convergence and consistency processing of near real-time tasks.

[0014] Alternatively, the data platform first performs a unified abstraction of all structural information in the lake warehouse data source system, then standardizes the fields based on the business meaning of the power grid physical quantities, and finally generates field-level lineage relationships for cross-source collaboration.

[0015] Optionally, the unified data semantic layer includes field naming conventions, attribute type constraints, cross-source field mapping rules, and field lineage links to ensure the uniformity of multi-source data in terms of structure, meaning, and processing path; the unified data semantic layer supports automatic parsing of metadata and rule-based processing of field conflicts.

[0016] Optionally, the near real-time access step includes monitoring the data access rate, source node latency, and timestamp drift, and achieving time consistency of multi-source data through dynamic offset calibration; the format normalization step involves mapping the input data fields to a unified field structure, and field normalization includes field padding, field pruning, and field reordering; the time alignment step involves correcting any input timestamp based on a standard time field.

[0017] Optionally, the standardized intermediate data retains field source and processing chain information to enable full-chain traceability.

[0018] Optionally, the fusion computing engine decomposes complex computing tasks into multiple sub-operator chains based on a unified semantic layer and field lineage; performs incremental processing on all physical quantities according to a unified time slice based on a standard time axis; and selects execution nodes based on data type and storage structure.

[0019] Alternatively, the fusion computing engine performs consistency checks through semantic layer constraints and field lineage; and performs high-concurrency fusion output of all results from the real-time task.

[0020] Optionally, it also includes a computation process monitoring step, which monitors latency, resource usage and data offset during the computation process, and dynamically adjusts operator links and scheduling strategies in combination with execution feedback to achieve adaptive performance optimization for different scenarios.

[0021] Optionally, the dynamic adjustment of operator links includes adjusting the execution order of operators, splitting high-load operators, merging lightweight operators, and dynamically adjusting the window size; the dynamic adjustment of scheduling strategies includes node migration, cross-storage path switching, and dynamic priority adjustment.

[0022] Alternatively, the adaptive performance optimization includes automatic rearrangement of operator chains, dynamic scaling of computation windows, execution node migration, and storage path switching.

[0023] This invention offers the following advantages: First, by constructing a unified semantic layer, it achieves unified modeling and consistent field definitions for multi-source power grid measurement data such as voltage, current, load, oil temperature, wind speed, temperature and humidity, and equipment status across lake areas, warehouse areas, and log systems. This effectively solves the problems of inconsistent physical quantity meanings and difficulties in cross-system collaboration in traditional lake-warehouse systems. Second, relying on standardized intermediate data structures and cross-source time alignment mechanisms, it enables near real-time access of streaming measurements, batch processing operation data, and equipment action logs in a low-latency, traceable manner, significantly improving the stability of the power grid online monitoring and analysis link. Third, by constructing a cross-storage fusion computing engine, combined with semantically driven execution plan splitting, time-slice incremental calculation, and heterogeneous storage scheduling, it enables collaborative execution of object storage, columnar warehouses, and log engines, improving the throughput and result consistency of high-concurrency tasks such as voltage limit detection, transformer temperature rise analysis, and line load rate calculation. Fourth, by real-time monitoring of latency, resource consumption, and multi-source data offset, and dynamically adjusting the operator chain structure, calculation window, and scheduling strategy based on feedback, it achieves adaptive optimization processing for complex power grid scenarios such as sudden traffic surges, drastic fluctuations in wind and solar power, and unstable equipment reporting. In summary, this invention significantly improves the real-time performance, stability, consistency, and adaptability of multi-source data processing in the power grid under the integrated lake-warehouse architecture, and has outstanding engineering application value. Attached Figure Description

[0024] Figure 1 The flowchart illustrates the method for constructing a near real-time lake warehouse fusion computing component based on a data middle platform, as provided in this invention. Detailed Implementation

[0025] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0026] Example

[0027] like Figure 1 As shown, this invention provides a method for constructing a near real-time lake warehouse fusion computing component based on a data middle platform, including:

[0028] S1: By integrating the metadata, table structure, and lineage information of the lake warehouse integrated architecture through the data middle platform, a unified data semantic layer is formed, providing a unified management and cross-source collaboration basic environment for near real-time fusion computing;

[0029] In this embodiment, physical measurement data such as voltage, current, load, power quality, oil temperature, temperature and humidity, and operating status generated by various power grid business systems (such as SCADA, DMS, OMS, metering acquisition, online monitoring, and meteorological monitoring) are uniformly entered into the data source system of the lake warehouse integrated architecture. The data platform first uniformly abstracts all structural information in the lake warehouse data source system, then standardizes the fields based on the business meaning of the power grid physical quantities, and finally generates field-level lineage relationships that can be used for cross-source collaboration.

[0030] (1) Unified abstraction of multi-source metadata

[0031] The data tables, fields, and partition attributes from the power grid multi-source measurement system and the lake-warehouse integrated architecture are abstracted into a unified metadata set:

[0032] M = {T i C ij ,P ij} (1)

[0033] Wherein: T i For the i-th table; C ij For table T i The j-th field; P ij This refers to the partition attributes corresponding to the table (such as collection date, equipment area, voltage level, etc.).

[0034] The basic attributes of a field are further represented as follows:

[0035] C ij =(n ij ,t ij ,d ij (2)

[0036] Where, n ij For field name; t ij For the field's primitive type; d ij Provide field descriptions (such as "transformer oil temperature", "phase A voltage", "ambient humidity", etc.).

[0037] (2) Standardization of field types and names

[0038] To ensure consistency across systems, this invention maps field types from different source systems to a unified set of types, and performs type normalization using the following formula:

[0039]

[0040] in, α represents the standardized field type; α is the type mapping coefficient, used to achieve unification such as VARCHAR→STRING, DATETIME→TIMESTAMP, etc.

[0041] Field names are also standardized, and their structure is represented as follows:

[0042]

[0043] in, β is the normalized field name (all lowercase and separated by "_"); β is the name normalization coefficient.

[0044] The standardized field structure is represented as follows:

[0045]

[0046] The standardized structure of the entire data table is as follows:

[0047]

[0048] (3) Construction of field-level lineage relationships

[0049] To support the near real-time fusion calculation of power grid physical quantities such as voltage, current, oil temperature, and load, this invention constructs a field-level lineage matrix by parsing SQL jobs, real-time streaming operator links, and integrated lakehouse processing tasks.

[0050]

[0051] Among them, C ab For source field; C cd The target field is L; L indicates the lineage relationship between the fields.

[0052] For a certain target field C cd All source fields are used to form a bloodline origin set:

[0053] S cd ={C ab |L(C ab C cd )=1} (8)

[0054] If multiple source fields jointly generate a single target field, their combined influence is expressed as follows:

[0055]

[0056] Among them, w ab This indicates the data quality weight of the source field, including update latency, null value rate, and measurement reliability.

[0057] (4) The composition of the unified semantic layer

[0058] After completing metadata abstraction, structural standardization, and lineage construction, this invention defines the semantic layer in a unified mathematical form:

[0059]

[0060] in, L is the set of standardized table structures; L is the set of field lineage relationships; d is semantic information such as field meaning, business definition, and indicator system.

[0061] The unique identifier for each field in the semantic layer is:

[0062]

[0063] Where h(·) is the field identifier generation method, used to ensure the consistency of the same physical quantity in different source systems.

[0064] S2: Based on the unified semantic layer of S1, it performs near real-time access, format normalization and time sequence alignment of streaming data, batch data and logs, and generates standardized intermediate data to ensure that the input side is controllable, low-latency and traceable.

[0065] In this embodiment, step S2, based on the unified semantic layer U established in step S1, performs near real-time access, field normalization, time base alignment, and intermediate data construction on multi-source measurement data (including SCADA voltage / current, power quality of distribution terminals, low-voltage user voltage of concentrators, transformer oil temperature, line temperature, wind speed, ambient temperature and humidity, user-side load event streams, and equipment action logs) from power grid field equipment and business systems, enabling it to support subsequent cross-lake warehouse fusion calculations. The specific process is described below.

[0066] (1) Quasi-real-time access and unified representation of multiple types of power grid data

[0067] Multi-source measurement data from different power grid business systems are synchronously entered into the data platform via real-time links, batch processing links, and equipment operation logs, and are uniformly abstracted into the following set of data streams:

[0068] X = {X f ,Xb ,X l} (12)

[0069] Among them, X f This represents real-time streaming data, such as SCADA bus voltage, current, feeder load, wind speed, oil temperature, etc., measured in seconds or sub-seconds; X b This represents batch data, such as topology, device records, operating modes, historical data flow files, etc.; X l This represents log-type data, such as protection action logs, alarm events, and device status change records.

[0070] Each input record is abstracted as a triple:

[0071]

[0072] Among them, v k This is a vector of observed values ​​of physical quantities in the power grid. The timestamps generated by field equipment or systems; src k This serves as an identifier for the source device or system, such as station number, feeder number, concentrator number, sensor ID, etc. This abstraction ensures that physical quantities from different data acquisition devices can be consistently represented.

[0073] (2) Normalization of power grid physical quantity fields based on semantic layer

[0074] To address the inconsistency in field naming across different power grid systems (e.g., "Ua", "voltage_A", "PhaseA_V"), this step relies on the standard field set constructed in S1. Map input data fields to a unified field structure:

[0075] y k =Φ(v k ,T * (14)

[0076] Among them, y k For the normalized records, such as the unified A-phase voltage, current, oil temperature, ambient temperature, load, etc.; Φ(·) represents the field normalization mapping rules, including field completion (missing oil temperature is automatically filled with null values), pruning (removing irrelevant columns), and order rearrangement, etc.; T * A set of standardized table structures from S1.

[0077] Field regularization includes field completion, field clipping, and field reordering, which can be further formalized as:

[0078]

[0079] in All of these are standardized physical parameters actually measured by the power grid, such as oil temperature (°C), wind speed (m / s), voltage (V), and current (A) after unifying the units.

[0080] (3) Time unification and multi-source timing alignment based on power grid clock system

[0081] The time bases of multi-source measurement systems in power grids differ: SCADA time may be offset by ±1 second; concentrator data is at the minute level; PMU is synchronized with GPS; and there is communication delay in the time of online monitoring equipment. To ensure accurate calculation of power flow, load trends, and over-limit detection, this invention performs time alignment. The standard time field t defined in the semantic layer of S1 is used. * Based on this, for any input timestamp Correction:

[0082]

[0083] Among them, t k This is the standard time after alignment; Data source src k The time offset is used to correct errors caused by communication link delays, network congestion, or device clock drift.

[0084] After unified time alignment, the physical quantities of the power grid form a strict time series:

[0085] Z = {y k |Press t k Ascending order} (17)

[0086] This ensures the consistency of physical processes such as voltage fluctuations, oil temperature changes, and sudden changes in wind speed over time.

[0087] (4) Generate standardized intermediate data with physical meaning

[0088] After field normalization and time alignment, this invention encapsulates multi-source power grid data into a unified intermediate structure:

[0089]

[0090] Among them, y k Standardized physical quantities such as voltage, current, oil temperature, and wind speed; t k For standardized timestamps; A unique code for the source device (e.g., "station number-interval-device ID").

[0091] To achieve cross-source traceability, intermediate data is associated with fields such as lineage, equipment origin, and data collection path.

[0092] Meta k ={Scd ,L,src k ,t k} (19)

[0093] Among them, S cd The set of physical quantity sources for generating this field; L represents the field's lineage; src k and t k Ensure traceability back to the original equipment and the time of data collection.

[0094] The final dataset obtained can be used for near real-time fusion computing:

[0095] Q = (D mid Meta) (20)

[0096] Meta contains the physical source, field lineage, and time information of all input records, ensuring the transparency, interpretability, and traceability of the power grid measurement data processing process.

[0097] S3: Utilizing the intermediate data from S2, a fusion engine supporting cross-storage computing in lake warehouses is designed. Through execution plan splitting, incremental processing, and heterogeneous scheduling, high concurrency and consistency processing of near real-time tasks are achieved.

[0098] In this embodiment, step S3 is based on the intermediate data output by S2. This step constructs a fusion engine that supports cross-storage collaborative computing in lake warehouses. Through execution plan splitting, time slice incremental processing, heterogeneous execution scheduling and semantic consistency verification, it realizes near real-time high-concurrency processing of power grid operation monitoring tasks (such as voltage over-limit detection, transformer temperature rise analysis, line load rate calculation, wind speed and temperature comparison analysis, low-voltage user voltage qualification rate calculation, etc.).

[0099] The fusion computing engine in this embodiment is driven by "semantic layer constraints + field lineage information + unified time slice structure", which unifies the data sources from different physical quantities such as voltage, current, oil temperature, and wind speed to achieve continuous computing across lake area, warehouse area, and log area.

[0100] (1) Semantic-driven execution plan splitting mechanism

[0101] Power grid operation analysis typically involves multiple sequential calculation steps of physical quantities, such as: deriving the line load rate from voltage and current; deriving the transformer temperature rise trend from oil temperature and load correlation; forming a wind turbine output prediction window from wind speed, temperature, and SCADA power; and calculating the voltage qualification rate from the three-phase voltage sequence. The fusion calculation engine, based on the unified semantic layer and field lineage relationships in step S1, decomposes the complex calculation task into multiple sub-calculation chains with clear boundaries and consistent semantics.

[0102] F = {O1,O2,...,O} n} (twenty one)

[0103] Among them O i The operators are minimal and include operators such as "voltage normalization operator," "oil temperature moving average operator," "wind speed-power conversion operator," and "low-voltage user voltage qualification judgment operator." Each operator corresponds to a clear physical meaning of the power grid. This decomposition process ensures the consistency of cross-system calculation logic; for example, semantic conflicts will not occur when SCADA voltage data and online monitoring voltage data are fused.

[0104] (2) Incremental processing mechanism based on unified time slice

[0105] Power grid measurement data has natural time series characteristics. Therefore, based on the standard time axis established in step S2, this invention performs incremental processing on all physical quantities according to a unified time slice.

[0106] Incremental window representation is as follows:

[0107] ΔWt=Wt-Wt -δ (twenty two)

[0108] Here, Wt represents the measured data such as voltage, current, oil temperature, and wind speed collected within the current window; δ is determined by the actual power grid task, such as 1 second, 5 seconds, 1 minute, etc.; incremental processing can quickly identify real-time characteristics such as sudden rises in oil temperature, sudden drops in voltage, and sudden changes in wind speed. Through a unified time-slice mechanism, data from different systems (such as SCADA, metering, distribution transformer monitoring, and weather stations) can be merged within the same window, ensuring the temporal integrity of the power grid's physical processes.

[0109] (3) Heterogeneous scheduling for lake warehouse integrated architecture

[0110] Grid data is stored in different types of lakehouse engines:

[0111] Frequently updated real-time voltage and current readings are more suitable for the lake region (Iceberg / Delta Lake).

[0112] Large-scale offline historical trend data → More suitable for warehouse areas (Hive / ClickHouse)

[0113] • Protect action logs and alarm logs → More suitable for the log area (object storage)

[0114] The fusion engine of this invention automatically selects different execution nodes based on factors such as data source type, field structure, and storage engine capabilities. For example:

[0115] Lake area data → Distributed to object storage scan operator nodes, suitable for high-concurrency data such as voltage and current;

[0116] Warehouse data → scheduled to SQL / MPP nodes for massive offline power flow calculations;

[0117] Log data is routed to the text parsing node for action backtracking and event detection.

[0118] This achieves the best match between the physical quantity processing and storage system capabilities.

[0119] (4) Cross-lake warehouse consistency guarantee

[0120] In a lake-warehouse integrated architecture, different engines may produce different results for the same computational logic (such as different window calculations, deduplication strategies, and data uploading cycles).

[0121] To ensure the consistency of power grid operation analysis results, this invention performs consistency checks through semantic layer constraints and field lineage relationships:

[0122] Consistent(F,t)=1 (23)

[0123] This indicates that the execution results of all subtasks of task F within window t conform to the semantic layer definition, field meaning, and lineage logic.

[0124] (5) Near real-time high-concurrency fusion output

[0125] The final fusion result can be expressed as:

[0126]

[0127] in, The fusion results for each sub-operator under window t include, for example, a list of voltage over-limit points, transformer temperature rise trends, line load rates, and wind speed sudden change alarms; all results are guaranteed to be consistent in physical meaning, consistent across storage links, and consistent in time slices.

[0128] S4: Monitor latency, resource and data offset during the computation process, and dynamically adjust operator links and scheduling strategies based on execution feedback to achieve adaptive performance optimization for different scenarios and improve the stability of fusion computing.

[0129] In this step, the "adaptive performance optimization layer" of the present invention is constructed to monitor the entire process of lake warehouse fusion computing. Based on real-time feedback from the power grid scenario, the operator chain structure, computing window size and cross-storage scheduling strategy are dynamically adjusted to ensure that near real-time fusion computing maintains stable, efficient and consistent processing capabilities even in environments with significant voltage fluctuations, frequent changes in equipment status, drastic fluctuations in wind and solar power, and rapid growth of multi-source heterogeneous data.

[0130] By real-time monitoring of the delay, drift, and resource usage of key power grid measurement data (such as voltage, current, load, oil temperature, wind speed, temperature and humidity, protection actions, etc.), this invention achieves adaptive optimization for online power grid analysis applications, including near real-time tasks such as voltage over-limit detection, transformer temperature rise trend prediction, line load rate calculation, and wind power analysis.

[0131] (1) Real-time monitoring mechanism for delays, resource and data offsets in power grid tasks

[0132] During the execution of fusion computing, this invention continuously collects the following key performance indicators through a monitoring system:

[0133] ①Latency

[0134] Real-time monitoring of the processing time of each operator during the execution phase, including data read latency, cross-memory access latency, network transmission latency, and operator queuing time, is used to determine whether operator congestion, node overload, or network anomalies occur. For example, when wind speed data latency exceeds a threshold, it will directly affect the real-time performance of wind power forecasting.

[0135] ② Resource Utilization

[0136] Monitoring includes CPU utilization, memory consumption, disk I / O, and network bandwidth. When certain node resources are detected to be approaching saturation, an automatic migration mechanism can be triggered.

[0137] ③ Data Drift

[0138] Combined with the unified time axis in step S2, the time consistency of multi-source data is monitored in real time, including: time drift of SCADA and online monitoring voltage sequences; lag in transformer oil temperature data; out-of-order wind speed / temperature data; delay in action recording caused by log accumulation; lag in uploading low-voltage user voltage batch data; once the offset expands, it will directly affect the physical validity of cross-source calculations (such as voltage + temperature + load relationship).

[0139] (2) Dynamic adjustment of operator links based on feedback

[0140] When the monitored indicators exceed the threshold, this invention automatically adjusts the operator chain through an optimization engine, including:

[0141] ① Adjust the operator execution order

[0142] If the delay of the voltage waveform cleaning operator increases, lightweight filtering and field clipping are automatically performed in advance; if the delay of the oil temperature exponential smoothing operator increases, the feature extraction operator is moved in advance to reduce the input size; thus reducing the pressure on downstream operators and improving overall throughput.

[0143] ② Splitting high-load operators

[0144] The calculation of "large-scale network load rate" is split into parallel operations by region and voltage level; the calculation of "station-level wind speed-power mapping model" is split into separate operations by measurement point to improve the parallelism of calculation.

[0145] ③ Merging lightweight operators

[0146] When node resources are plentiful or business is stable, multiple lightweight operators can be merged to reduce scheduling frequency and context switching, thereby improving overall execution efficiency.

[0147] ④ Dynamically adjust window size

[0148] When wind speed fluctuates drastically, load suddenly increases, or oil temperature changes abruptly, the window is automatically shortened to improve response speed; when the system load is low and the data is stable, the window is expanded to improve throughput and batch processing efficiency, ensuring that the results are both real-time and stable.

[0149] (3) Dynamic optimization of heterogeneous scheduling strategy for lake warehouse integrated architecture

[0150] This invention achieves optimal allocation of computing resources in cross-lake warehouse scenarios through dynamic scheduling of execution plans:

[0151] ① Node migration (cross-node scheduling)

[0152] When SQL nodes are queuing due to a large number of power flow calculations, the voltage / oil temperature operators are migrated to idle nodes; when object storage nodes are saturated with I / O, the wind speed data reading operators are migrated to the distributed caching system to achieve load balancing.

[0153] ②Cross-storage path switching

[0154] The system automatically switches execution paths between the lake area and the warehouse area based on latency and resource status:

[0155] When I / O pressure in the lake area increases, some tasks are switched to the warehouse area for execution.

[0156] • When there is severe queuing in the warehouse area, switch to lake area object storage for execution.

[0157] Implement "storage structure-aware scheduling".

[0158] ③ Dynamic priority adjustment

[0159] When a voltage over-limit event occurs, increase the priority of the "voltage compliance analysis operator"; during typhoons or rainstorms, increase the priority of "wind speed change detection" and "line icing analysis"; ensure that critical power grid safety tasks are executed first.

[0160] (4) Adaptive performance optimization closed loop

[0161] This invention constructs the above-mentioned monitoring, adjustment, and scheduling optimization into a continuously operating closed-loop mechanism, including:

[0162] Monitor: Collects latency, resource, and offset information.

[0163] Analyze: Identify bottlenecks, drift, out-of-order loading, and load anomalies.

[0164] Adjustment: Rearrange operator links, split and combine operators, adjust window.

[0165] Optimize: Migrate tasks, switch paths, adjust priorities

[0166] Execute: Continue executing the fused computing according to the optimized plan.

[0167] This closed loop will run continuously throughout the entire converged computing process, enabling real-time self-repair, self-optimization, and self-adaptation.

[0168] This invention also provides a storage medium storing a computer program. When executed by a processor, the computer program implements some or all of the steps in the various embodiments of the method for constructing a near real-time lakehouse fusion computing component based on a data platform provided by this invention. The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0169] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or certain parts of the embodiments of the present invention.

[0170] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for constructing a near real-time lakehouse fusion computing component based on a data middle platform, characterized in that, Includes the following steps: Collect measurement data from various power grid business systems to form a data source system with an integrated lake-warehouse architecture; By integrating the metadata, table structure, and lineage information of the lake warehouse integrated architecture through the data middle platform, a unified data semantic layer is formed; Based on the unified data semantic layer, streaming data, batch data and logs are accessed in near real-time, formatted and time-series aligned, and standardized intermediate data is generated. Using the aforementioned intermediate data, a converged computing engine supporting cross-storage computing in lake warehouses is designed. The converged computing engine has the functions of execution plan splitting, incremental processing, and heterogeneous scheduling. The converged computing engine realizes high-concurrency convergence and consistency processing of near real-time tasks.

2. The method for constructing a near real-time lake warehouse fusion computing component based on a data middle platform according to claim 1, characterized in that, The data platform first performs a unified abstraction of all structural information in the lake warehouse data source system, then standardizes the fields based on the business meaning of the power grid physical quantities, and finally generates field-level lineage relationships for cross-source collaboration.

3. The method for constructing a near real-time lake warehouse fusion computing component based on a data middle platform according to claim 1, characterized in that, The unified data semantic layer includes field naming conventions, attribute type constraints, cross-source field mapping rules, and field lineage links, which are used to ensure the uniformity of multi-source data in terms of structure, meaning, and processing path. The unified data semantic layer supports automatic parsing of metadata and rule-based processing of field conflicts.

4. The method for constructing a near real-time lake warehouse fusion computing component based on a data middle platform according to claim 1, characterized in that, The near real-time access step includes monitoring the data access rate, source node latency, and timestamp drift, and achieving time consistency of multi-source data through dynamic offset calibration; the format normalization step involves mapping the input data fields to a unified field structure, and field normalization includes field padding, field pruning, and field reordering; the time alignment step involves correcting any input timestamp based on a standard time field.

5. The method for constructing a near real-time lake warehouse fusion computing component based on a data middle platform according to claim 1, characterized in that, The standardized intermediate data retains field source and processing chain information to enable full-chain traceability.

6. The method for constructing a near real-time lake warehouse fusion computing component based on a data middle platform according to claim 1, characterized in that, The fusion computing engine decomposes complex computing tasks into multiple sub-operator chains based on a unified semantic layer and field lineage. Based on the standard timeline, all physical quantities are processed incrementally according to a unified time slice; the execution node is selected based on the data type and storage structure.

7. The method for constructing a near real-time lake warehouse fusion computing component based on a data middle platform according to claim 1, characterized in that, The fusion computing engine performs consistency checks through semantic layer constraints and field lineage; and performs high-concurrency fusion output of all results from aligning with real-time tasks.

8. The method for constructing a near real-time lake warehouse fusion computing component based on a data middle platform according to claim 1, characterized in that, It also includes a computation process monitoring step, which monitors latency, resource usage and data offset during the computation process, and dynamically adjusts operator links and scheduling strategies in combination with execution feedback to achieve adaptive performance optimization for different scenarios.

9. The method for constructing a near real-time lake warehouse fusion computing component based on a data middle platform according to claim 8, characterized in that, The dynamic adjustment of operator links includes adjusting the execution order of operators, splitting high-load operators, merging lightweight operators, and dynamically adjusting the window size; the dynamic adjustment of scheduling strategies includes node migration, cross-storage path switching, and dynamic priority adjustment.

10. The method for constructing a near real-time lake warehouse fusion computing component based on a data middle platform according to claim 8, characterized in that, The adaptive performance optimization includes automatic rearrangement of operator chains, dynamic scaling of computation windows, execution node migration, and storage path switching.