Power plant information data intelligent analysis system and method for regional power grid

CN122838680APending Publication Date: 2026-09-29ZHONGNENG ZHIWANG (CHENGDU) NEW ENERGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611036674.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

现有方法多依赖固定模型或离线训练结果,运行模式切换时原有模型立即终止,分析结果易出现跳变;数值分析与语义推理通常独立或松耦合工作,缺乏自动化的双向交互机制,难以实现从“数据异常”到“因果推理”再到“增强识别”的认知迭代

Benefits of technology

[0017]与现有技术相比,本发明的有益效果包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838680A_ABST
    Figure CN122838680A_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent analysis system and method for power plant information data in regional power grids, belonging to the field of data processing technology. The system includes a data semantic unification and adaptive storage organization module, a situational awareness and multi-mode analysis and reasoning module, and a forward-looking inference and closed-loop feedback module. These three modules work collaboratively at the data layer, analysis layer, and decision layer. The data layer forms a unified view through ontology-driven semantic unification and spatiotemporal alignment, constructs a spatial fractal self-similar index, and uses index hit frequency weighting to drive multi-temperature storage preheating. The analysis layer extracts multi-dimensional features to identify operating modes and dynamically selects models, combining knowledge graph semantic reasoning to achieve numerical and semantic fusion through a bidirectional enhanced closed loop. The decision layer uses a reduced-order digital twin inference scheme, layered asynchronous feedback to adjust parameters at each layer, and establishes cross-layer coordination. This invention improves data fusion efficiency, query response speed, and decision accuracy, and is suitable for intelligent dispatching and operation management of regional power grids.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to an intelligent analysis system and method for power plant information data in regional power grids. Background Technology

[0002] With the deepening of the construction of new power systems, regional power grids have deployed multiple information systems, including SCADA, PMU, smart meters, meteorological monitoring, and online equipment monitoring. These systems differ significantly in data format, semantic expression, and spatiotemporal reference, forming a typical multi-source heterogeneous data environment. Existing semantic fusion solutions mostly use static ontologies or predefined mapping files. The ontologies cannot adaptively evolve with the access of new devices, and spatiotemporal alignment operations often lose the original information of high-frequency data sources due to simple resampling. At the same time, semantic fusion, index construction, and storage management operate independently, lacking a connected data flow and feedback flow, making it difficult to form a complete link from data access to efficient access.

[0003] At the data access level, power grid operation data exhibits significant spatiotemporal multidimensional characteristics. Dispatchers need to grasp the overall network situation from a regional perspective, as well as drill down to specific equipment and even sensors for microscopic positioning. Existing R-tree or B-tree index structures have pre-fixed hierarchies, requiring the creation of different indexes or full scans for cross-granularity queries, resulting in long response times. When the topology changes, the index often needs to be completely rebuilt, leading to high maintenance costs. Although multi-temperature storage solutions have been applied in the power sector, their migration decisions rely solely on the historical access statistics of the data objects themselves, failing to perceive changes in upper-level analysis and query patterns, and lacking predictive data warming capabilities based on index semantics.

[0004] At the situational awareness and analysis level, the high penetration rate of new energy sources has led to increasingly complex and variable power grid operation modes. Existing methods mostly rely on fixed models or offline training results. When the operation mode changes, the original model immediately terminates, and the analysis results are prone to abrupt changes. Numerical analysis and semantic reasoning usually work independently or loosely coupled, lacking an automated two-way interaction mechanism, making it difficult to achieve cognitive iteration from "data anomalies" to "causal reasoning" and then to "enhanced recognition." The construction of knowledge graphs mostly relies on manual annotation and offline updates, and the timeliness and accuracy of the reasoning results are difficult to meet the requirements of online operation.

[0005] At the decision-making and closed-loop optimization levels, high-fidelity digital twin models suffer from high computational overhead, and simplified proxy models have limited generalization capabilities, making it difficult to cover various operating modes such as strong fluctuations in new energy sources, equipment maintenance, and fault recovery. The parameter update cycles of the data, analysis, and decision layers are singular and evolve independently, lacking cross-layer coordination mechanisms and prone to inter-layer mismatch. The utilization of feedback signals is also rather crude, failing to implement layered asynchronous optimization based on the temporal characteristics of parameter changes at different levels.

[0006] In view of the above, this application is hereby submitted. Summary of the Invention

[0007] To address the aforementioned problems in existing technologies, this paper provides an intelligent analysis system and method for power plant information data in regional power grids, aiming to solve at least one of the above problems.

[0008] The technical solution to achieve the purpose of this invention is as follows:

[0009] This invention provides an intelligent analysis system for power plant information data in regional power grids, which includes a data semantic unification and adaptive storage organization module, a situational awareness and multi-mode analysis and reasoning module, and a forward-looking inference and closed-loop feedback module. The three modules work collaboratively according to the data layer, analysis layer, and decision layer, and form a self-optimizing closed-loop system through cross-layer feedback channels.

[0010] The data semantic unification and adaptive storage organization module, as the core of the data layer, includes a semantic unification and spatiotemporal alignment unit, a spatial fractal self-similar indexing unit, a multi-temperature storage unit driven by access frequency prediction, and a cross-unit collaborative mechanism. This module transforms the raw data from multi-source heterogeneous power plants into a unified semantic data view and supports efficient access for upper-layer analysis with an adaptive indexing and storage structure.

[0011] The situational awareness and multi-mode analysis and reasoning module, as the core of the analysis layer, includes a situational adaptive model selection engine and a knowledge graph semantic reasoning unit, which collaborate through a bidirectional enhanced closed loop. This module intelligently analyzes the regional power grid's operational status from both numerical and semantic dimensions.

[0012] The forward-looking simulation and closed-loop feedback module, serving as the core of the decision-making layer, includes a digital twin simulation unit and a hierarchical asynchronous feedback adaptive update unit. This module simulates the effects of intervention schemes in a virtual space and drives continuous optimization of system parameters through hierarchical asynchronous feedback.

[0013] This invention also provides an intelligent analysis method for power plant information data oriented towards regional power grids, which is implemented using an intelligent analysis system for power plant information data oriented towards regional power grids, and includes the following steps:

[0014] Step S1: Receive heterogeneous data from multiple sources, and form a unified semantic data view through semantic unification and spatiotemporal alignment; construct a spatial fractal self-similar index; under a three-level storage system, generate weighted increments based on the hit frequency of index branches to drive data migration and achieve storage preheating;

[0015] Step S2: Extract multi-dimensional operating features, identify power grid operation modes and dynamically select analysis models; perform semantic reasoning through knowledge graphs, and use the reasoning results as prior features to form a two-way enhanced closed loop;

[0016] Step S3: Construct a digital twin and simulate the effects of the intervention plan; collect user feedback and adjust the parameters of the data layer, analysis layer, and decision layer in a layered and asynchronous manner; set up a cross-layer coordination mechanism to maintain inter-layer compatibility.

[0017] Compared with the prior art, the beneficial effects of the present invention include:

[0018] (1) This invention automatically maps semantically equivalent but dissimilar indicators from multi-source heterogeneous data to a unified semantic layer by using a built-in power terminology ontology library in the semantic unification and spatiotemporal alignment unit and a semantic matching algorithm based on graph embedding. At the same time, multi-scale interpolation alignment is performed on low-frequency data sources using the highest frequency data source as the time reference, avoiding high-frequency information loss caused by simple resampling and effectively reducing the information loss rate of multi-source heterogeneous data in the semantic fusion process. The ontology self-expansion sub-unit generates expansion suggestions for unmatched indicators through vector embedding and clustering with existing concepts, enabling the ontology library to adaptively evolve with changes in data sources. When new power stations or new equipment are added to the regional power grid, there is no need to manually reconfigure the mapping rules, and the semantic mapping accuracy is maintained in long-term operation.

[0019] (2) This invention abstracts the spatial topology of the regional power grid into five progressive levels: regional layer, sub-regional layer, station layer, equipment layer, and sensor layer through spatial fractal self-similar index units. Each level uses a homogeneous graph structure to organize index nodes. The cross-granularity query optimizer automatically selects the execution path from the top layer or from the bottom layer according to the query granularity, realizing penetrating query for macro-situation analysis and micro-equipment location. When the topology changes, only the affected local index branches are incrementally rebuilt. The double buffer verification mechanism rebuilds the shadow index and performs hash comparison while the main index is kept online. This controls the index maintenance overhead and ensures index consistency, significantly shortening the average response time of cross-granularity queries.

[0020] (3) This invention constructs a three-level storage system of hot, warm, and cold storage units driven by access frequency prediction. The data temperature adaptive migration scheduler receives the branch hit frequency signal from the spatial fractal self-similar index unit, and uses the inverse of the branch's depth in the index tree as a weighting factor. The shallower the depth, the greater the weight. This encodes the macroscopic analysis requirements into a data warming prediction signal, enabling data serving the regional or sub-regional aggregation query to reside in the hot storage area in advance. This mechanism transforms the analysis query mode into a storage preheating action, reducing the overall storage cost while maintaining the microsecond-level read / write latency of hot data. At the same time, the cost model of the query optimizer is dynamically adjusted according to the actual data distribution through the feedback channel, reducing the jitter of query latency.

[0021] (4) This invention extracts operational feature vectors from a unified semantic data view through a situational adaptive model selection engine. These vectors include eight dimensions: voltage stability margin, frequency stability index, load curve shape, new energy output fluctuation intensity, equipment health comprehensive score, regional supply and demand balance, protection action margin, and communication link quality. An unsupervised classification framework combining a Gaussian mixture model and a hidden Markov model is used to identify the current power grid operation mode. Based on the posterior probability of the mode, an appropriate model is automatically selected from a set of preset lightweight time-series prediction models, high-precision deep neural network models, physical constraint-based optimization scheduling models, and root cause diagnosis models based on association rule mining. During mode switching, a hot-switching mechanism is used, with the old and new models running in parallel and the output weighted average transition, avoiding analytical jumps at mode boundaries and effectively controlling analytical errors.

[0022] (5) This invention constructs and maintains a weighted directed power knowledge graph through a knowledge graph semantic reasoning unit. Nodes include entity nodes and event nodes. Directed edges represent direct influence, mutual backup, electrical coupling, temporal accompaniment, and operational constraint relationships, with dynamically updated weight coefficients, thus automating root cause tracing and impact range prediction. The transfer learning cold start enhancement subunit loads a general power knowledge graph base pre-trained from multiple typical regional power grids, freezes the general semantic feature extraction layer, and only fine-tunes the domain-specific relationship classification layer, reducing the dependence of the newly deployed system on local labeled corpora. The semantic reasoning subunit performs weighted depth-first search or breadth-first search along directed edges starting from event nodes, compressing the reasoning time from minutes to seconds. More importantly, the bidirectional enhancement closed-loop mechanism enables numerical anomalies to automatically trigger semantic reasoning. The semantic reasoning results are fed back as prior features to the feature vector of the next time window to participate in pattern classification, realizing mutual enhancement between numerical analysis and semantic reasoning.

[0023] (6) This invention constructs a reduced-order state space model through a digital twin inference unit and maintains a candidate state transition matrix model library corresponding to five operating modes. Integrated prediction is performed using the posterior probability distribution of the current mode as weights, achieving a balance between computational efficiency and inference accuracy. The hierarchical asynchronous feedback adaptive update unit adjusts the data layer storage upgrade and downgrade thresholds in the first cycle, adjusts the analysis layer mode classification parameters and knowledge graph edge weights in the second cycle, and calibrates the digital twin state transition matrix in the third cycle. The duration of the three cycles increases progressively, respectively matching the time scales of rapid changes in data access patterns, slow changes in analysis models, and the slowest changes in system dynamic characteristics. Cross-layer coordination triggers the linkage update of other layers when the parameter changes in the analysis layer or decision layer exceed the threshold, ensuring that the parameter evolution of the data layer, analysis layer, and decision layer are mutually adapted and maintaining inter-layer parameter consistency during long-term operation. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the accompanying drawings without creative effort.

[0025] Figure 1 This is a flowchart illustrating a method for intelligent analysis of power plant information data for regional power grids. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0027] Therefore, the following detailed description of embodiments of the present invention is not intended to limit the scope of the claimed invention, but merely illustrates some embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0028] It should be noted that, unless otherwise specified, the embodiments and features and technical solutions in the present invention can be combined with each other.

[0029] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0030] The present invention will be further described in detail below with reference to embodiments.

[0031] This invention provides an intelligent analysis system for power plant information data in regional power grids, which includes a data semantic unification and adaptive storage organization module, a situational awareness and multi-mode analysis and reasoning module, and a forward-looking inference and closed-loop feedback module. The three modules work collaboratively in a data layer → analysis layer → decision layer link, and form a self-optimizing closed-loop system through cross-layer feedback channels.

[0032] The data semantic unification and adaptive storage organization module, as the core of the system's data layer, transforms raw data from multi-source heterogeneous power plants into a unified semantic data view and supports efficient access for upper-layer analysis with an adaptive indexing and storage structure. This module includes a semantic unification and spatiotemporal alignment unit, a spatial fractal self-similar indexing unit, a multi-temperature storage unit driven by access frequency prediction, and a cross-unit collaborative mechanism connecting the three.

[0033] The semantic unification and spatiotemporal alignment unit has a built-in power terminology ontology library for regional power grids. It extracts the indicator name, unit of measurement, and enumerated value range for each access heterogeneous data source. It uses a semantic matching algorithm to automatically map semantically equivalent but differently expressed indicators to a unified semantic layer. Using the data source with the highest sampling frequency in the system as the time reference, it performs multi-scale interpolation alignment on low-frequency data sources. It adds multi-dimensional labels, including at least time labels, spatial coordinate labels, data source labels, quality confidence labels, and semantic type labels, to each aligned data record to form the unified semantic data view.

[0034] The spatial fractal self-similar indexing unit abstracts the spatial topology of the regional power grid into multiple progressive levels: regional layer, sub-regional layer, station layer, equipment layer, and sensor layer. The organization of index nodes at each level maintains fractal self-similarity with the overall structure. Based on the spatial range and granularity involved in the query request, the unit adaptively selects the execution path from the top layer to the bottom layer to achieve penetrating query from macroscopic analysis to microscopic positioning. When the topology changes, incremental reconstruction is only performed on the affected local index branches.

[0035] A multi-temperature storage unit driven by access frequency prediction constructs a three-tiered storage system of hot, warm, and cold storage. A data temperature adaptive migration scheduler performs cross-temperature migration based on access frequency statistics of data objects and collaborative signals from spatial fractal self-similar indexing units. Specifically, after executing a query, the spatial fractal self-similar indexing unit records the hit frequency of each index branch and sends the high-frequency branch identifier and weighted increment to the data temperature adaptive migration scheduler in real time. The weighted increment is determined by the inverse of the branch's depth in the index hierarchy; the shallower the depth, the greater the weight, thus proactively preheating the data used for macro-analysis. Based on this, the data temperature adaptive migration scheduler increases the weighted access frequency prediction value of the corresponding data object, transforming the analysis query mode into a storage preheating action.

[0036] The cross-unit collaboration mechanism connects the three units mentioned above into a functionally mutually supportive whole. The unified semantic data view output by the semantic unification and spatiotemporal alignment unit serves as the sole data input source for the spatial fractal self-similar index unit. The multi-dimensional label attached to each record directly determines its hierarchical affiliation in the fractal index. The hit frequency information of the spatial fractal self-similar index unit drives the temperature migration decision of the multi-temperature storage unit. Changes in data access latency in the multi-temperature storage unit indirectly affect the query optimizer parameters of the index unit through the feedback channel, forming an internal closed loop of the data layer: query execution → heat encoding → migration decision → access latency change → query optimization adjustment.

[0037] The situational awareness and multi-mode analysis and reasoning module, as the core of the system's analysis layer, receives the unified semantic data view and performs intelligent analysis of the regional power grid's operational status from both numerical and semantic dimensions. This module includes a situational adaptive model selection engine and a knowledge graph semantic reasoning unit, which collaborate through a bidirectional enhanced closed loop.

[0038] The adaptive model selection engine extracts multi-dimensional operational feature vectors from the unified semantic data view by time window. The feature vectors integrate electrical quantities, equipment health status, and communication quality information. Based on the feature vectors, an unsupervised classification framework is used to identify the current power grid operation mode and automatically selects the appropriate model from a variety of preset analysis models. When switching modes, a hot model switching mechanism is used to keep the old and new models running in parallel until the output of the new model stabilizes before releasing the resources of the old model to eliminate analysis jumps at the mode boundary.

[0039] The knowledge graph semantic reasoning unit constructs and maintains a weighted directed power knowledge graph. Graph nodes include entity nodes and event nodes, and directed edges represent semantic relationships such as direct impact, mutual backup, electrical coupling, temporal accompaniment, and operational constraints, with dynamically updated weight coefficients. This unit automatically extracts entity relationships from unstructured texts such as power regulations and operation and maintenance manuals to expand the graph. It dynamically updates node states and edge weights based on real-time operational data in the unified semantic data view, and performs graph search on query requests initiated by the model selection engine, outputting root cause tracing or impact range prediction results.

[0040] A bidirectional enhanced closed-loop mechanism connects the aforementioned engines and units. Among the multi-dimensional operational features extracted by the situational adaptive model selection engine, the anomaly dimension triggers the knowledge graph semantic reasoning unit to perform root cause tracing. The set of possible root causes produced by the knowledge graph semantic reasoning unit serves as semantic prior features, which are fed back to the feature extraction stage of the situational adaptive model selection engine to participate in the construction of feature vectors for the next time window. This forms an internal closed loop within the analysis layer where numerical anomalies trigger semantic reasoning, and semantic conclusions enhance numerical features.

[0041] The forward-looking simulation and closed-loop feedback module, as the core of the system's decision-making layer, receives the situational identification results output by the situational adaptive model selection engine and the semantic constraints output by the knowledge graph semantic reasoning unit. It simulates the effects of intervention schemes in virtual space and drives continuous optimization of system parameters through hierarchical asynchronous feedback. This module includes a digital twin simulation unit and a hierarchical asynchronous feedback adaptive update unit, which work collaboratively through a closed-loop link: simulation output → user decision → feedback collection → parameter update → simulation model calibration.

[0042] The digital twin simulation unit constructs a reduced-order simulation model synchronized with the real power grid as a digital twin, retaining key state variables related to operational decisions. The state transition matrix is ​​obtained from historical operational data through system identification and updated with a rolling time window. Simultaneously, it maintains a library of candidate state transition matrix models corresponding to different operational modes. This unit receives the situation identification results and semantic constraints, generates candidate intervention schemes, and uses an ensemble prediction method to weighted average the simulation results of multiple candidate matrices according to the posterior probability distribution of the current mode, outputting the prediction curve, confidence interval, and no-intervention baseline curve.

[0043] The hierarchical asynchronous feedback adaptive update unit collects users' actual decision-making behavior as feedback signals. These signals are routed to the data layer, analysis layer, and decision layer according to their target, and updated at different intervals: the first interval adjusts the up / down transition thresholds of the data layer's data temperature adaptive migration scheduler; the second interval, longer than the first, adjusts the pattern classification parameters, model selection mapping rules, and knowledge graph edge weights of the analysis layer; and the third interval, longer than the second, calibrates the state transition matrix of the digital twin. This unit also includes a cross-layer coordination trigger: when the change in the analysis layer's pattern classification parameters or the decision layer's state transition matrix exceeds a preset threshold, an additional update to the data layer's threshold is triggered to ensure that the data access pattern adapts to changes in the upper layers.

[0044] The data semantic unification and adaptive storage organization module provides the situation awareness and multi-modal analysis inference module with a unified semantic data view with quality labels, and significantly reduces data access latency for analysis queries through index-aware storage warm-up. The situation awareness and multi-modal analysis inference module, through a bidirectional enhanced closed loop, outputs situation information and operational constraints with causal explanations to the look-ahead inference and closed-loop feedback modules, providing high-value boundary conditions for candidate solution generation. The hierarchical asynchronous feedback of the look-ahead inference and closed-loop feedback modules optimizes the storage threshold of the data layer, the pattern recognition and model selection parameters of the analysis layer, and the inference model of the decision layer, respectively. Cross-layer coordination triggers maintain parameter consistency across the three levels of the system, forming a global evolutionary closed loop that runs through the data, analysis, and decision layers. These three modules functionally support each other and jointly achieve an evolutionary situational closed loop.

[0045] It should be noted that this invention constructs a regional power grid power station information intelligent analysis system with a three-layer linkage of data, analysis, and decision-making, and multi-level closed-loop self-evolution capability, breaking the bottleneck of the separation between data storage, intelligent analysis, and decision inference in traditional power monitoring systems. Through a collaborative architecture of "spatial fractal self-similar index" and "multi-temperature storage driven by access frequency prediction," it not only accelerates queries with indexes but also injects the hit frequency of index branches, weighted by hierarchical depth, into the temperature migration scheduler of the storage layer in real time. This allows shallow index queries serving macro-analysis to proactively warm up relevant data from cold storage to hot storage. Conversely, changes in storage access latency feed back into adjusting the query optimizer parameters of the index. This forms an internal closed loop within the data layer: "query execution → heat encoding → migration decision → access latency → query optimization," enabling autonomous adaptation of storage layout to the analysis mode. The situational adaptive model selection engine and the knowledge graph semantic reasoning unit are designed as a bidirectionally enhanced collaborative entity. This approach not only automatically triggers root cause tracing using the knowledge graph upon detecting anomalies in the numerical model, but more importantly, it feeds back the set of possible root causes generated by the knowledge graph as "semantic prior features" into the feature vector construction for the next time window. This breaks the traditional unidirectional "analysis first, reasoning later" serial model, forming an internal closed loop within the analysis layer where "numerical anomalies trigger semantic reasoning, and semantic conclusions enhance numerical features," enabling the system to continuously strengthen its ability to identify fuzzy faults and cascading disturbances. The layered asynchronous feedback adaptive update mechanism uses the user's actual decision regarding the deduced solution as a feedback signal, optimizing the parameters of the data, analysis, and decision layers at different time periods, with the period increasing layer by layer. The key lies in introducing a "cross-layer coordination trigger": when the model or parameters of the upper layer (analysis or decision layer) change significantly, it proactively triggers additional updates to the data layer's storage strategy. This ensures that the three layers maintain parameter consistency even when evolving at different paces, forming a global evolution closed loop throughout the entire system, enabling the system to move from "passive response" to "proactive evolution."

[0046] Furthermore, in some embodiments of the present invention, the data semantic unification and adaptive storage organization module serves as the core of the data layer in an intelligent analysis system for power plant information data of a regional power grid. It transforms multi-source heterogeneous power plant raw data into a unified semantic data view and supports efficient access for upper-layer analysis with an adaptive index and storage structure. This module includes a semantic unification and spatiotemporal alignment unit, a spatial fractal self-similar index unit, a multi-temperature storage unit driven by access frequency prediction, and a cross-unit collaborative mechanism. Each unit works collaboratively in the order of semantic unification, index construction, and hierarchical storage, and forms a closed-loop optimization within the data layer through a feedback channel.

[0047] The semantic unification and spatiotemporal alignment unit is used to eliminate inconsistencies in semantic expression and spatiotemporal reference between multi-source heterogeneous data. This unit incorporates a power terminology ontology for regional power grids. This ontology defines core concepts and their attribute mapping relationships in the form of a directed graph, including voltage level, load type, generator status, protection action type, and renewable energy station category. For each accessed heterogeneous data source, the semantic common-reference resolution subunit reads the data dictionary or protocol specification of that data source, extracts the indicator name, unit of measurement, and enumerated value range, and uses a graph embedding-based semantic matching algorithm to calculate the similarity between the extracted local indicators and the standard concepts in the ontology. Indicators with similarity exceeding a first threshold are automatically mapped; indicators with similarity between a second and first threshold are generated into a candidate mapping list for confirmation; and indicators with similarity below the second threshold are marked as needing expansion. The spatiotemporal alignment subunit uses the data source with the highest sampling frequency in the system as the time reference, performs multi-scale interpolation alignment on low-frequency data sources, and adds multi-dimensional labels, including at least a time label, spatial coordinate label, data source label, quality confidence label, and semantic type label, to each aligned data record, forming a unified semantic data view. The self-expansion subunit of the ontology clusters the indicators marked as to be expanded with existing concepts through vector embedding. If the frequency of an unknown indicator exceeds a preset number and its similarity with all existing concepts is lower than the third threshold, an ontology expansion suggestion is generated. After confirmation, it is included in the ontology library. The similarity threshold of the candidate recognition is dynamically adjusted according to the matching accuracy after expansion, so that the ontology library can adaptively evolve with changes in the data source.

[0048] The spatial fractal self-similar indexing unit, based on the principle of fractal self-similarity, abstracts the spatial topology of the regional power grid into five progressive levels: regional layer, sub-regional layer, station layer, equipment layer, and sensor layer. The organization of index nodes at each level maintains self-similarity with the overall structure; that is, each level uses the same type of graph structure to organize index nodes, with only the assignment rules for node granularity and edge weights changing with each level. After receiving a query request, the cross-granularity query optimizer parses the spatial range qualifiers and time span parameters: if the query granularity belongs to the regional or sub-regional layer, it performs an aggregation query from the top-level index node, accessing only aggregated information; if it belongs to the equipment or sensor layer, it directly locates the target node through hash mapping and backtracks upwards along the index tree to obtain the context; if both macroscopic and microscopic constraints are included, it estimates the input-output overhead and computational overhead of both the top-down aggregation and bottom-up backtracking execution paths, automatically selecting the path with lower overhead. When the power grid topology changes, the incremental update maintenance mechanism determines the smallest fractal branch affected by the change based on the electrical connectivity domain analysis results, and only performs local index reconstruction on that branch. The verification and self-repair sub-units trigger global consistency verification periodically or after major topology changes, using a double-buffered index structure: while the main index remains online, a shadow index is rebuilt based on the current topology snapshot, and the key branches of the shadow index are hashed and compared with the main index; if the number of inconsistent nodes exceeds a preset threshold, the shadow index is used to replace the main index, otherwise incremental repair is performed only on the differing branches.

[0049] Access frequency prediction-driven multi-temperature storage units construct a three-tiered storage system: hot, warm, and cold. The hot storage area uses in-memory databases or NVMe solid-state drive arrays, providing microsecond-level read / write latency; the warm storage area uses solid-state drives or conventional relational databases, providing millisecond-level latency; and the cold storage area uses high-density hard disk drives combined with columnar compressed storage, reducing unit storage costs. The data temperature adaptive migration scheduler maintains a statistical table of the actual access frequency for each data object. The statistical window length and attenuation coefficient are configurable, and it receives cooperative signals from the spatial fractal self-similar indexing unit. After executing a query, the spatial fractal self-similar indexing unit records the hit frequency of each index branch and sends the identifier of the high-frequency branch and its weighted increment to the data temperature adaptive migration scheduler in real time. The weighted increment is equal to the reciprocal of the branch's depth in the index tree; the shallower the depth, the greater the weight, indicating that the branch serves a more macroscopic analytical need, and its associated data should maintain higher access frequency. Based on this, the data temperature adaptive migration scheduler increases the weighted access frequency prediction value of the corresponding data object. When making migration decisions, for data objects located in the cold storage area, if their weighted access frequency estimate exceeds the hot threshold, they are migrated to the hot storage area; if it falls between the warm and hot thresholds, they are migrated to the warm storage area. For data objects located in the hot storage area, if their weighted access frequency estimate is lower than the cold threshold, they are first compressed into a columnar storage format and then migrated to the cold storage area; if it falls between the cold and warm thresholds, they are migrated to the warm storage area. This mechanism enables the storage strategy not only to respond to actual access patterns but also to predict macro-analysis needs and proactively warm up data through the index structure.

[0050] The cross-unit collaboration mechanism connects the three units mentioned above into a functionally mutually supportive whole. The unified semantic data view output by the semantic unification and spatiotemporal alignment unit serves as the sole data input source for the spatial fractal self-similar index unit. The multi-dimensional labels attached to each record directly determine its hierarchical position and node affiliation in the fractal index tree. The branch hit frequency information generated after the spatial fractal self-similar index unit executes a query is sent in real-time, unidirectionally to the data temperature adaptive migration scheduler of the multi-temperature storage unit driven by access frequency prediction. This scheduler adjusts the weighted access frequency prediction of the corresponding data object, thus encoding the branch's analytical heat as the basis for data temperature migration. Changes in the actual data access latency in the multi-temperature storage unit indirectly affect the cross-granularity query optimizer parameters of the index unit through feedback channels: when the access latency decreases after a certain type of data migrates from a cold zone to a hot zone, the query optimizer adjusts the cost estimation model parameters of the index branch corresponding to that type of data, making subsequent query path selection more inclined towards that branch. This forms a complete closed loop from query execution to heat encoding, then to migration decision, which in turn affects access latency, ultimately impacting query optimization adjustments. The three units support and interact with each other functionally.

[0051] It should be further explained that, through semantic coreference resolution and spatiotemporal alignment, the information loss rate of multi-source heterogeneous data is reduced, and the ontology self-expansion mechanism ensures that the semantic mapping accuracy is maintained in long-term operation; the spatial fractal self-similar index unit realizes penetrating query from macroscopic analysis to microscopic positioning, and the incremental update and double buffer verification mechanism ensure index consistency while controlling index maintenance overhead; the multi-temperature storage unit, through index-driven active preheating, enables the data involved in macroscopic analysis queries to reside in the hot zone in advance, reducing the overall storage cost while maintaining microsecond-level latency for hot data; the cross-unit collaborative closed loop enables semantic expansion, index hit and storage migration to work together, and the system can automatically adjust its internal strategies according to changes in data access patterns, achieving continuous optimization in long-term operation.

[0052] Furthermore, in some embodiments of the present invention, the situational awareness and multi-mode analysis reasoning module serves as the core of the analysis layer of the intelligent analysis system for power plant information data in a regional power grid. It receives the unified semantic data view output by the data semantic unification and adaptive storage organization module, and performs intelligent analysis of the regional power grid's operating status from both numerical and semantic dimensions. This module includes a situational adaptive model selection engine and a knowledge graph semantic reasoning unit. The two collaborate through a two-way enhancement loop—where numerical anomalies trigger semantic reasoning and semantic conclusions enhance numerical features—forming a logical link within the analysis layer.

[0053] The situation-adaptive model selection engine is used to perceive the power grid's operational status in real time from a unified semantic data view and automatically select an appropriate analysis model based on the identified operational mode. This engine includes a situation feature extractor, an operational mode classifier, and a dynamic model selector.

[0054] The situation feature extractor extracts multi-dimensional operational feature vectors from a unified semantic data view using a configurable fixed time window. These feature vectors include at least eight dimensions: voltage stability margin, frequency stability index, load curve shape, intensity of renewable energy output fluctuations, comprehensive equipment health score, regional supply-demand balance, protection action margin, and communication link quality. Voltage stability margin is represented by the root mean square of the deviation between the voltage amplitude and rated value of each bus; frequency stability index is represented by the time integral of the absolute value of the instantaneous deviation of the system frequency from the standard frequency; load curve shape is composed of peak-to-valley difference rate, load factor, and peak occurrence time; renewable energy output fluctuation intensity is represented by the variance of the rate of change of wind and solar power generation on a minute-level time scale; comprehensive equipment health score is obtained by weighted fusion of the equivalent values ​​of the main transformer oil chromatographic analysis and partial discharge monitoring values; regional supply-demand balance is represented by the ratio of the real-time difference between the total regional power generation and total load power to the regional rated capacity; protection action margin is represented by the relative distance between the measured electrical quantities and setting values ​​of each protection channel; and communication link quality is represented by the statistical values ​​of data acquisition success rate and end-to-end transmission delay. The raw data for each of the above dimensions all come from a unified semantic data view, and the extractor obtains real-time input by subscribing to the incremental data stream of this view.

[0055] The operation mode classifier receives the feature vectors output by the situation feature extractor and uses an unsupervised classification framework combining a Gaussian mixture model (GMM) and a Hidden Markov Model (HMM) for classification. The GMM models the probability distribution of the feature vectors, dividing the continuous feature space into several Gaussian components. The HMM receives the output sequence of the GMM and captures the state transition patterns of the feature vectors over time. The classifier classifies the power grid operation status within the current time window into one of the following: normal stable operation mode, strong fluctuation mode of new energy sources, high load mode during peak summer season, planned equipment maintenance mode, or fault occurrence or recovery mode. It outputs the posterior probability of each mode, using the mode with the highest posterior probability as the current identification result, while retaining the second highest probability mode as a candidate.

[0056] The dynamic model selector has a pre-built set of analysis models, including a lightweight time-series forecasting model, a high-precision deep neural network model, a physical constraint-based optimization scheduling model, and a root cause diagnosis model based on association rule mining. The selector chooses a model based on the current mode output by the operating mode classifier and its posterior probability: when the mode is a normal, stable operating mode and the posterior probability exceeds a first threshold, the lightweight time-series forecasting model is activated for medium- to long-term trend analysis, and the call frequency of other models is reduced; when the mode is a highly volatile renewable energy mode, the high-precision deep neural network model is activated for ultra-short-term power forecasting, and the hot update frequency of model parameters is increased; when the mode is a peak summer load mode, the physical constraint-based optimization scheduling model is activated, focusing on regional load distribution analysis and demand-side response assessment; when the mode is a fault occurrence or recovery mode, the root cause diagnosis model based on association rule mining is activated, and the query priority of the knowledge graph semantic reasoning unit is increased to the highest level. When switching modes, the selector adopts a model hot-switching mechanism: instead of immediately terminating the original model instance, the old and new models are kept running in parallel. The new model continuously receives data streams and outputs results. Only when the confidence level of the output of the new model exceeds the preset threshold in three consecutive time windows is the computing resources of the old model released. During the switching period, the front-end analysis results adopt a weighted average of the outputs of the old and new models. The weight coefficients transition linearly with time to eliminate analysis jumps at mode boundaries.

[0057] The knowledge graph semantic reasoning unit constructs and maintains a weighted directed power knowledge graph, providing semantic-level causal reasoning capabilities for numerical analysis. Nodes in the graph are divided into entity nodes and event nodes. Entity nodes include physical entities such as power plants, equipment units, transmission lines, busbars, load centers, and protection devices, as well as abstract entities such as dispatching rules and operating procedures. Event nodes include operational events such as equipment tripping, protection actions, voltage exceeding limits, frequency exceeding limits, and active power imbalance. Directed edges represent semantic relationships between nodes, including direct influence relationships, mutual backup relationships, electrical coupling relationships, temporal association relationships, and operational constraint relationships. Each directed edge is accompanied by an initial weight coefficient assigned by power domain expert rules.

[0058] This unit includes an entity relationship automatic extraction subunit, a dynamic update subunit, a semantic reasoning subunit, and a transfer learning cold start enhancement subunit.

[0059] The automatic entity relation extraction subunit adopts a relation extraction architecture based on a pre-trained language model, using power regulations documents, operation and maintenance manuals, and historical accident reports as input corpora. This subunit performs the following operations: it performs sentence segmentation and dependency parsing on the input corpus to identify sentences containing two or more power entity nouns; it inputs the identified entity noun pairs and their context into a pre-trained language model fine-tuned with power domain corpora, outputting the semantic relation category and confidence score between entity pairs; it directly loads triples with confidence scores exceeding a first threshold into the knowledge graph, stores triples between the first and second thresholds in a candidate queue for manual sampling and verification before loading, and discards triples below the second threshold.

[0060] The dynamic update subunit uses real-time operational data from the unified semantic data view as a source of structured facts to continuously update the state attributes of entity nodes and the weight coefficients of edges in the graph. When the health score of a device drops below the threshold, the state attribute of that device node is updated from normal to warning and a warning timestamp is added; when a device malfunctions, the weight coefficients of the edges directly affecting downstream affected nodes from that device node are increased; when the device operates stably for a period of time after the warning, the weight coefficients of the relevant edges are gradually reduced.

[0061] The semantic reasoning subunit receives query requests from the dynamic model selector. For root cause tracing queries, a weighted depth-first search is performed in the reverse direction along the directed edges, starting from the event node. The upper limit of the search depth is configurable. The output consists of several entity nodes with the highest weights along the path, serving as a set of possible root causes. Each root cause is accompanied by a cumulative weight value as its confidence level. For impact range prediction queries, a breadth-first search is performed in the forward direction along the directed edges, starting from the faulty equipment node. All load centers and downstream equipment nodes visited during the search are output as the impact range and sorted by electrical distance.

[0062] The transfer learning cold start enhancement subunit is used to provide initial inference capabilities when local corpora are scarce in newly deployed systems. This subunit loads a general power knowledge graph base pre-trained from multiple typical regional power grid regulations documents and operation logs, and adopts a parameterized adaptation method: freezing the first few layers of the base model responsible for general semantic feature extraction, and only fine-tuning the last few layers responsible for domain-specific relation classification using a small number of local labeled samples; after fine-tuning, the node and edge weights of the local knowledge graph are initialized to the values ​​after knowledge transfer from the base; as local corpora accumulate, more base layers are gradually unfrozen to participate in training, enabling the local graph to transition from general knowledge to proprietary knowledge.

[0063] The aforementioned engine and units collaborate through a bidirectional enhanced closed-loop mechanism. This closed loop comprises two pathways. The first pathway is a numerically driven semantic pathway: when the value of a certain dimension in the feature vector output by the situation feature extractor exceeds a preset threshold, the system automatically generates a semantic reasoning trigger signal, which, along with the identifier of the abnormal dimension, is sent to the semantic reasoning subunit. The semantic reasoning subunit selects the corresponding root cause tracing query template based on the abnormal dimension, performs a graph search, and returns a set of possible root causes. The second pathway is a semantically enhanced numerical pathway: the set of possible root causes output by the semantic reasoning subunit is fed back to the situation feature extractor and added as an additional feature dimension to the feature vector of the next time window. This includes Boolean features indicating whether an entity is in the root cause set and continuous features representing its cumulative weight in the root cause set. Its semantic prior features and original numerical features are input together into the running pattern classifier, making pattern recognition simultaneously constrained by numerical anomalies and semantic causality. This closed loop realizes the cognitive iteration of recognition, reasoning, and enhanced recognition.

[0064] It should be further explained that the situational adaptive model selection engine can identify multiple operating modes of the power grid and automatically adapt the analysis model through multi-dimensional operational feature extraction and unsupervised classification. The model hot switching mechanism avoids analysis jumps during mode conversion. The knowledge graph semantic reasoning unit realizes the automation of root cause tracing and impact range prediction through automatic extraction and dynamic updating. The transfer learning cold start enhancement subunit reduces the dependence of new scenario deployment on labeled corpora. The bidirectional enhancement closed-loop mechanism enables numerical analysis and semantic reasoning to enhance each other, improving the accuracy and robustness of the overall analysis.

[0065] Furthermore, in some embodiments of the present invention, the forward-looking simulation and closed-loop feedback module serves as the core of the decision-making layer of the intelligent analysis system for power plant information data in a regional power grid. It receives the situational awareness and multi-mode analysis reasoning module's output of situational identification results and semantic constraints, simulates the effects of intervention schemes in virtual space, and drives continuous optimization of system parameters through a hierarchical asynchronous feedback mechanism. This module includes a digital twin simulation unit and a hierarchical asynchronous feedback adaptive update unit. The two work collaboratively through a closed-loop link of simulation output, user decision-making, feedback acquisition, parameter updates, and simulation model calibration, forming a global evolutionary closed loop within the decision-making layer and throughout the data and analysis layers.

[0066] The digital twin simulation unit constructs a lightweight simulation environment synchronized with the real regional power grid to predict the effects of different intervention schemes before control measures are implemented. This unit includes a digital twin builder, a candidate intervention scheme generator, a multi-scenario integrated simulation unit, and a comparative evaluation and recommendation unit.

[0067] The digital twin builder establishes a digital twin based on a reduced-order state-space model. The state vector includes the voltage amplitude and phase angle of each bus, the active and reactive power flow of each branch, the output setpoint and actual output of each generator unit, the state of charge of each energy storage unit, and the regulation margin of each controllable load. The state transition matrix adopts a subspace identification algorithm, estimating the system matrix, input matrix, and output matrix of the state-space model using the input and output data of the regional power grid under normal operation. The inputs are control quantities such as generator active power setpoint, energy storage charging and discharging commands, and load shedding ratios, while the outputs are controlled quantities such as voltage, frequency, and power flow. The state transition matrix is ​​re-identified using a rolling time window to adapt to the slow changes in the power grid's operating characteristics. The builder also maintains a model library containing multiple candidate state transition matrices. Each candidate matrix corresponds to one of five operating modes: normal stable operation, strong fluctuations in new energy sources, high load during peak summer season, planned equipment maintenance, and fault occurrence or recovery. Each is identified through independent historical data fragments under the corresponding mode, enabling the sandbox to call the matching state transition matrix under different operating modes.

[0068] The candidate intervention scheme generator receives two types of input from the situational awareness and multi-modal analysis reasoning module: the current operating mode and analysis conclusions output by the situational adaptive model selection engine, and the operational constraints output by the knowledge graph semantic reasoning unit. The generator has a pre-defined action space, with action types including adjusting the generator setpoint's active power output, changing the energy storage unit's charging and discharging power, reducing a certain proportion of the load at the load center, switching capacitor or reactor branches, and adjusting the tap position of the on-load tap changer. Starting from the current operating state and using the operational constraints as hard boundaries, the generator employs a heuristic search algorithm to generate a set of candidate intervention schemes, each accompanied by a label describing the expected control objective.

[0069] The multi-scenario integrated inference engine employs an ensemble prediction method for inference. For each candidate intervention scheme, the inference engine selects the top few candidate state transition matrices with the highest probabilities from the model library, using the posterior probability distribution of the current operating mode as weights. Each matrix independently infers the state sequence for multiple future time steps, with the step size adaptively selected based on whether the problem is voltage or frequency stability. The inference results of each matrix are weighted and averaged according to their normalized weights in the posterior probability distribution to generate an ensemble prediction curve. Simultaneously, the inference engine performs multiple Monte Carlo runs by applying random perturbations to the state transition matrices, outputting the confidence interval of the prediction curve. The inference engine also calculates the dispersion between the inference results of each matrix as an ensemble confidence index. When the ensemble confidence index is lower than a preset threshold, a suggestion for manual review is indicated in the output. After the inference is completed, the engine outputs the predicted change curves of each state variable, the final values ​​of key performance indicators, and the corresponding uncertainty intervals. It also outputs a baseline curve assuming no intervention is applied.

[0070] The comparative evaluation and recommender assigns a weighted comprehensive score to each candidate solution based on the improvement in key performance indicators (KPIs). The improvement in a KPI is defined as the positive change in the predicted final value of a KPI after intervention relative to the baseline curve without intervention; the weighting coefficients are preset by the user or learned through historical feedback. The evaluator sorts the comprehensive scores in descending order, marks the top three solutions with the highest scores as recommended solutions, and presents the predicted curves, confidence intervals, KPI comparison tables, and the baseline curve without intervention for each solution in a visual format.

[0071] The hierarchical asynchronous feedback adaptive update unit collects user decision-making behavior during actual scheduling as feedback signals, which are then applied to the system's data layer, analysis layer, and decision layer parameters according to different update cycles, and a cross-layer coordination mechanism is set up. This unit includes a feedback signal collector, a hierarchical parameter updater, and a cross-layer coordination trigger.

[0072] The feedback signal collector records three types of user decision-making behaviors: marking the adoption or rejection of the analysis conclusions; recording the selection of recommended solutions and the parameter differences before and after modification; and the manual correction amount of the predicted curve or key indicator value of the inference results. The above feedback signals are organized into timestamped sample pairs, in a format that shows the correspondence between system output and user decisions.

[0073] The hierarchical parameter updater routes feedback sample pairs to three different level update sub-modules according to the target module, using different update cycles. The first level is the data layer update, executed in the first cycle. It summarizes the query latency records in the feedback samples and the user's response time to the analysis results into the statistical pool, which is used to adjust the hot and cold threshold parameters of the data temperature adaptive migration scheduler in the data semantic unification and adaptive storage organization module. If a certain type of data frequently appears in the queries corresponding to the analysis conclusions adopted by users, but is currently in the cold storage area and each query triggers migration, the hot threshold is lowered to allow this type of data to enter the hot storage area in advance. The second level is the analysis layer update, executed in the second cycle, which is longer than the first cycle. User rejection records of dynamic model selection results in the feedback samples are used as negative samples. An incremental expectation-maximization algorithm is used to update the mean, covariance, and weights of each Gaussian component in the Gaussian mixture model of the running mode classifier. User confirmation or denial of knowledge graph reasoning results is used to update the weight coefficients of the corresponding directed edges; confirmation increases the weight, and denial decreases it, with the weights constrained within a preset range. User evaluation of the stability of the output after model hot switching is used to adjust the window length parameter for parallel operation in the hot switching mechanism. The third level is the decision-making layer update, executed in the third cycle, which is longer than the second cycle. The deviation between the intervention plan finally selected by the user in the feedback sample and the actual system response is used to calibrate the state transition matrix of the digital twin. The calibration process adopts the recursive least squares algorithm, defining the error vector as the difference between the actual response state and the inferred predicted state. This error is used as a correction term to update the system matrix and input matrix in the form of Kalman filtering. The user's implicit rating of the recommended plan from the plan selection behavior is used as a reinforcement signal to adjust the weight coefficients of each key performance indicator in the comparison evaluator.

[0074] A cross-layer coordination trigger monitors the magnitude of changes in model parameters in the analysis layer updater. When the mean shift of any Gaussian component in the running pattern classifier exceeds a set threshold, an additional update of the data layer threshold is triggered to adapt to changes in data access patterns caused by changes in the pattern classification boundary. When the update magnitude of the decision layer state transition matrix exceeds a threshold, a rapid recalibration of the knowledge graph edge weights in the analysis layer is triggered to reflect changes in causal relationships revealed by the new system dynamic model. This cross-layer coordination mechanism enables the evolution of the data layer, analysis layer, and decision layer to adapt to each other, forming a global evolution loop that runs through all three layers.

[0075] The following connections exist between this module and the data semantic unification and adaptive storage organization module, and the situational awareness and multi-modal analysis and reasoning module: The initial state snapshot of the digital twin obtained from the data layer comes from the latest data record of the unified semantic data view, and the data temperature adaptive migration scheduler ensures that the above data is in the hot storage area when the inference is initiated; the current operating mode and its posterior probability distribution, analysis conclusions, and knowledge graph semantic constraints are obtained from the analysis layer. The above information determines the action space boundary of the candidate scheme generation and the weighting strategy of the integrated inferencer; the user feedback signal output by this module is applied to the parameters of the data layer, analysis layer, and decision layer respectively through the hierarchical asynchronous feedback update unit to realize the overall self-evolution of the system.

[0076] It should be further explained that the digital twin simulation unit enables dispatchers to rehearse the effects of different intervention schemes in a virtual environment before implementing control measures. The weighted simulation of the multi-mode model library improves the simulation accuracy compared to a single fixed model. The confidence assessment mechanism enables the system to identify the reliability boundaries of its own simulation results. The hierarchical asynchronous feedback adaptive update unit enables the parameters of the data layer, analysis layer, and decision layer to be continuously optimized according to the actual use scenario. The different update cycles of the three levels are matched with their respective technical characteristics. The cross-layer coordination trigger makes the evolution of the three layers mutually adaptable. When the mode classification boundary of the analysis layer shifts, the data layer automatically adjusts the storage threshold. When the dynamic model of the decision layer system is updated, the weight of the knowledge graph of the analysis layer is recalibrated accordingly, ensuring that the system maintains parameter consistency in long-term operation.

[0077] like Figure 1 As shown, the present invention also provides an intelligent analysis method for power plant information data oriented towards regional power grids, which is implemented using the aforementioned intelligent analysis system for power plant information data oriented towards regional power grids. The method includes the following steps:

[0078] Step S1: Receive raw power plant data from multiple heterogeneous data sources, and form a unified semantic data view with multi-dimensional labels through semantic unification and spatiotemporal alignment; based on fractal self-similarity, abstract the spatial topology of the regional power grid into multiple progressive levels and construct a spatial fractal self-similarity index, which supports cross-granularity queries; under a three-level storage system of hot, warm, and cold, generate a weighted increment based on the hit frequency of each index branch in the spatial fractal self-similarity index, the weighted increment being inversely proportional to the depth of the branch in the index level; update the weighted access frequency estimate of the data object based on the weighted increment to perform data migration between different temperature levels and achieve storage preheating;

[0079] Step S2: Extract multi-dimensional operation feature vectors from the unified semantic data view according to time windows; identify the current power grid operation mode using an unsupervised classification framework; dynamically select an appropriate model from a variety of pre-set analysis models based on the identified operation mode; and adopt a model hot-switching mechanism that allows the old and new models to run in parallel when switching modes. Construct and maintain a weighted directed power knowledge graph; use the knowledge graph to perform semantic reasoning on queries initiated by the dynamically selected model; and feed the reasoning results as semantic prior features to feature extraction in the next time window. The semantic prior features and the multi-dimensional operation feature vectors are jointly input into the unsupervised classification framework to form a two-way enhancement closed loop in which numerical anomalies trigger semantic reasoning and semantic conclusions enhance numerical features.

[0080] Step S3: Construct a digital twin using a reduced-order state-space model, receive the current operating mode and posterior probability distribution, and semantic constraints output by the knowledge graph, generate candidate intervention schemes, and use an integrated prediction method to deduce and output the prediction effect; collect user decision feedback signals, and adjust the storage rise and fall transition thresholds in the data layer processing step, the pattern recognition and model selection parameters in the analysis layer processing step, and the state transition matrix of the digital twin in a hierarchical asynchronous manner, and set up a cross-layer coordination mechanism. When the parameter changes in the analysis layer or decision layer exceed the threshold, parameter updates in other layers are triggered, forming a global evolution closed loop that runs through the data layer, analysis layer, and decision layer.

[0081] Furthermore, in some embodiments of the present invention, step S1 belongs to the data layer processing stage of a power plant information data intelligent analysis method for regional power grids, and its execution entity is the data semantic unification and adaptive storage organization module. This step receives raw power plant data from multiple heterogeneous data sources, and processes it through a series of three sub-steps: semantic unification and spatiotemporal alignment, spatial fractal self-similar index construction, and multi-temperature storage preheating based on index hit frequency weighted increments. The result is a unified semantic data view that can be efficiently accessed by the upper-layer analysis module. There are data dependencies and feedback coordination relationships among the three sub-steps: the output of semantic unification and spatiotemporal alignment determines the input data structure for index construction; the branch hit frequency generated during index construction drives the preheating decision for multi-temperature storage as a weighted increment; and changes in the access latency of multi-temperature storage adjust the cost model parameters of the index query optimizer through a feedback channel.

[0082] Step S11: Generate a unified semantic data view with multi-dimensional labels from multi-source heterogeneous data through semantic unification and spatiotemporal alignment. This sub-step takes various heterogeneous data sources connected to the regional power grid as input, including SCADA systems, PMU devices, smart meters, and meteorological monitoring equipment. Each data source has inherent differences in indicator naming, units of measurement, sampling frequency, and data format.

[0083] First, the system pre-constructs a power terminology ontology for the regional power grid. This ontology defines core concepts and their attribute mappings in the form of a directed graph, including voltage levels, load types, generator status, protection action types, and new energy station categories. For each accessed heterogeneous data source, its data dictionary or protocol specification is read, and the indicator names, units of measurement, and enumerated value ranges defined in that source are extracted. A semantic matching algorithm based on graph embedding is used to calculate the similarity between the extracted local indicators and the standard concepts in the ontology. Indicators with similarity exceeding a first threshold are automatically mapped; indicators with similarity between a second and first threshold generate a candidate mapping list for confirmation; indicators with similarity below the second threshold trigger an ontology self-expansion process, i.e., similarity clustering with existing concepts through vector embedding. If the frequency of an unknown indicator exceeds a preset number and its similarity with all existing concepts is below a third threshold, an ontology expansion suggestion is generated, confirmed, and added to the ontology, enabling the ontology to adaptively evolve with changes in data sources.

[0084] Secondly, the data source with the highest sampling frequency in the system is used as the time reference. For SCADA data with a second-level cycle, cubic spline interpolation is used between adjacent valid sampling points to generate estimates aligned with the time reference; for meteorological data with a minute-level cycle, linear interpolation is used within the sampling interval to generate aligned estimates. All data streams are mapped onto a unified time axis, forming time-aligned data records.

[0085] Finally, multidimensional labels are attached to each aligned data record. These multidimensional labels include at least a time label, a spatial coordinate label, a data source label, a quality confidence label, and a semantic type label. After the above processing, the originally scattered, heterogeneous, and time-inconsistent raw data is transformed into a unified, fully labeled structured data set, i.e., a unified semantic data view. This view serves as the data foundation for all subsequent indexing, storage, and analysis operations.

[0086] Step S12: Using the unified semantic data view output in step S11 as input, construct a spatial fractal self-similar index that supports cross-granularity queries.

[0087] First, the spatial topology of the regional power grid is abstracted into five progressive levels: regional layer, sub-regional layer, station layer, equipment layer, and sensor layer. The regional layer corresponds to the boundary range of the entire target regional power grid, storing aggregated statistical information for that region. The sub-regional layer is divided according to voltage level or geographical partition. The station layer stores key operating parameters for each individual power station. The equipment layer stores real-time status and health scores for major equipment such as transformers, circuit breakers, generator sets, and transmission lines. The sensor layer stores raw instantaneous value sampling sequences at the measurement point level. In each of these levels, the organization of index nodes maintains self-similarity with the overall structure; that is, each level uses the same type of graph structure to organize index nodes, with only the node granularity and edge weight assignment rules changing with each level.

[0088] Secondly, the optimizer receives query requests and parses the spatial scope qualifiers and time span parameters involved. The query requests use an extended form of Structured Query Language. The optimizer's decision logic is as follows: If the query granularity belongs to the region or sub-region level, the clustered query is executed starting from the top-level index node, accessing only high-level aggregated information without delving down to the device or sensor level; if the query granularity belongs to the device or sensor level, the target device node is directly located through hash mapping, and the context information of the parent node is obtained by traversing upwards along the index tree; if the query conditions contain both macro-level scope constraints and micro-level feature constraints, the optimizer estimates the input-output overhead and computational overhead of the two execution paths—clustering from the top-level down and traversing from the bottom-level up—and selects the path with lower overhead.

[0089] Furthermore, when the power grid topology changes, the system determines the smallest fractal branch affected by the changed region based on the electrical connectivity analysis results. The electrical connectivity analysis is based on the node-branch model of the power grid and uses a depth-first search algorithm to label connected components. Only the affected branches are locally indexed and reconstructed, while the remaining branches remain unchanged. The time complexity of local reconstruction is proportional to the number of nodes in the affected branches.

[0090] Finally, a verification and self-repair subunit is included to trigger global consistency checks periodically or after significant topology changes. A double-buffered index structure is adopted: while the main index remains online, a shadow index is rebuilt based on the current topology snapshot, and the key branches of the shadow index are hashed and compared with those of the main index. If the number of inconsistent nodes exceeds a preset threshold, the main index is replaced by the shadow index; otherwise, incremental repair is performed only on the differing branches.

[0091] Step S13: Storage preheating is achieved in a multi-temperature storage system based on weighted increments of index hit frequency. This sub-step runs in parallel with step S12 and has feedback coordination. Its inputs are the branch hit frequency information generated during the index query execution in step S12, and the original data records in the unified semantic data view output in step S11. This sub-step constructs a three-tiered storage system of hot, warm, and cold, and drives the active migration of data between different temperature levels through weighted increments of index hit frequency.

[0092] First, the hot storage area uses in-memory databases or NVMe solid-state drive arrays, providing microsecond-level read / write latency, with storage capacity configured as a portion of the total data volume. The warm storage area uses solid-state drives or conventional relational databases, providing millisecond-level latency. The cold storage area uses high-density hard disk drives with a columnar compressed storage format, resulting in lower unit storage costs than the hot storage area.

[0093] Secondly, the system maintains a statistical table of the actual access frequency for each data object, with a sliding time window and a configurable decay coefficient. Simultaneously, after each query, the spatial fractal self-similar index unit records the identifier of the index branch hit by the query and the branch's depth in the index tree. The system calculates a weighted increment, which is equal to the product of a preset coefficient and the reciprocal of the branch's depth. This weighted increment is sent in real-time to the data temperature adaptive migration scheduler, which updates the weighted access frequency estimate for the corresponding data object accordingly. Shallower index branches correspond to more macroscopic analytical queries, and their larger weighted increments indicate that their associated data should maintain higher access frequency.

[0094] Furthermore, the data temperature adaptive migration scheduler periodically scans the weighted access frequency estimates of data objects and executes the following decision logic: For data objects located in the cold storage area, if their weighted access frequency estimate exceeds the hot threshold, they are migrated to the hot storage area; if it is between the warm and hot thresholds, they are migrated to the warm storage area. For data objects located in the hot storage area, if their weighted access frequency estimate is lower than the cold threshold, they are first compressed into columnar storage format and then migrated to the cold storage area; if it is between the cold and warm thresholds, they are migrated to the warm storage area. The migration process employs a double-buffering mechanism: first, a complete copy of the data is written to the new storage area; after successful verification, the metadata pointer is updated; and then the copy in the old storage area is deleted to ensure uninterrupted data access during the migration process.

[0095] Finally, when the access latency of a data object is significantly reduced due to its migration to a hot zone, this latency change is recorded in the query execution statistics. The cross-granularity query optimizer periodically reads its statistics and updates the storage latency parameters in its cost model, lowering the input and output cost coefficients of the index branches corresponding to hot zone data and raising the cost coefficients of the branches corresponding to cold zone data. This makes the optimizer more inclined to use the index branches corresponding to hot zone data when selecting subsequent query paths. This feedback channel forms a closed loop between steps S12 and S13: the query generates a hit frequency, the frequency drives the data to warm up, the latency decreases after warming up, and the optimizer adjusts the cost model accordingly, making subsequent queries more inclined to use the warmed-up branches.

[0096] After sequential processing through steps S11 to S13, the original multi-source heterogeneous data is transformed into a data foundation with unified semantics, efficient indexing, and adaptive hot and cold storage, providing support for subsequent situational awareness, multi-modal analysis, and forward-looking inference.

[0097] It should be noted that there are multiple solutions in existing technologies for semantic unification and spatiotemporal alignment. Early research on power information integration adopted an XML-based middleware model, converting fields from different data sources into a standard format through predefined mapping files. However, these mapping files require manual maintenance and cannot adapt to changes in data sources. In recent years, some patents have proposed ontology-based semantic fusion methods for power data. These methods construct domain ontology and use inference engines to automatically discover mapping relationships. However, these ontology systems are statically constructed, and the ontology library cannot automatically expand when new equipment or concepts emerge, requiring manual updates by domain experts. The ontology self-expansion process in this step achieves dynamic evolution of the ontology library by embedding unknown indicators into a vector space and clustering them, identifying new concepts in real time, and generating expansion suggestions. Furthermore, the spatiotemporal alignment operation uses a multi-scale interpolation method based on the highest frequency data source. Compared to the traditional timestamp alignment approach, which simply resamples data from different frequencies to the same frequency, this method preserves the original sampling information of the high-frequency data source and reduces information loss.

[0098] In terms of index construction, existing power grid data indexing technologies mostly employ R-trees and their variants for spatial data retrieval, or B-trees for time-series data retrieval. The hierarchical structure of these indexes is pre-fixed, requiring different indexes to be built or full scans to be performed for queries of different granularities. The spatial fractal self-similar index proposed in this step utilizes the principle of self-similarity to achieve a homogeneous index structure across all levels. The cross-granularity query optimizer can automatically select the optimal execution path based on the query granularity, achieving seamless switching from macro to micro levels without the need to maintain multiple independent indexes. The electrical connectivity domain analysis in the incremental update maintenance mechanism limits the scope of influence, avoiding the overhead of full reconstruction of traditional indexes during topology changes. A double-buffered verification mechanism further ensures index consistency.

[0099] In data storage, the concept of multi-temperature storage has already been applied in the power industry. A patent proposes a cold / hot tiered storage method for power data, dividing data into cold and hot data based on historical access frequency and storing them separately. However, this approach relies solely on the access statistics of the data objects themselves and cannot detect changes in upper-level analytical query patterns. This step utilizes a storage preheating mechanism based on index hit frequency-weighted increments, using the inverse of the index branch depth as a weighting factor. This allows macro-analysis needs to predictively drive data warming before query execution. Furthermore, a feedback channel ensures that the query optimizer's cost model aligns with the actual data distribution, forming a closed-loop optimization of query, heat, latency, cost, and query again.

[0100] Regarding the coordination of the three sub-steps, existing technologies typically design and operate semantic fusion, index building, and storage management independently, lacking seamless data and feedback flow. This step, however, uses a unified semantic data view as an intermediary, directly using the output of the semantic unification sub-step as input for index building; it dynamically couples the indexing and storage sub-steps through a weighted increment of index branch hit frequency; and it adjusts the query optimizer cost model through feedback, allowing the warm-up effect of the storage sub-step to influence the query path selection of the indexing sub-step. This three-layer collaborative architecture ensures that the semantic, indexing, and storage functions of the data layer support each other and have a clear interaction relationship.

[0101] It should be further explained that, regarding data fusion quality, this step significantly reduces the information loss rate of multi-source heterogeneous data through semantic unification and spatiotemporal alignment operations. Employing an ontology-guided semantic matching algorithm combined with an ontology self-expansion process, the semantic information loss rate is controlled at a low level in scenarios involving at least four heterogeneous data sources, significantly outperforming traditional field-mapping-based methods. The spatiotemporal alignment operation uses the highest frequency data source as a benchmark, interpolating and filling in low-frequency data to ensure that the fused time series retains the original high-frequency signal characteristics within the Nyquist frequency range, avoiding the loss of high-frequency information caused by simple resampling.

[0102] In terms of data access efficiency, the spatial fractal self-similar index significantly reduces the average response time for cross-granularity queries. Switching from macroscopic statistical queries at the regional level to microscopic location queries at the equipment level results in latency controlled within the hundreds of milliseconds, and no index rebuilding is required during the switch. For regional power grids containing a large number of stations and equipment, both regional-level cluster queries and equipment-level point queries can achieve rapid responses. The incremental update maintenance mechanism ensures that the index maintenance time during power grid topology changes is proportional to the number of affected branch nodes, significantly reducing the time compared to a full rebuild. The double-buffered verification mechanism, while ensuring index consistency, restricts the triggering conditions for global rebuilding to scenarios of drastic topology changes, avoiding the risk of accumulated errors from conventional local updates.

[0103] In balancing storage cost and access performance, a storage preheating mechanism based on index hit frequency-weighted increments achieves predictive data warming. Under typical operating scenarios, the hit rate of the hot storage area is high, and the average access latency for hot data remains at the microsecond level. Compared to multi-temperature storage solutions without preheating, index-driven preheating significantly reduces the number of passive data migrations for macro-analysis queries, reducing query latency fluctuations. Overall storage costs are significantly lower than all-hot storage solutions, while the average response speed for critical queries is significantly improved compared to all-cold storage solutions.

[0104] Regarding system adaptability, the feedback channel enables the cost model of the index query optimizer to dynamically adjust according to the actual data distribution. Through online learning, the optimizer's accuracy in selecting query paths is improved. The parameters of the data temperature adaptive migration scheduler are continuously optimized through the feedback mechanism, making the storage distribution increasingly approximate the actual access pattern, adapting to the slow changes in power grid operating characteristics without manual intervention.

[0105] This step resolves the semantic conflicts of multi-source heterogeneous data, the contradiction between cross-granularity query efficiency and storage cost and access performance at the data level through the cascading collaboration and feedback loop of three sub-steps: semantic unification and spatiotemporal alignment, spatial fractal self-similar indexing, and multi-temperature storage preheating driven by index hit frequency. It provides a high-quality, high-efficiency, and adaptive data foundation for subsequent situational awareness and multi-modal analysis and reasoning, forward-looking inference and closed-loop feedback steps.

[0106] Furthermore, in some embodiments of the present invention, step S2 belongs to the analysis layer processing stage of an intelligent analysis method for power plant information data in a regional power grid, and its execution entity is the situation awareness and multi-mode analysis reasoning module. This step follows the unified semantic data view output from step S1, performs intelligent analysis on the operating status of the regional power grid from both numerical and semantic dimensions, and outputs the analysis results to step S3. Step S2 includes two sub-steps: situation adaptive model selection and knowledge graph semantic reasoning. The two are coordinated through a two-way enhancement closed loop of semantic reasoning triggered by numerical anomalies and semantic conclusions enhancing numerical features.

[0107] Step S21: Extract multi-dimensional operating features from the unified semantic data view by time window, identify the power grid operating mode and dynamically select the analysis model; This sub-step takes the unified semantic data view output by step S1 as input, perceives the power grid operating status in real time, and automatically selects the appropriate analysis model according to the identified operating mode.

[0108] First, multidimensional runtime feature vectors are extracted from the unified semantic data view using a configurable fixed time window. Let the current time window be t, and the extracted feature vector be f(t)∈R.n The feature vector includes at least eight dimensions: voltage stability margin, frequency stability index, load curve shape, intensity of renewable energy output fluctuation, comprehensive equipment health score, regional supply and demand balance, protection action margin, and communication link quality. Voltage stability margin is represented by the root mean square of the deviation between the voltage amplitude of each bus and its rated value; frequency stability index is represented by the time integral of the absolute value of the instantaneous deviation of the system frequency relative to the standard frequency; load curve shape is composed of peak-to-valley difference rate, load factor, and peak occurrence time; renewable energy output fluctuation intensity is represented by the variance of the rate of change of wind and solar power generation on a minute-level time scale; comprehensive equipment health score is obtained by weighted fusion of the equivalent values ​​of the main transformer oil chromatographic analysis and the partial discharge monitoring values; regional supply and demand balance is represented by the ratio of the real-time difference between the total regional power generation and total load power to the regional rated capacity; protection action margin is represented by the relative distance between the measured electrical quantities and setting values ​​of each protection channel; and communication link quality is represented by the statistical values ​​of data acquisition success rate and end-to-end transmission delay. The raw data for each of the above dimensions all come from a unified semantic data view, and real-time input is obtained by subscribing to the incremental data stream of this view.

[0109] Secondly, the feature vector f(t) output by the situation feature extractor is received and classified using an unsupervised classification framework that combines a Gaussian mixture model and a hidden Markov model.

[0110] Gaussian mixture models depict the probability distribution of an eigenvector as a weighted sum of K Gaussian components, with the following probability density function:

[0111]

[0112] Where p(f(t)) represents the total probability density of the power grid operation feature vector f(t) within the time window t; K represents the total number of Gaussian components (set to 5 in this scheme, corresponding to five power grid operation modes: normal stable operation, strong fluctuations in new energy sources, high load during peak summer, planned equipment maintenance, and fault occurrence or recovery); π k This represents the mixing weight of the k-th Gaussian component (all π). k The sum is 1); N(f(t)|μ k ,Σ k Let ) denote the density function of the k-th Gaussian distribution, where μ k Σ is the mean vector of this distribution (representing the central characteristic of this operating pattern). k It is the covariance matrix (representing the degree of dispersion between features under this pattern).

[0113] Hidden Markov Models (HMMs) use the output of Gaussian mixture models as the observation sequence to capture the state transition patterns of feature vectors over time. HMMs start with an initial state probability vector π.HMM State transition matrix A HMM and observation probability matrix B HMM Define , where the elements a of the state transition matrix A are... {ij} =P(q {t+1} =S j |q t =S i ), indicating that from state S i Transition to state S j The probability of q, where q t q represents the hidden state of the system at time t. t =S i This indicates that the power grid is in the i-th operating mode at time t.

[0114] Given an observation sequence O={f(1),f(2),...,f(T)}, using the forward variable α t(i) and backward variable β t(i) Calculate the posterior probability of being in each state at each time t:

[0115] γ t(i) =P(q t =S i |O,λ)

[0116] Where λ=(π) HMM A HMM B HMM ) represents the complete parameter set of the Hidden Markov Model, where π HMM Let A be the initial state probability vector. HMM Let B be the state transition matrix. HMM The observation probability matrix is ​​given. The running pattern classifier outputs the posterior probabilities P(m) of the five patterns. k |f(t)), k=1,...,5, take the pattern corresponding to the highest posterior probability as the current recognition result, and retain the second highest probability pattern as a candidate.

[0117] Furthermore, an internal set of pre-defined analysis models M={M1,M2,M3,M4} is provided, where M1 is a lightweight time-series prediction model, M2 is a high-precision deep neural network model, M3 is a physical constraint-based optimization scheduling model, and M4 is a root cause diagnosis model based on association rule mining. Model selection is based on the current mode output by the operating mode classifier and its posterior probability: when m = normal stable operating mode and P(m|f(t)) exceeds the first threshold, M1 is activated for medium- to long-term trend analysis, and the call frequency of other models is reduced; when m = strong fluctuation mode of new energy, M2 is activated for ultra-short-term power prediction, and the hot update frequency of model parameters is increased; when m = peak summer high load mode, M3 is activated, focusing on regional load distribution analysis and demand-side response assessment; when m* = fault occurrence or recovery mode, M4 is activated, and the query priority of the knowledge graph semantic reasoning unit is increased to the highest level.

[0118] When switching modes, a hot-swap model mechanism is used. The original model is M. old The new selection model is M. new The output y during the switching period transition (t) is obtained by weighted averaging the outputs of the old and new models:

[0119] y transition (t)=(1-α(t))·y old( t)+α(t)·y new (t)

[0120] Among them, y transition (t) represents the front-end analysis results released to the public. By weighting and fusing the outputs of the old and new models, the continuity of the output curve is ensured during the switching process, and no abrupt changes occur. α(t) is the time-varying weight coefficient of the new model output, reflecting the contribution ratio of the new model in the mixed output. Its value range is [0,1], and it is calculated as α(t)=min(t / T). transition ,1), where: t is the cumulative time calculated from the time of the mode switch trigger, T transition The preset transition time window length (typically set to 3 to 5 times the model's single inference cycle; for example, if the model's inference cycle is 1 second, then T...) transition (Set to 5 seconds); the min(...) operation ensures that the weight coefficients stop increasing after reaching 1, meaning that after the transition period, the new model outputs completely independently; y old(t) Indicates the original model M old The analysis output value at time t shows that the model has been validated in the current operating mode and exhibits high output stability; y new(t) Represents the new choice model M newThe analysis output value at time t shows that the model is running in parallel in the background and receiving the same data stream, but its output has not yet been validated over a long period of time under current operating conditions. Only when M... new M is only released after the output confidence level exceeds the preset threshold for three consecutive time windows. old This mechanism eliminates analytical jumps at mode boundaries, thus conserving computational resources.

[0121] Step S22: Perform semantic reasoning using the weighted directed power knowledge graph, and feed the reasoning results back to the feature extraction of the next time window as semantic prior features; this sub-step constructs and maintains the weighted directed power knowledge graph, provides semantic causal reasoning capabilities for numerical analysis, and forms a bidirectional enhanced closed loop with step S21.

[0122] First, the knowledge graph is represented by a directed graph G=(V,E,W), where V is the set of nodes, E is the set of directed edges, and W is the set of edge weights. Nodes are divided into entity nodes V. entity and event node V event Entity nodes include physical entities such as power plants, equipment units, transmission lines, busbars, load centers, and protection devices, as well as abstract entities such as scheduling rules and operating procedures. Event nodes include operational events such as equipment tripping, protection actions, voltage exceeding limits, frequency exceeding limits, and active power imbalance. A directed edge e=(u,v)∈E represents the semantic relationship from node u to node v. Relationship types include direct influence relationships, mutual backup relationships, electrical coupling relationships, temporal accompaniment relationships, and operational constraint relationships. Each directed edge e is accompanied by a weight coefficient w. e ∈[0,1], with initial values ​​assigned by power industry expert rules.

[0123] Secondly, a relation extraction architecture based on a pre-trained language model is adopted, using power regulations documents, operation and maintenance manuals, and historical accident reports as input corpora. Sentence segmentation and dependency parsing are performed on the input corpora to identify sentences containing two or more power entity nouns. The identified entity noun pairs and their context are input into a pre-trained language model fine-tuned with power domain corpora, outputting the semantic relationship category and confidence score between entity pairs. Triples with confidence scores exceeding a first threshold are directly loaded into the knowledge graph, triples between the first and second thresholds are stored in a candidate queue and loaded after manual sampling verification, and triples below the second threshold are discarded.

[0124] Furthermore, real-time operational data from the unified semantic data view is used as a source of structured facts to continuously update the state attributes of entity nodes and the weight coefficients of edges in the graph. When the health score h(t) of a certain device drops to the threshold τ... hWhen the following occurs, update the status attribute of the device node from normal to alert and attach an alert timestamp; when a device fails, increase the weight coefficient w of the direct impact edge between that device node and its downstream affected nodes. e (t+1)=min(w e (t)+Δw up ,1), where w e (t) represents the weight coefficient of edge e before the update, characterizing the strength of this direct influence relationship before the failure occurs, and its value ranges from [0,1]; w e (t+1) represents the weight coefficient of edge e after the failure event triggers the update, reflecting the degree to which the relationship is strengthened after the failure occurs; Δw up This represents the preset weight increment step size, typically ranging from 0.1 to 0.3. The selection of this step size is based on the following: a step size that is too small will result in a slow reinforcement speed of the fault propagation path, failing to promptly increase the priority of related inference paths after a fault occurs; a step size that is too large may cause a single fault event to excessively alter the graph structure, reducing the smoothness of the graph's gradual learning of multiple events. `min(·,1)` represents the minimum value operation, ensuring that the updated weight does not exceed 1. A weight of 1 indicates that the direct influence relationship has the highest strength in the knowledge graph, meaning the certainty of fault propagation to downstream nodes is extremely high. After the device operates stably for a period of time following the warning, the weight coefficient w of the relevant edges is gradually reduced. e (t+1)=max(w e (t)-Δw down ,w min ), where Δw down To preset the decay step size, w min This is the lower limit of the weight.

[0125] For newly deployed systems with limited local corpora, a transfer learning cold start enhancement method is adopted: a general power knowledge graph base G, pre-trained from multiple typical regional power grid regulations documents and operation logs, is loaded. base The first few layers of the base model responsible for general semantic feature extraction are frozen, and only the last few layers responsible for domain-specific relation classification are fine-tuned using a small number of local labeled samples. After fine-tuning, the node and edge weights of the local knowledge graph are initialized to the values ​​after the base knowledge transfer. As the local corpus accumulates, more base layers are gradually unfrozen to participate in training, so that the local graph transitions from general knowledge to proprietary knowledge.

[0126] Then, it receives query requests from the dynamic model selector, and for root cause tracing queries, it uses event nodes v. event Starting from the directed edge, perform a weighted depth-first search in the opposite direction, with an upper limit D for the search depth. max Configurable. During the search process, the cumulative weight of each path is calculated. Output the top N (N is a natural number) entity nodes with the largest cumulative weights as the possible root cause set R={v1,v2,...,v... N Each root cause is accompanied by a normalized confidence score. Where, i∈N, j∈N. For impact range prediction queries, take the faulty device node v as an example. fault Starting from the positive direction of the directed edge, perform a breadth-first search, outputting all load centers and downstream device nodes visited during the search as the scope of influence and sorting them by electrical distance.

[0127] Finally, a bidirectional enhancement loop is formed, which contains two pathways. The first pathway is the numerically driven semantic pathway: in the feature vector f(t) output by the situation feature extractor, when the value of a certain dimension i exceeds the preset threshold range, i.e., f... i (t)>τ i _upper or f i (t)<τ i _lower, where τ i _upper represents the preset upper limit value of the i-th feature dimension within the normal operating range; τ i `_lower` represents the preset lower limit value of the i-th feature dimension within the normal operating range; the system automatically generates a semantic inference trigger signal, which, along with the identifier i of the abnormal dimension, is sent to the semantic inference subunit; the semantic inference subunit selects the corresponding root cause tracing query template according to the abnormal dimension, performs a graph search, and returns the set of possible root causes R(t). The second path is the semantic enhancement numerical path: the set of possible root causes R(t) output by the semantic inference subunit is fed back to the situation feature extractor and added as an additional feature dimension to the feature vector of the next time window. Define the semantic enhancement feature vector f. aug (t+1) is the concatenation of the original numerical feature vector f(t+1) and the semantic prior feature vector s(t):

[0128] f aug (t+1)=[f(t+1);s(t)]

[0129] Where s(t) is an M-dimensional vector, M = |V entity | represents the number of entity nodes in the knowledge graph, s i (t)=c(v i )·I(v i ∈R(t)), c(v i ) is entity v i The normalized confidence score in the root cause set, I(·), is the indicator function. Its semantic prior features and original numerical features are jointly input into the running pattern classifier, so that pattern recognition is simultaneously constrained by numerical anomalies and semantic causality. This closed loop realizes the cognitive iteration of recognition, reasoning, and enhanced recognition.

[0130] After the collaborative processing of steps S21 and S22, the numerical information in the unified semantic data view is transformed into operational situation awareness results with semantic enhancement, including the current operational mode and its posterior probability, the adapted analysis model and its output, and the root cause or scope of influence derived from knowledge graph reasoning. The results are output to step S3 to provide situation boundaries and semantic constraints for the scheme inference of digital twins.

[0131] It should be noted that, regarding model selection, existing technologies include power equipment condition assessment methods based on multi-mode data fusion, and schemes that use machine learning to classify the power grid operating status and then select analysis strategies. However, these models are mostly based on offline training and static deployment, and the original model stops immediately when switching modes, posing a risk of analysis jumps. The model hot-switching mechanism in this step allows the old and new models to run in parallel until the output of the new model stabilizes before releasing the old resources. During the switching period, the output uses a weighted average transition, where the transition weight α(t) changes linearly with time, eliminating analysis jumps at the mode boundary.

[0132] In terms of knowledge graph construction, there are existing methods for power grid dispatch auxiliary decision-making and fault handling strategy generation based on knowledge graphs. However, their graph construction mainly relies on manual annotation and offline updates, lacking the ability to dynamically update node states and edge weights from real-time data. This step uses formula w e (t+1)=min(w e (t)+Δw up ,1) and w e (t+1)=max(w e (t)-Δw down ,w min This enables adaptive adjustment of weights, allowing the knowledge graph to reflect the real-time status of the power grid operation.

[0133] Regarding cold start, existing knowledge graph construction techniques generally face the problem of insufficient labeled corpora when applied to new fields. The transfer learning cold start enhancement method in this step loads a pre-trained general power knowledge graph base and performs parameterized fine-tuning, enabling the newly deployed system to still provide basic reasoning capabilities when corpora are scarce.

[0134] In the fusion of numerical and semantic data, existing technologies mostly involve independent operation or loosely coupled integration of the two, lacking automated bidirectional information interaction. The bidirectional enhanced closed-loop mechanism established in this step, through the concatenation of formula f... aug (t+1)=[f(t+1);s(t)] injects the semantic reasoning result as a priori feature into the numerical feature vector, so that numerical anomalies automatically trigger semantic reasoning, and the semantic reasoning result is automatically fed back to the numerical feature extraction, forming a cognitive iterative closed loop of recognition, reasoning, and enhanced recognition.

[0135] It should be further explained that, in terms of operating condition adaptability, the situational adaptive model selection engine, through real-time extraction of eight-dimensional feature vectors and a classification framework combining Gaussian mixture models and hidden Markov models, can accurately identify five typical operating modes of the regional power grid. Compared to existing systems using a single fixed model, the average absolute percentage error of the analysis results is reduced by 22% to 35% under different operating conditions. The model hot-switching mechanism eliminates analytical jumps during mode transitions through weighted average transitions, ensuring the continuity and stability of situational awareness.

[0136] In terms of semantic reasoning capabilities, the knowledge graph semantic reasoning unit automates fault root cause tracing and impact range prediction through automatic entity relation extraction, dynamic updating, and weighted graph search. Path accumulation weights are used in root cause tracing.

[0137] The path set `path` represents the complete sequence of directed edges traversed from the query origin to the target node in a weighted directed power knowledge graph. For root cause tracing queries: the origin is the event node that triggered the query (e.g., bus voltage exceeding limits), and the path extends in the reverse direction of the directed edges (i.e., upstream of the causal relationship). For impact range prediction queries: the origin is the faulty equipment node, and the path extends in the forward direction of the directed edges (i.e., downstream of causal propagation). The edge index `e` represents a specific directed edge in the path set `path` (i.e., the semantic connection between two adjacent nodes). The edge weight `w`... e This represents the current weight coefficient of the directed edge e, with a value ranging from [0,1]. This weight quantifies the strength of the semantic relationship between nodes (e.g., the strength of "direct influence") and is dynamically updated. When a device malfunctions, the weight of the affected edge will be increased by w. e (t+1)=min(w e (t)+Δw up 1) Once the equipment is running smoothly, the weights of the relevant edges will gradually decrease. e (t+1)=max(w e (t)-Δw down ,w min ). Path cumulative weight W path This represents the sum of the weights of all edges along the path. Its physical meaning is the overall confidence strength of the inference link: the larger the value, the stronger the causal relationship propagating along the link from the starting point, and the more credible the inference conclusion.

[0138] Regarding the depth of numerical and semantic fusion, the bidirectional enhancement closed-loop mechanism uses formula f augThe expression (t+1)=[f(t+1);s(t)] achieves the fusion of semantic prior features and numerical features, subjecting pattern recognition to both numerical anomalies and semantic causality. This closed loop enables the system's analytical capabilities to transcend single-dimensional numerical statistics or semantic reasoning, filling the gap in the deep integration of numerical intelligence and semantic intelligence in existing technologies.

[0139] In terms of system deployment and evolution capabilities, the transfer learning cold start enhancement method reduces the dependence of new regional power grid deployments on locally labeled corpora, shortening the online preparation cycle. As runtime increases and local corpora accumulate, the knowledge graph gradually transitions from general knowledge to proprietary knowledge, achieving domain adaptation. The pre-built model library of the dynamic model selector can be replaced and expanded online based on actual operational results, without requiring downtime maintenance.

[0140] Furthermore, in some embodiments of the present invention, step S3 belongs to the decision-making layer processing stage, and its execution entity is the forward-looking deduction and closed-loop feedback module. This step inherits the unified semantic data view output from step S1 and the situation identification results and semantic constraints output from step S2. It uses a digital twin model as a carrier to perform forward-looking deduction of intervention schemes and drives the continuous optimization of parameters in the data layer, analysis layer, and decision layer through a hierarchical asynchronous feedback mechanism. There are causal dependencies and feedback synergies among the three sub-steps: the digital twin deduction relies on the situation and semantic constraints provided upstream to generate candidate schemes; user decision feedback acts as a supervision signal to update parameters at the three levels respectively; and the cross-layer coordination mechanism ensures the mutual adaptation of parameter evolution at each level.

[0141] Step S31: Receive situational and semantic constraints using a digital twin model, generate candidate intervention schemes, and extrapolate and predict the effects.

[0142] This sub-step takes the real-time state snapshot in the unified semantic data view output by step S1, the current operating mode and posterior probability distribution output by step S2, and the set of operational constraint relationships output by the knowledge graph semantic reasoning unit as input.

[0143] Construct and maintain a digital twin model, which is described by a reduced-order state-space equation, and its mathematical expression is:

[0144] x(t+1)=A sys x(t)+B sys u(t)

[0145] y(t)=C sys x(t)+D sys u(t)

[0146] Where x(t) is the state vector, including the voltage magnitude and phase angle of each bus, the active and reactive power flow of each branch, the setpoint and actual output of each generator unit, the state of charge of each energy storage unit, and the regulation margin of each controllable load; u(t) is the control input vector, including the generator active power setpoint increment, energy storage charging and discharging commands, and load shedding ratio; y(t) is the output vector, including voltage, frequency, and power flow. System matrix A sys Input matrix B sys Output matrix C sys and direct transfer matrix D sys It is estimated from historical running data through a subspace identification algorithm and re-identified using a rolling time window method.

[0147] The digital twin model maintains a model library containing multiple candidate state transition matrices. Each candidate matrix corresponds to one of five operating modes: normal stable operation, strong fluctuations in new energy sources, high load during peak summer season, planned equipment maintenance, and fault occurrence or recovery. Each is identified through independent historical data fragments under the corresponding mode.

[0148] The system receives the current operating mode and analysis conclusions output from step S2, as well as the set of operational constraint relationships output by the knowledge graph semantic reasoning unit. It has an internally preset action space, with action types including adjusting the generator set's active power output, changing the energy storage unit's charging and discharging power, reducing a certain proportion of the load at the load center, switching capacitor or reactor branches, and adjusting the tap position of the on-load tap-changing transformer. Starting from the current operating state and using the operational constraint relationships as hard boundaries, a heuristic search algorithm generates a set of candidate intervention schemes, each accompanied by a label describing the expected control objective.

[0149] For each candidate intervention scheme, the top few candidate state transition matrices with the highest probabilities are selected from the model library and weighted using the posterior probability distribution of the current operating mode as weights for weighted inference. The inference results of each matrix are then weighted and averaged according to normalized weights to obtain the ensemble prediction curve.

[0150]

[0151] The formula for calculating the normalized weights of the k-th candidate matrix is ​​as follows:

[0152]

[0153] Here, the time variable t represents the current time step (or time window) of the inference, i.e., the moment of prediction. The number of candidate matrices L represents the number of candidate state transition matrices selected from the digital twin model library, specifically the top few matrices with the highest probabilities (i.e., the first L matrices truncated after sorting by posterior probability from high to low), rather than all 5 pattern matrices. Matrix indices k and j: k is used to traverse the index (from 1 to L) of the first L selected candidate matrices, identifying the k-th matrix; j is the loop index, also used to traverse the L candidate matrices and perform summation and normalization. Normalization weight ω k ω represents the weighting coefficient of the k-th candidate state transition matrix in the ensemble prediction. Its value is calculated using the second formula and is equal to the proportion (i.e., normalized) of the posterior probability of the corresponding operating mode in the sum of the posterior probabilities of all L selected matrices; ω of all L matrices k The sum is 1. The single-matrix derivation outputs y. (k) (t) represents the independent inference output of the k-th candidate state transition matrix at time t (i.e., the predicted curves of state variables such as voltage, frequency, or power flow obtained from the simulation of this matrix alone). The mode posterior probability p(m) k |f(t)) represents the posterior probability that the system is in the k-th operating mode (such as normal operation, strong fluctuations in new energy sources, or failure) given the feature vector f(t) of the current time window; this probability value comes from the output of the operating mode classifier in step S2. The ensemble prediction output y ensemble (t) represents the final comprehensive prediction result at time t after multi-model weighted integration. Its physical meaning is: the inference results of the selected L candidate state transition matrices are weighted and averaged according to the probability of their corresponding modes occurring, thus taking into account the inference accuracy under different operating conditions and avoiding the bias caused by a single fixed model. The inference step size is adaptively selected according to the problem type: millisecond-level step size is used for frequency stability problems, and second-level step size is used for voltage stability and load balancing problems. After the inference is completed, the predicted change curves of each state variable, the final values ​​of key performance indicators, and the corresponding uncertainty intervals are output. Simultaneously, a no-intervention baseline curve assuming no control action is applied is also output.

[0154] The simulation results of each candidate solution are weighted and comprehensively scored according to the improvement of key performance indicators. The comprehensive scoring formula is as follows:

[0155]

[0156] Among them, scheme s iThe i-th candidate intervention scheme (i.e., a set of control measures to be evaluated generated by a heuristic search algorithm); the total number of indicators m represents the total number of key performance indicators (KPIs) used in evaluating the scheme (e.g., voltage compliance rate, frequency deviation, load reduction, economic cost, etc.); the indicator index j represents the cyclic index number (from 1 to m) of the traversed indicators, used to identify a specific evaluation dimension; the improvement amount ΔI ij Scheme s i The improvement relative to the no-intervention baseline on the j-th indicator; w j The weight coefficient for the j-th indicator is preset by the user or learned through historical feedback. The schemes are sorted in descending order of comprehensive score, recommended schemes are marked, and the prediction curves, confidence intervals, key indicator comparison tables, and no-intervention baseline curves for each scheme are presented visually.

[0157] Step S32: Collect user decision feedback and adjust the storage transition thresholds of the data layer, the pattern recognition and model selection parameters of the analysis layer, and the state transition matrix of the decision layer in a hierarchical and asynchronous manner.

[0158] This sub-step collects user decision-making behavior during actual scheduling as feedback signals, and applies them to the system's data layer, analysis layer, and decision layer parameters according to three different time scales.

[0159] In terms of feedback signal acquisition, three types of user decision-making behaviors are recorded: marking the adoption or rejection of analysis conclusions; recording the selection of recommended solutions and the parameter differences before and after modification; and the manual correction amount of the predicted curve or key indicator value of the inference results. Feedback signals are organized into timestamped sample pairs.

[0160] The first level is the data layer update, executed in the first cycle T1. It summarizes the query latency records from the feedback samples and the user response time to the analysis results, and uses this information to adjust the hot threshold Th of the data temperature adaptive migration scheduler in the data semantic unification and adaptive storage organization module. hot and cold threshold Th cold If a certain type of data object experiences more query response times than a preset number within a statistical period due to being in a cold zone, the hot threshold for that data type is lowered; conversely, if a certain type of data object remains unaccessed for several consecutive periods, its cold threshold is raised. This update strategy ensures that the data storage distribution continuously approximates the actual distribution of users' analytical interests.

[0161] The second level is the analysis layer update, executed in the second cycle T2, where T2 is greater than T1. User rejection records of dynamic model selection results in the feedback samples are used as negative samples. An incremental expectation-maximization algorithm is used to update the mean, covariance, and weights of each Gaussian component in the Gaussian mixture model of the running mode classifier. User confirmation or denial of knowledge graph reasoning results is used to update the weight coefficients of corresponding directed edges; confirmation increases the weight, and denial decreases it, with the weights constrained within a preset range. User evaluation of the model's output stability after hot switching is used to adjust the window length parameter for parallel operation in the hot switching mechanism.

[0162] The third level is the decision-making layer update, executed in the third cycle T3, where T3 is greater than T2. ​​The deviation between the intervention plan finally selected by the user in the feedback sample and the actual system response is used to calibrate the state transition matrix (system matrix A) of the digital twin model. sys Input matrix B sys Output matrix C sys and direct transfer matrix D sys The calibration process employs a recursive least squares algorithm, using the error between the actual response and the projected prediction as a correction term to update the matrix parameters. User scores extracted from their scheme selection behavior serve as reinforcement signals, adjusting the weighting coefficients w in the comprehensive scoring formula. j If a user frequently selects a scheme with a lower weight for a certain indicator, then the weight of that indicator will be increased, and all weights will be normalized.

[0163] Step S33: Establish inter-layer linkage rules, monitor the parameter change range of the analysis layer and decision layer, and trigger additional updates in other layers when the change exceeds the threshold.

[0164] When the mean offset of any Gaussian component in the running pattern classifier exceeds a set threshold, an additional update of the data layer threshold is triggered to adapt to changes in data access patterns caused by changes in the pattern classification boundary. When the update magnitude of the decision layer state transition matrix exceeds a threshold, a rapid recalibration of the knowledge graph edge weights in the analysis layer is triggered to reflect changes in causal relationships revealed by the digital twin model.

[0165] The additional updates triggered by the above actions adhere to the original asynchronous cycle constraints during execution. State variables within each layer retain historical versions with each update, allowing for rollback in case of misjudgment. Through this cross-layer coordination mechanism, the parameter evolution of the data layer, analysis layer, and decision layer remains mutually compatible, forming a closed loop throughout the entire system: the decision layer's deduction generates user feedback, layered asynchronous updates affect parameters at each layer, and cross-layer coordination ensures inter-layer compatibility, improving the quality of the next deduction.

[0166] This step has a clear connection with steps S1 and S2: the initial state snapshot of the digital twin model obtained from step S1 comes from the latest data record of the unified semantic data view, and the data temperature adaptive migration scheduler ensures that the above data is in the hot storage area when the inference is initiated; the current operating mode and its posterior probability distribution, analysis conclusions and knowledge graph semantic constraints are obtained from step S2. The above information determines the action space boundary of the candidate scheme generation and the weighting strategy of the integrated inference; the user feedback signal output in this step is applied to the parameters of steps S1 and S2 respectively through the hierarchical asynchronous feedback update unit to realize the self-evolution of the system as a whole.

[0167] It should be further explained that, in terms of decision-making simulation, existing technologies are mainly divided into two categories: high-fidelity digital twin methods and simplified surrogate model methods. High-fidelity digital twin methods have high computational overhead, requiring several seconds to several minutes for a single simulation; simplified surrogate model methods have limited generalization ability. This step adopts an integrated simulation method that combines a reduced-order state-space model with a multi-mode model library. Weighted integration of multiple models is achieved through integrated prediction curves, striking a balance between simulation time and the generalization ability of multiple models covering different operating conditions. Regarding the feedback mechanism, existing online learning methods typically update parameters for a single model with a fixed update cycle. This step divides the feedback signal into three levels according to its target: data layer, analysis layer, and decision layer, setting short, medium, and long update cycles for each level, matching the time characteristics of rapid changes in power grid data access patterns, slow changes in analysis models, and the slowest changes in system dynamic characteristics. In terms of cross-layer coordination, the parameters of each layer in the existing system evolve independently, which easily leads to mismatch problems. The cross-layer coordination mechanism established in this step triggers data layer updates when the parameters of the analysis layer deviate beyond the limit, and triggers knowledge graph weight recalibration when the parameters of the decision layer change beyond the limit, so that the evolution of the three layers can be mutually adapted.

[0168] It should be noted that technical features that are not fully explained will be addressed using conventional technical methods.

[0169] The above embodiments are only used to illustrate the present invention and are not intended to limit the technical solutions described herein. Although the present invention has been described in detail with reference to the above embodiments, the present invention is not limited to the specific embodiments described above. Therefore, any modifications or equivalent substitutions to the present invention, as well as all technical solutions and improvements that do not depart from the spirit and scope of the invention, are covered within the scope of the claims of the present invention.

Claims

1. A power plant information data intelligent analysis system for regional power grids, characterized in that, It includes a data semantic unification and adaptive storage organization module, a situational awareness and multi-modal analysis and reasoning module, and a forward-looking inference and closed-loop feedback module. The three modules work together in a collaborative manner according to the data layer, analysis layer, and decision layer, and form a closed loop through cross-layer feedback. The data semantic unification and adaptive storage organization module includes: a semantic unification and spatiotemporal alignment unit, used to transform multi-source heterogeneous data into a unified semantic data view; a spatial fractal self-similar index unit, used to construct a progressive hierarchical adaptive index that supports cross-granularity queries; a multi-temperature storage unit, used to construct a hierarchical storage system and drive data migration between different storage levels according to access frequency; and a cross-unit collaboration mechanism, used to make the view the input of the index unit, drive multi-temperature storage migration with index hit information, and adjust the index query strategy based on storage access performance feedback. The situation awareness and multi-mode analysis and reasoning module includes: a situation adaptive model selection engine, used to extract operating features from the view, identify the power grid operating mode and dynamically select the analysis model, and adopt a model hot switching mechanism when switching; a knowledge graph semantic reasoning unit, used to construct and maintain the knowledge graph and update it dynamically according to real-time operating data, and perform semantic reasoning on queries; the engine and unit cooperate through a two-way enhanced closed loop, with semantic reasoning triggered by numerical analysis and the semantic reasoning results fed back to enhance numerical analysis; The forward-looking simulation and closed-loop feedback module includes: a digital twin simulation unit, which is used to construct a power grid simulation model as a digital twin, generate candidate schemes based on situation identification results and semantic constraints, and perform simulations; and a hierarchical asynchronous feedback adaptive update unit, which is used to collect feedback signals and adjust the parameters of the data layer, analysis layer and decision layer at different periods, and has a cross-layer coordination mechanism to trigger the linkage update of other layers when the layer parameters change. The data layer provides the view to the analysis layer, the analysis layer outputs situation information and constraints to the decision layer, and the feedback from the decision layer optimizes the parameters of each layer and maintains inter-layer adaptation through the cross-layer coordination mechanism, forming a global evolution closed loop that runs through the three layers.

2. The intelligent analysis system for power plant information data oriented towards regional power grids according to claim 1, characterized in that, The semantic unification and spatiotemporal alignment unit has a built-in power terminology ontology library for regional power grids. It extracts the indicator name, unit of measurement, and enumerated value range for each accessed heterogeneous data source, and uses a graph embedding-based semantic matching algorithm to automatically map the indicators to a unified semantic layer. Using the data source with the highest sampling frequency in the system as the time reference, it performs multi-scale interpolation alignment on low-frequency data sources, and adds multi-dimensional labels, including at least time labels, spatial coordinate labels, data source labels, quality confidence labels, and semantic type labels, to each aligned data record to form the unified semantic data view. This unit also has an ontology self-expansion subunit, which clusters unmatched indicators with existing concepts through vector embedding, generates ontology expansion suggestions, and incorporates them into the ontology library after confirmation.

3. The intelligent analysis system for power plant information data oriented towards regional power grids according to claim 1, characterized in that, The spatial fractal self-similar indexing unit abstracts the spatial topology of the regional power grid into multiple progressive levels: regional layer, sub-regional layer, station layer, equipment layer, and sensor layer. The organization of index nodes at each level maintains fractal self-similarity with the overall structure. The unit is equipped with a cross-granularity query optimizer, which estimates the overhead of two execution paths—aggregation from the top layer and backtracking from the bottom layer—based on the spatial range and granularity involved in the query request, and automatically selects the path with lower overhead. When the topology changes, incremental reconstruction is performed only on the affected local index branches. The unit also has a double-buffered verification subunit, which is used to rebuild the shadow index and perform consistency verification while keeping the main index online.

4. The intelligent analysis system for power plant information data in a regional power grid according to claim 1, characterized in that, The access frequency prediction-driven multi-temperature storage unit constructs a three-level storage system of hot, warm, and cold. The data temperature adaptive migration scheduler performs cross-temperature layer migration based on the weighted access frequency prediction of data objects and the cooperative signal from the spatial fractal self-similar index unit. After executing the query, the spatial fractal self-similar index unit records the hit frequency of each index branch and sends the high-frequency branch identifier and weighted increment to the data temperature adaptive migration scheduler in real time. The weighted increment is determined by the inverse of the depth of the branch in the index level, with shallower depths having greater weights. The data temperature adaptive migration scheduler updates the weighted access frequency prediction of the corresponding data object based on the weighted increment. For data objects located in the cold storage area, if their weighted access frequency prediction exceeds the hot threshold, they are migrated to the hot storage area; if it is between the warm and hot thresholds, they are migrated to the warm storage area. For data objects located in the hot storage area, if their weighted access frequency prediction is lower than the cold threshold, they are migrated to the cold storage area; if it is between the cold and warm thresholds, they are migrated to the warm storage area.

5. The intelligent analysis system for power plant information data oriented towards regional power grids according to claim 1, characterized in that, The situational adaptive model selection engine extracts multi-dimensional operational feature vectors from the unified semantic data view according to time windows. The feature vectors include at least eight dimensions: voltage stability margin, frequency stability index, load curve shape, intensity of new energy output fluctuation, comprehensive equipment health score, regional supply and demand balance, protection action margin, and communication link quality. An unsupervised classification framework combining Gaussian mixture models and hidden Markov models is used to identify the current power grid operation mode. These operation modes include normal stable operation mode, strong fluctuation mode of new energy sources, high load mode during peak summer season, planned equipment maintenance mode, and fault occurrence or recovery mode. The engine automatically selects the appropriate model from pre-set lightweight time series prediction models, high-precision deep neural network models, physical constraint-based optimization scheduling models, and root cause diagnosis models based on association rule mining. A hot model switching mechanism is used when switching modes. During the switching period, the front-end analysis results are a weighted average of the outputs of the old and new models, with the weight coefficients transitioning linearly over time. The old model resources are released only when the confidence level of the output of the new model exceeds a preset threshold in multiple consecutive time windows.

6. The intelligent analysis system for power plant information data oriented towards regional power grids according to claim 1, characterized in that, The knowledge graph semantic reasoning unit constructs and maintains a weighted directed power knowledge graph. Graph nodes include entity nodes and event nodes. Directed edges represent direct influence, mutual backup, electrical coupling, temporal association, and operational constraint relationships, with dynamically updated weight coefficients. The unit automatically extracts entity relationships from unstructured text to expand the graph and dynamically updates node states and edge weights based on real-time operating data. When a device malfunctions, the unit increases the weight coefficients of the direct influence edges between the malfunctioning device node and the downstream affected nodes by a preset step size, but the weight coefficients do not exceed the upper limit. After the device is running smoothly, the weight coefficients of the relevant edges are gradually reduced by a preset decay step size, but not lower than the preset lower limit. For root cause tracing queries, the unit performs a weighted depth-first search along the directed edges in the opposite direction from the event node, outputting the entity nodes with the largest weight sum on the path as a set of possible root causes. The unit also includes a transfer learning cold start enhancement subunit, used to load a pre-trained general power knowledge graph base and perform parameterized fine-tuning.

7. The intelligent analysis system for power plant information data in a regional power grid according to claim 1, characterized in that, The digital twin simulation unit constructs a reduced-order state-space digital twin model synchronized with the real power grid. This model describes the system dynamics using state vectors, control input vectors, and output vectors. Its system matrix, input matrix, output matrix, and direct transmission matrix are obtained from historical operating data through a subspace identification algorithm and updated with a rolling time window. Simultaneously, it maintains a candidate state transition matrix model library corresponding to different operating modes. The unit receives the situation identification results and semantic constraints, generates candidate intervention schemes, and uses an integrated prediction method to select several candidate state transition matrices according to the posterior probability distribution of the current mode. After independent simulation, the integrated prediction result is obtained by weighted averaging with normalized weights. The unit also outputs the confidence interval of the prediction curve and the no-intervention baseline curve.

8. The intelligent analysis system for power plant information data in a regional power grid according to claim 1, characterized in that, The hierarchical asynchronous feedback adaptive update unit collects the user's actual decision-making behavior as feedback signals, routes the feedback signals to the data layer, analysis layer, and decision layer according to their target, and applies different update cycles: the first cycle adjusts the hot and cold thresholds of the multi-temperature storage unit driven by the access frequency prediction of the data layer; the second cycle, which is longer than the first cycle, adjusts the pattern classification parameters, model selection mapping rules, and knowledge graph edge weights of the analysis layer; and the third cycle, which is longer than the second cycle, calibrates the state transition matrix of the digital twin model. The unit also has a cross-layer coordination trigger, which triggers an additional update of the data layer threshold when the change in the pattern classification parameters of the analysis layer exceeds a preset threshold, and triggers a rapid recalibration of the knowledge graph edge weights of the analysis layer when the update magnitude of the state transition matrix of the decision layer exceeds a preset threshold.

9. A method for intelligent analysis of power plant information data for regional power grids, implemented using the intelligent analysis system for power plant information data for regional power grids as described in any one of claims 1 to 8, characterized in that, It includes the following steps: Step S1: Receive heterogeneous data from multiple sources, and form a unified semantic data view with multi-dimensional labels through semantic unification and spatiotemporal alignment; based on fractal self-similarity, abstract the spatial topology of the regional power grid into multiple progressive levels and construct a spatial fractal self-similarity index; In a three-tiered storage system of hot, warm, and cold, a weighted increment is generated based on the hit frequency of each index branch in the index. The weighted increment is inversely proportional to the depth of the branch in the index level. The weighted access frequency estimate of the data object is updated based on the weighted increment to perform data migration between different temperature levels, thereby achieving storage preheating. In this system, the branch hit frequency generated by the query execution of the index drives the data to warm up. After the data is warmed up, the access latency decreases. The query optimizer of the index adjusts the cost model according to the change in access latency, forming a closed loop within the data layer. Step S2: Extract multi-dimensional operation feature vectors from the unified semantic data view according to time windows, identify the current power grid operation mode using an unsupervised classification framework, and dynamically select an appropriate model from a variety of pre-set analysis models based on the identified operation mode. When switching modes, a model hot-switching mechanism is adopted, in which the old and new models run in parallel and output a weighted average transition. Construct and maintain a weighted directed power knowledge graph, use the knowledge graph to perform semantic reasoning on queries, and use the set of possible root causes obtained by reasoning as semantic prior features to feed back to the feature extraction of the next time window. The semantic prior features and the multi-dimensional operation feature vectors are jointly input into the unsupervised classification framework to form a two-way enhancement closed loop in which numerical anomalies trigger semantic reasoning and semantic conclusions enhance numerical features. Step S3: Construct a digital twin using a reduced-order state-space model. Receive the current operating mode and posterior probability distribution, along with the semantic constraints output from the knowledge graph. Generate candidate intervention schemes and use an integrated prediction method to perform a weighted average of the deduction results of multiple candidate state transition matrices based on the current mode's posterior probability distribution, outputting the prediction effect. Collect user decision feedback signals and adjust the storage transition thresholds in the data layer processing steps, the pattern recognition and model selection parameters in the analysis layer processing steps, and the state transition matrix of the digital twin in a hierarchical asynchronous manner in the first, second, and third cycles. The duration of the first, second, and third cycles increases sequentially. A cross-layer coordination mechanism is set up so that when the parameter changes in the analysis layer or decision layer exceed the threshold, parameter updates in other layers are triggered, forming a global evolutionary closed loop that runs through the data layer, analysis layer, and decision layer.

10. The method according to claim 9, characterized in that, The semantic unification and spatiotemporal alignment described in step S1 include: building a built-in power terminology ontology for regional power grids; using a semantic matching algorithm to calculate the similarity between local indicators and standard concepts in the ontology and establish a mapping; performing multi-scale interpolation alignment on low-frequency data sources using the highest frequency data source as the time benchmark; attaching multi-dimensional labels to each aligned data record to form a unified semantic data view; and clustering unmatched indicators with existing concepts through vector embedding to generate ontology expansion suggestions, which are then incorporated into the ontology after confirmation.