LLM-based supply and demand factor anomaly identification traceability method and apparatus, and storage medium

By using an LLM-based approach for power system anomaly monitoring and tracing, the problems of multi-source heterogeneous data fusion and model mechanism disconnection are solved, enabling efficient anomaly detection and interpretable tracing analysis in the power market, and supporting market risk decision-making.

CN121998682APending Publication Date: 2026-05-08ANHUI ELECTRIC POWER TRADING CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI ELECTRIC POWER TRADING CENT CO LTD
Filing Date
2026-01-29
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies for monitoring and tracing anomalies in the power market suffer from problems such as insufficient fusion of multi-source heterogeneous data, disconnect between model mechanisms and data-driven approaches, low automation in tracing the causes of anomalies, and system rigidity, making them unable to meet the complex risk analysis needs of the power market.

Method used

We employ an LLM-based approach for deep semantic parsing of unstructured data, construct an unstructured text-structured factor association knowledge base, combine an anomaly detection algorithm based on power system physical constraints, construct a factor coupling network using mutual information and Granger causality tests, and use a three-layer progressive tracing logic for causal analysis to generate interpretable anomaly tracing results.

Benefits of technology

It achieves deep fusion and dynamic correlation of multi-source heterogeneous data, improves the reliability of anomaly detection and the interpretability of tracing conclusions, is suitable for complex scenario analysis in the power market, and supports market risk decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998682A_ABST
    Figure CN121998682A_ABST
Patent Text Reader

Abstract

The invention relates to an LLM-based supply and demand factor anomaly identification traceability method and device and a storage medium, and is applied to the technical field of power operation monitoring, and the method comprises the steps: achieving the deep semantic analysis and dynamic association of unstructured data through a field fine tuning model, constructing an unstructured text-structured factor association knowledge base, and obtaining an unstructured text-structured factor association knowledge base; the problems of'information islands' and'semantic gaps' are thoroughly solved, and panoramic and high-quality feature input is provided for anomaly analysis; a detection framework integrating mechanism embedding and data driving is innovated, complex modes in data are utilized, physical constraints of a power system and limitation of a market mechanism are also utilized, the detection reliability in a complex scene is remarkably improved, and false alarms and missed alarms are avoided; an LLM-driven three-layer progressive traceability mechanism is constructed, a clear and credible causal chain atlas is automatically generated, manual analysis by experts is not needed, and the interpretability and consistency of a traceability conclusion are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power operation monitoring technology, specifically to a method, device, and storage medium for identifying and tracing anomalies in supply and demand factors based on LLM. Background Technology

[0002] With the deepening of power market reforms and the continuous increase in the proportion of renewable energy installed capacity, the operating environment of the power system is becoming increasingly complex, and the frequency and magnitude of electricity price fluctuations are significantly increasing. The clearing price in the electricity spot market is influenced by multiple factors, including fluctuations in primary energy prices, uncertainty in renewable energy output, changes in weather conditions, maintenance schedules for transmission and transformation equipment, policy adjustments, and the game-playing strategies of market participants. Abnormal electricity prices can disrupt the rational allocation of market resources, reduce the enthusiasm of market participants to participate in transactions, and may also lead to losses for power generation companies, affecting their subsequent investment and production enthusiasm and increasing market operation risks.

[0003] Anomaly monitoring and root cause analysis of power supply and demand factors (i.e., the thermal power bidding space, which is the difference between total power demand and total non-thermal power output) are the core of power market risk management. Existing technologies in this field suffer from four key shortcomings: First, the fusion of multi-source heterogeneous data is superficial. Structured data (such as SCADA and clearing results) and unstructured data (such as weather warning texts and maintenance information) are often mapped using simple rules or feature splicing, failing to achieve deep semantic understanding and dynamic correlation, resulting in "rich data but poor information." Second, the predictive and anomaly detection model mechanisms are disconnected from data-driven approaches. Purely data-driven models ignore the physical constraints of the power system, while purely rule-based models are rigid and inflexible, with weak scenario generalization capabilities and high false alarm rates. Third, the automation level of anomaly cause tracing is low, relying on expert manual analysis, and often remaining at the level of correlation analysis, lacking interpretable causal reasoning capabilities and unable to generate clear causal chains. Fourth, the system is rigid and inflexible, lacking the ability to continuously self-evolve from operational feedback and unable to adapt to the dynamic evolution of power market structure and rules.

[0004] Existing technologies cannot meet the needs of complex risk analysis in the power market. Therefore, there is an urgent need for a technical solution with deep data fusion, accurate anomaly identification, intelligent causal tracing, and continuous evolution capabilities. Summary of the Invention

[0005] In view of this, the purpose of this invention is to provide a method, device and storage medium for identifying and tracing anomalies in supply and demand factors based on LLM, so as to solve the problems in the prior art where the fusion of multi-source heterogeneous data is superficial, the prediction and anomaly detection model mechanism is disconnected from data-driven, the degree of automation in tracing the causes of anomalies is low and it relies on manual analysis by experts.

[0006] According to a first aspect of the present invention, a method for identifying and tracing anomalies in supply and demand factors based on LLM is provided, the method comprising: Acquire structured data on supply and demand factors from multiple channels, as well as unstructured data on factors affecting supply and demand from multiple dimensions; A pre-trained model is used to perform semantic parsing on the unstructured data, extract structured related data, and construct a knowledge base linking unstructured text to structured factors. The structured data of the supply and demand factors and the structured related data are subjected to data standardization processing to generate a standardized supply and demand factor dataset; An anomaly detection algorithm based on decision trees with embedded power system physical constraints is adopted, combined with the scenario adaptive dynamic prediction interval generated by the time series prediction model, to calculate the anomaly confidence score of each factor in the standardized supply and demand factor dataset, and to distinguish between abnormal and normal states based on the anomaly confidence score of each factor. The mutual information method is used to measure the nonlinear correlation strength between factors, and the Granger causality test is used to determine the causal direction in time series. A directional factor coupling network is constructed. Typical abnormal patterns are identified based on the directional factor coupling network. An abnormal report containing the core information of the abnormality is generated based on the identification results. Based on the importance and scope of influence of the factors, the analytic hierarchy process is used to prioritize the abnormal factors. Based on a professional knowledge base in the power industry, a reasoning template is constructed through prompt word engineering. Standardized supply and demand factor datasets and anomaly reports are integrated, and a three-layer progressive tracing logic is used to complete the cause analysis. After verification by an expert rule base, an interpretable anomaly tracing result is output.

[0007] Preferably, The data standardization process includes: Interpolation methods are used to fill in the missing values ​​in the structured data and structured correlation data of the supply and demand factors; Outliers in the structured data and structured correlation data of the supply and demand factors are removed using the 3σ principle; Based on Z-score standardization, the dimensions of different dimension factor data in the structured data and structured correlation data of the supply and demand factors are unified. The logical consistency between textual information and numerical data is verified using a pre-trained LLM model. Inconsistent data is marked and then manually reviewed.

[0008] Preferably, The step of using a pre-trained model to perform semantic parsing on the unstructured data includes: The unstructured data is preprocessed using a Qwen-Max large model finely tuned with data from the power sector, including word segmentation, stop word filtering, and domain terminology normalization. Using a template for extracting information from the power sector, structured relational data containing region, time, factor type, numerical value, and affected object are extracted from preprocessed text.

[0009] Preferably, it further includes: Standardized supply and demand factor datasets, anomaly reports, and anomaly tracing results are stored in a historical case library for model parameter iteration and related knowledge base updates.

[0010] Preferably, The decision tree-based anomaly detection algorithm is an improved isolated forest algorithm, and the time series prediction model is an LSTM model. The physical constraints of the power system include the unit ramp rate limit, the constraint that the output of new energy sources cannot change abruptly, and the power balance boundary. The anomaly confidence score is 0–100. A score greater than or equal to the first confidence threshold is considered a high-confidence anomaly, a score between the first and second confidence thresholds is considered a low-confidence anomaly, and a score less than the second confidence threshold is considered a normal state.

[0011] Preferably, The fused standardized supply and demand factor dataset and anomaly report include: By using a pre-trained LLM model, semantic alignment and context fusion are performed on the standardized supply and demand factor dataset and anomaly reports to obtain standard inference input.

[0012] Preferably, The three-level progressive tracing logic used to complete the causal analysis includes: Based on real-time factor changes and event logs, locate the direct triggering events of anomalies; By leveraging the causal reasoning capabilities of pre-trained LLM models, we can uncover the dynamic coupling relationships between multiple factors and obtain anomaly propagation paths. Retrieve semantically similar historical cases from the power industry knowledge base, compare evolutionary patterns, and generate a causal chain map.

[0013] According to a second aspect of the present invention, an LLM-based supply and demand factor anomaly identification and tracing device is provided, the device comprising: Multi-source data acquisition module: used to acquire structured data of supply and demand factors from multiple channels, as well as unstructured data of factors affecting supply and demand from multiple dimensions; Unstructured data processing module: Used to perform semantic parsing on the unstructured data using a pre-trained model, extract structured related data, and construct a knowledge base linking unstructured text to structured factors; Data standardization module: used to perform data standardization processing on the structured data of the supply and demand factors and the structured related data to generate a standardized supply and demand factor dataset; Anomaly detection module: It is used to calculate the anomaly confidence score of each factor in the standardized supply and demand factor dataset by using an anomaly detection algorithm based on decision tree embedded with power system physical constraints, combined with the scenario adaptive dynamic prediction interval generated by the time series prediction model, and to distinguish between abnormal and normal states based on the anomaly confidence score of each factor. Anomaly report generation module: It is used to measure the nonlinear correlation strength between factors by mutual information method, and to determine the causal direction in time series by Granger causality test, and to construct a directional factor coupling network; it identifies typical anomaly patterns based on the directional factor coupling network, generates anomaly reports containing core information of anomalies based on the identification results, and prioritizes the anomalous factors based on the importance and scope of influence of the factors using the analytic hierarchy process. Source tracing module: Based on a professional knowledge base in the power field, it constructs reasoning templates through prompt word engineering, integrates standardized supply and demand factor datasets and anomaly reports, and completes cause analysis using a three-layer progressive source tracing logic. After verification by an expert rule base, it outputs interpretable anomaly source tracing results.

[0014] According to a third aspect of the present invention, a storage medium is provided, the storage medium storing a computer program, which, when executed by a host controller, implements the steps of the above-described method.

[0015] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This application achieves deep semantic parsing and dynamic association of unstructured data through a domain-fine-tuned model, constructing a knowledge base of "unstructured text - structured factors" association, completely solving the problems of "information silos" and "semantic gaps," and providing panoramic, high-quality feature input for anomaly analysis. It innovates a detection framework that integrates mechanism embedding and data-driven approaches, utilizing complex patterns in the data while being constrained by the physical constraints of the power system and market mechanisms, significantly improving detection reliability in complex scenarios and avoiding false alarms and missed alarms. It constructs an LLM-driven three-layer progressive tracing mechanism, automatically generating clear and reliable causal chain graphs without relying on expert manual analysis, greatly improving the interpretability and consistency of tracing conclusions. Focusing on anomaly analysis of core economic indicators in the power market (thermal power bidding space), it upgrades from serving "equipment safety operation and maintenance" to supporting "market risk decision-making," filling the gap in existing technologies for market risk domain analysis, and is applicable to multiple scenarios such as the power spot market and new energy grid connection and consumption.

[0016] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0018] Figure 1 This is a flowchart illustrating an LLM-based supply and demand factor anomaly identification and tracing method according to an exemplary embodiment. Figure 2 This is a schematic diagram of a supply and demand factor anomaly identification and tracing device based on LLM, according to another exemplary embodiment. In the attached diagram: 1-Multi-source data acquisition module, 2-Unstructured data processing module, 3-Data standardization module, 4-Anomaly detection module, 5-Anomaly report generation module, 6-Source tracing module. Detailed Implementation

[0019] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0020] Example 1 Figure 1 This is a flowchart illustrating an LLM-based supply and demand factor anomaly identification and tracing method according to an exemplary embodiment, such as... Figure 1 As shown, the method includes: S1, obtain structured data on supply and demand factors from multiple channels, and obtain unstructured data on factors affecting supply and demand from multiple dimensions; S2, use a pre-trained model to perform semantic parsing on the unstructured data, extract structured related data, and construct a knowledge base linking unstructured text to structured factors; S3, perform data standardization processing on the structured data of the supply and demand factors and the structured related data to generate a standardized supply and demand factor dataset; S4. An anomaly detection algorithm based on decision tree with embedded power system physical constraints is adopted. Combined with the scenario adaptive dynamic prediction interval generated by the time series prediction model, the anomaly confidence score of each factor in the standardized supply and demand factor dataset is calculated. The anomaly confidence score of each factor is used to distinguish between abnormal and normal states. S5. The strength of nonlinear correlation between factors is measured by mutual information method, and the causal direction in time series is determined by Granger causality test. A directional factor coupling network is constructed. Typical abnormal patterns are identified according to the directional factor coupling network. An abnormal report containing the core information of the abnormality is generated according to the identification results. Based on the importance and scope of influence of the factors, the analytic hierarchy process is used to prioritize the abnormal factors. S6, based on a professional knowledge base in the power field, constructs reasoning templates through prompt word engineering, integrates standardized supply and demand factor datasets and anomaly reports, and completes cause analysis using a three-layer progressive tracing logic. After verification by the expert rule base, it outputs interpretable anomaly tracing results. It is understood that this embodiment specifically includes: Multi-source data acquisition and preprocessing: The system collects four types of structured data: power source side, load side, grid side, and market side. The data collection channels include the power system SCADA platform, new energy power plant monitoring system, power spot trading system, and power grid dispatch EMS system. The data collection frequency is set to once every 15 minutes. Simultaneously, it captures unstructured data such as meteorological warning texts, equipment operation and maintenance work orders, dispatch instruction logs, and historical anomaly case reports.

[0021] For unstructured data, a Qwen-Max model finely tuned for the power sector is used as the core of semantic parsing. Power sector information extraction templates are constructed through prompt word engineering. Unstructured data undergoes word segmentation, stop word filtering, and domain terminology normalization to extract structured relational data such as region, time, factor type, value, and affected objects. This data is then semantically mapped to the factor dimensions of structured data. For example, for the text "#3 main transformer tripped due to lightning strike, resulting in a loss of 300MW of output," the entity "#3 main transformer" and the event "tripped" are identified. The attribute "loss of 300MW of output" is extracted, and semantic associations are established with structured data such as grid reserve capacity, constructing an "unstructured text-structured factor" association knowledge base.

[0022] Data standardization processing: Linear interpolation is used to fill missing values, and extreme outliers are removed using the 3σ principle (|x-μ|>3σ is considered an extreme outlier). Z-score standardization (z=(x-μ) / σ, where μ is the mean and σ is the standard deviation) is used to unify the dimensions of data with different dimensions. An LLM model is used to verify the logical consistency between textual information and numerical data. For example, the time consistency between "typhoon landfall" in meteorological warning texts and the sudden drop in wind turbine output data in the SCADA system is verified; inconsistent data is marked and manual review is triggered. Finally, a standardized supply and demand factor dataset containing "factor ID, timestamp, value / status, associated text label, and data quality score" is generated, stored in an HBase distributed database, and pushed to the anomaly detection module via API. Original data and processing logs are also retained, supporting full-process traceability.

[0023] Identification of supply and demand anomalies: This embodiment addresses the shortcomings of traditional anomaly detection methods that "only consider a single indicator and rely on fixed thresholds." By integrating power system operation mechanisms with data-driven technology, it achieves full-scale identification from single-point anomalies to systemic risks. The entire process is divided into three progressive levels: first, focusing on intelligent anomaly detection of individual factors; second, analyzing the dynamic correlation between multiple factors; and finally, providing structured output and prioritization of anomalies. Specifically, it includes: Single-factor intelligent anomaly detection: The isolated forest algorithm is improved by embedding power system physical constraints (the upper limit of unit ramp rate is set to 5% / minute, the constraint that the output of new energy sources cannot change abruptly, and the power balance boundary) into the decision tree splitting rules to ensure that the anomaly judgment conforms to the operation pattern. At the same time, the LSTM model is used to predict the factor values ​​for the next few hours. The input features include the time series data of the past 3 days and the quantitative features corresponding to the associated text labels to generate a normal fluctuation range that is adaptively adjusted based on the current operation scenario (such as high wind power penetration, holiday load, extreme weather, etc.).

[0024] Combining the anomaly detection results of the improved isolated forest method with the LSTM dynamic prediction interval, anomaly confidence scores (0–100 points) are calculated for each factor: scores ≥80 indicate high-confidence anomalies, scores 40 ≤ scores <80 indicate low-confidence anomalies, and scores <40 indicate normal conditions. For example, if the actual value of the wind power output factor exceeds three standard deviations of the LSTM prediction interval, and the improved isolated forest method determines that it violates the "non-mutational constraint on new energy output," then the anomaly confidence score is 85 points, and it is judged as a high-confidence anomaly.

[0025] Multifactor association analysis: Using four key factors—power source, load, power grid, and market—as nodes, a directional "factor coupling network" is constructed. This network measures the strength of nonlinear correlations between factors using mutual information (mutual information ≥ 0.6 indicates a strong correlation), and combines this with Granger causality tests to determine the direction of time-series causality (p < 0.05 indicates a causal relationship). Based on this network, the system can identify whether an anomaly has spread from one factor to multiple related factors. While graph neural networks (GNNs) can be used to model such propagation processes, simpler rule-based propagation or ensemble tree models can be used instead in practical engineering. Regardless of whether GNNs are used, the core objective is to identify five typical anomaly patterns: Single-factor numerical mutation: A single factor changes by more than 20% within 15 minutes; Single-factor trend drift: A single factor shifts in the same direction for four consecutive data collection periods (1 hour), with a cumulative shift exceeding 15%; Two-factor mismatch: The load factor and the non-thermal power output factor show opposite trends, and the mutual information value is ≥0.7; Multifactor chain fluctuations: ≥3 causally related factors show abnormalities sequentially, with a propagation delay of ≤1 hour; Overall supply and demand imbalance: There are ≥3 types of abnormal factors on the power supply side, load side, grid side, and market side, and the overall supply and demand balance deviation exceeds 10%; Exceptional structured output: The system integrates detection results to generate anomaly reports, including anomaly ID, occurrence time, list of involved factors, anomaly type, confidence score, deviation magnitude, and related factor chains. Using the Analytic Hierarchy Process (AHP), anomalies are prioritized in a five-level P1–P5 hierarchy (P1 being the highest priority and P5 the lowest) based on factor importance (calculated using SHAP values, with SHAP values ​​≥ 0.5 indicating important factors) and impact scope (affecting ≥ 100,000 users indicating large-scale impact). Anomaly reports are pushed to the source tracing module via a Kafka message queue, while samples and tags are written to the historical case database. The system aims for an anomaly detection accuracy of at least 95% and a recall rate of at least 92%, ensuring reliable and accurate anomaly detection even in complex operating environments.

[0026] The origins of LLM-driven architecture: Domain knowledge injection and reasoning template construction: Load a professional knowledge base in the power field, including core content such as supply and demand balance mechanism, power market trading rules (spot market clearing mechanism, ancillary service pricing rules), and equipment operation specifications (unit start-up and shutdown constraints, line transmission limits); encode knowledge into reasoning templates through prompt word engineering, for example: "Given anomaly factors are {anomaly factor list}, associated text information is {text tag}, based on supply and demand balance mechanism and market trading rules, reason about the direct cause, transmission path and root cause of the anomaly".

[0027] Multi-source heterogeneous information fusion and intelligent reasoning: It receives anomaly reports from the anomaly identification module and simultaneously integrates standardized supply and demand factor datasets (such as wind power output time series data, load demand data, and electricity price data) with unstructured text information (such as weather warning notices "typhoon landfall, wind speed drops sharply" and equipment fault logs "#3 main transformer trips"). Through the cross-modal understanding capabilities of LLM, it achieves semantic alignment and context fusion to form a unified inference input.

[0028] Three-level progressive tracing logic: First layer: Direct cause matching. Based on the temporal change characteristics of abnormal factors and the timestamps of event logs, the direct triggering events can be quickly located. For example, if the anomaly report shows "Wind power output factor (ID: WF-001) showed a high confidence anomaly (score 88 points) at 14:00, with a deviation of 35%", combined with the meteorological text "Typhoon made landfall in area A at 13:50, and the wind speed dropped below 3m / s", the cause can be directly matched as "The typhoon caused a sharp drop in wind power output in area A".

[0029] The second layer: Coupling causal analysis. Leveraging the causal reasoning capabilities of LLM, this layer uncovers dynamic coupling relationships among multiple factors, generating causal transmission chains. For example, based on "sudden drop in wind power output," combined with unit ramp-up constraints and market clearing rules, the transmission path is inferred as: "sudden drop in wind power output (WF-001 anomaly) → increased demand for gas turbine start-up and shutdown (GT-003 factor fluctuation) → triggering of transmission limits on regional B lines (TL-002 anomaly) → exacerbated node congestion (BL-001 anomaly) → abnormally tight bidding space for thermal power (PB-001 anomaly) → electricity price increase (PR-001 anomaly)."

[0030] The third layer: Historical case analogy verification. Semantically similar cases are retrieved from the historical case library, such as "the abnormal bidding space caused by the sudden drop in wind power output in area C due to the typhoon on August 10, 2023". The abnormal evolution patterns (abnormal order of factors, deviation magnitude) and the treatment effects (the bidding space recovered after increasing the gas turbine output) of the two cases are compared to enhance the robustness of the current inference and generate a causal chain diagram.

[0031] Expert rule base validation: The expert rule base is built based on power system operation procedures and typical fault modes, including rules such as "When wind power output drops by ≥30%, the gas turbine ramp-up response delay should be ≤30 minutes, otherwise it will cause a reserve shortage." A rule-evidence matching mechanism is used to verify the logical consistency of LLM inference results. For example, if the LLM inference "insufficient gas turbine ramp-up" is an intermediate cause, the rule base's rule "gas turbine ramp-up rate ≥5% / minute is a normal response" is verified. Combined with gas turbine output data (actual ramp-up rate 3% / minute), the rationality of this cause is verified, and the deviation result is corrected (by supplementing the detailed description of "gas turbine ramp-up rate not meeting the standard"), ultimately outputting an interpretable source tracing conclusion.

[0032] Closed-loop optimization: Standardized supply and demand factor datasets, anomaly reports, source tracing conclusions, and expert review results are stored in a historical case library. On a regular (weekly) basis, the prompt word templates of LLM, the constraint weights of the improved isolated forest, and the parameters of the LSTM model are iteratively optimized. At the same time, new causal relationship patterns (such as bidding space anomalies caused by new market pricing strategies) are extracted from new cases, and the "unstructured text-structured factor" association knowledge base and expert rule base are updated to achieve a closed-loop evolution of "perception-cognition-decision".

[0033] To make the technical solution of this invention clearer and easier to understand, the following detailed description is provided in conjunction with specific implementation examples. The examples are based on actual operating data from a provincial electricity spot market on July 15, 2024, as follows: Implementation scenario: This provincial-level electricity spot market includes 10 new energy power stations (6 wind power and 4 photovoltaic), 5 coal-fired power plants, and 3 gas-fired power plants. The grid side includes 20 key transmission lines, and the market side covers 100 electricity-consuming enterprises. From 14:00 to 15:00 on July 15, 2024, the system detected an abnormal tightening of the thermal power bidding space (PB-001), requiring the anomaly identification and cause tracing to be completed using the method of this invention.

[0034] Implementation steps: Multi-source data acquisition and preprocessing: Structured data acquisition: Data on wind power output (WF-001 to WF-006), photovoltaic power output (PV-001 to PV-004), coal-fired power output (CL-001 to CL-005), gas-fired power output (GT-001 to GT-003), load demand (LD-001), and line transmission power (TL-001 to TL-020) are collected from 14:00 to 15:00 through the SCADA platform at a frequency of 15 minutes / time; the clearing price at 14:00 (PR-001) is collected through the electricity spot trading system; and the grid reserve capacity (RS-001) is collected through the EMS system.

[0035] Unstructured data collection: Capture the warning text issued by the meteorological department at 13:50: "Typhoon Maria made landfall on the eastern coast of the province with wind force reaching level 10, expected to last for 2 hours"; the maintenance work order issued by the power grid dispatch center at 14:05: "Due to the impact of the typhoon, the #12 transmission line (TL-012) is temporarily shut down for maintenance"; and the case report from the historical case database on August 10, 2023: "Typhoon caused a sharp drop in wind power output, resulting in abnormal bidding space".

[0036] Parsing and Knowledge Base Construction: The Qwen-Max large model parses meteorological warning texts, extracting the "typhoon landfall" event, time "13:50", affected area "eastern coastal area", and duration "2 hours"; it parses maintenance work orders, extracting the "#12 line outage" event, time "14:05", and affected object "TL-012"; it establishes semantic associations between "typhoon landfall" and wind power output factors, and between "line outage" and transmission power factors, and updates the associated knowledge base.

[0037] Data standardization: The original data for WF-001 at 14:00 was missing. Linear interpolation (based on 800MW at 13:45 and 300MW at 14:15) was used to fill in the missing data to 550MW. The transmission power of TL-012 at 14:15 was detected to be 0MW, which meets the 3σ principle (historical average of 200MW, standard deviation of 50MW, |0-200|>3×50). It was determined to be a reasonable outlier (due to shutdown and maintenance) and was not removed. The Z-score standardization was used to unify the dimensions of all structured data. LLM verification confirmed the logical consistency between "typhoon landfall (13:50)" and "WF-001 decreased from 800MW at 13:45 to 550MW at 14:00". The consistency was confirmed to be consistent, and a standardized supply and demand factor dataset was generated.

[0038] Identification of supply and demand anomalies: Single-factor intelligent anomaly detection: The improved Isolation Forest algorithm incorporates constraints such as unit ramp rate (5% / minute) and power balance to process standardized data; the LSTM model is input with historical 3-day wind power output, load demand and other data, as well as the textual quantitative features of "typhoon landfall", predicting that WF-001's normal fluctuation range at 14:00 is 650MW-900MW. The actual value is 550MW, exceeding the range. The improved Isolation Forest algorithm determines that it violates the "new energy output cannot change abruptly" constraint (a decrease of 31.25% within 15 minutes > 20%), and the anomaly confidence score of WF-001 is 89 points (high confidence anomaly); similarly, the anomaly confidence score of TL-012 is 92 points (high confidence anomaly), and the actual value of PB-001 (thermal bidding space) is 120MW, with an LSTM prediction range of 200MW-350MW, and an anomaly confidence score of 87 points (high confidence anomaly).

[0039] Multi-factor correlation analysis: The mutual information value between WF-001 and PB-001 was calculated to be 0.75 (strong correlation), and the mutual information value between TL-012 and PB-001 was 0.68 (strong correlation). Granger causality test showed that the WF-001 anomaly (14:00) was the cause of the PB-001 anomaly (14:15) (p=0.03<0.05), and the TL-012 anomaly (14:05) was the cause of the PB-001 anomaly (14:15) (p=0.02<0.05). A factor coupling network was constructed: WF-001→PB-001, TL-012→PB-001. The anomaly pattern was identified as "multi-factor chain fluctuation" (three related factors are sequentially abnormal, with a propagation delay ≤1 hour).

[0040] Anomaly Structured Output: Generate an anomaly report with an anomaly ID of ABN-20240715-001, occurrence time 14:15, involving the factor list [WF-001, TL-012, PB-001], anomaly type of multi-factor chain fluctuation, confidence score of 87, deviation magnitude of 43.75% ((200-120) / 200), and the associated factor chain is WF-001 anomaly → TL-012 anomaly → PB-001 anomaly; based on the AHP method, PB-001 is the core market factor (SHAP value 0.65), affecting 120 home appliance companies (≥100,000 users), and the judgment priority is P1 (highest priority). The anomaly report is pushed to the source tracing module via Kafka.

[0041] The origins of LLM-driven architecture: Domain knowledge injection and reasoning template construction: Load supply and demand balance mechanism, power market clearing rules (electricity price increase under congestion), equipment operation specifications (line outage leading to transmission capacity reduction); construct reasoning template: "Given the anomaly factors are [WF-001, TL-012, PB-001], the anomaly type is multi-factor chain fluctuation, the associated text information is 'typhoon landfall, #12 line outage', based on the supply and demand balance mechanism and market transaction rules, reason about the direct cause, transmission path and root cause of the anomaly."

[0042] Multi-source information fusion: LLM integrates WF-001 output drop data, TL-012 shutdown data, PB-001 tightening data, meteorological warning texts, maintenance work orders and other information to achieve semantic alignment.

[0043] Three-tiered progressive tracing: Direct cause matching: The direct cause of WF-001 anomaly is "the typhoon landfall caused a sudden drop in wind speed at wind farms along the eastern coast"; the direct cause of TL-012 anomaly is "the typhoon caused a fault in the #12 transmission line, resulting in a temporary shutdown for maintenance"; the direct cause of PB-001 anomaly is "the combined effect of a sudden drop in wind power output and line shutdown".

[0044] Coupling Cause Analysis: Reasoning Transmission Path: "Typhoon landfall → wind speed at wind farms along the eastern coast drops sharply → WF-001 output decreases from 800MW to 550MW (14:00) → total non-thermal power output decreases → thermal power bidding space initially tightens; at the same time, the typhoon causes the #12 transmission line to fail and shut down (14:05) → regional transmission capacity decreases → node congestion intensifies → gas turbine ramping response (GT-001 ramping rate 3% / minute < 5% / minute, not meeting the standard) → insufficient reserve capacity (RS-001 decreases from 500MW to 200MW) → thermal power bidding space further tightens (decreases to 120MW at 14:15) → clearing price increases (PR-001 increases from 0.3 yuan / kWh to 0.5 yuan / kWh)."

[0045] Historical case analogy verification: A similar case from August 10, 2023 was retrieved (typhoon caused a sharp drop in wind power output + line shutdown → abnormal bidding space). In this case, after the output of the gas turbine units was increased urgently (the ramp rate was increased to 6% / minute), the bidding space returned to the normal range within 30 minutes, verifying the rationality of the current transmission path and generating a causal chain diagram.

[0046] Expert rule base verification: The rule "When the wind power output drops by ≥30%, the gas turbine ramp-up response delay should be ≤30 minutes and the ramp-up rate should be ≥5% / minute" was called. The actual ramp-up rate of GT-001 was verified to be 3% / minute. It was confirmed that "insufficient gas turbine ramp-up" was the key intermediate cause. The source tracing conclusion was revised and the detail that "the gas turbine unit's ramp-up rate did not meet the standard, which aggravated the shortage of reserves" was added.

[0047] Closed-loop optimization: The standardized dataset, anomaly report, source tracing conclusions (including causal chain diagrams), and expert review opinions (confirming the accuracy of the source tracing conclusions) will be stored in the historical case library. On July 22, 2024 (weekly iteration), the input feature weights of the LSTM model will be optimized based on this case (the weight of the quantitative feature "typhoon duration" will be increased), the association strength between "typhoon" and "wind power output" and "line fault" in the associated knowledge base will be updated, and the "emergency response standard for gas turbine ramp rate (≥6% / minute under typhoon weather)" in the expert rule base will be added to achieve closed-loop evolution of the system.

[0048] Implementation results: In this implementation, the method of the present invention completed anomaly identification within 15 minutes and output an interpretable cause-tracing conclusion within 30 minutes. The anomaly detection accuracy rate was 96.5%, the recall rate was 93.2%, and the consistency between the cause-tracing conclusion and the expert manual analysis results reached 98%. Based on the cause-tracing conclusion, the dispatch center took emergency measures (increasing the gas turbine unit ramp rate to 6.5% / minute and activating the backup line TL-018), and the competitive bidding space for thermal power plants was restored to 220MW (normal range) within 1 hour, effectively avoiding market risks.

[0049] Example 2 Figure 2 This is a schematic diagram of a supply and demand factor anomaly identification and tracing device based on LLM, according to another exemplary embodiment. The device includes: Multi-source data acquisition module 1: used to acquire structured data of supply and demand factors from multiple channels, as well as unstructured data of factors affecting supply and demand from multiple dimensions; Unstructured data processing module 2: Used to perform semantic parsing on the unstructured data using a pre-trained model, extract structured related data, and construct a knowledge base linking unstructured text to structured factors; Data standardization module 3: used to perform data standardization processing on the structured data of the supply and demand factors and the structured related data to generate a standardized supply and demand factor dataset; Anomaly detection module 4: It is used to employ an anomaly detection algorithm based on decision tree with embedded power system physical constraints, combined with the scenario adaptive dynamic prediction interval generated by the time series prediction model, to calculate the anomaly confidence score of each factor in the standardized supply and demand factor dataset, and to distinguish between abnormal and normal states based on the anomaly confidence score of each factor. Anomaly report generation module 5: It is used to measure the nonlinear correlation strength between factors by mutual information method, combine Granger causality test to determine the causal direction in time series, and construct a directional factor coupling network; identify typical anomaly patterns according to the directional factor coupling network, generate anomaly reports containing core information of anomalies according to the identification results, and prioritize the anomaly factors based on the importance and scope of influence of factors using the analytic hierarchy process. Source tracing module 6: Based on a professional knowledge base in the power field, it constructs reasoning templates through prompt word engineering, integrates standardized supply and demand factor datasets and anomaly reports, and completes cause analysis using a three-layer progressive source tracing logic. After verification by an expert rule base, it outputs interpretable anomaly source tracing results.

[0050] Example 3: This embodiment provides a storage medium storing a computer program, which, when executed by a host controller, implements the various steps in the above method. It is understood that the storage medium mentioned above can be a read-only memory, a hard disk, or an optical disk, etc.

[0051] It is understood that the same or similar parts in the above embodiments can be referred to each other, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.

[0052] It should be noted that in the description of this invention, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this invention, unless otherwise stated, "a plurality of" means at least two.

[0053] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0054] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0055] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0056] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0057] The storage media mentioned above can be read-only memory, disk, or optical disk, etc.

[0058] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0059] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A supply and demand factor anomaly identification and tracing method based on LLM, characterized in that, The method includes: Acquire structured data on supply and demand factors from multiple channels, as well as unstructured data on factors affecting supply and demand from multiple dimensions; A pre-trained model is used to perform semantic parsing on the unstructured data, extract structured related data, and construct a knowledge base linking unstructured text to structured factors. The structured data of the supply and demand factors and the structured related data are subjected to data standardization processing to generate a standardized supply and demand factor dataset; An anomaly detection algorithm based on decision trees with embedded power system physical constraints is adopted, combined with the scenario adaptive dynamic prediction interval generated by the time series prediction model, to calculate the anomaly confidence score of each factor in the standardized supply and demand factor dataset, and to distinguish between abnormal and normal states based on the anomaly confidence score of each factor. The mutual information method is used to measure the nonlinear correlation strength between factors, and the Granger causality test is used to determine the causal direction in time series. A directional factor coupling network is constructed. Typical abnormal patterns are identified based on the directional factor coupling network. An abnormal report containing the core information of the abnormality is generated based on the identification results. Based on the importance and scope of influence of the factors, the analytic hierarchy process is used to prioritize the abnormal factors. Based on a professional knowledge base in the power industry, a reasoning template is constructed through prompt word engineering. Standardized supply and demand factor datasets and anomaly reports are integrated, and a three-layer progressive tracing logic is used to complete the cause analysis. After verification by an expert rule base, an interpretable anomaly tracing result is output.

2. The method according to claim 1, characterized in that, The data standardization process includes: Interpolation methods are used to fill in the missing values ​​in the structured data and structured correlation data of the supply and demand factors; Outliers in the structured data and structured correlation data of the supply and demand factors are removed using the 3σ principle; Based on Z-score standardization, the dimensions of different dimension factor data in the structured data and structured correlation data of the supply and demand factors are unified. The logical consistency between textual information and numerical data is verified using a pre-trained LLM model. Inconsistent data is marked and then manually reviewed.

3. The method according to claim 2, characterized in that, The step of using a pre-trained model to perform semantic parsing on the unstructured data includes: The unstructured data is preprocessed using a Qwen-Max large model finely tuned with data from the power sector, including word segmentation, stop word filtering, and domain terminology normalization. Using a template for extracting information from the power sector, structured relational data containing region, time, factor type, numerical value, and affected object are extracted from preprocessed text.

4. The method according to claim 3, characterized in that, Also includes: Standardized supply and demand factor datasets, anomaly reports, and anomaly tracing results are stored in a historical case library for model parameter iteration and related knowledge base updates.

5. The method according to claim 4, characterized in that, The decision tree-based anomaly detection algorithm is an improved isolated forest algorithm, and the time series prediction model is an LSTM model. The physical constraints of the power system include the unit ramp rate limit, the constraint that the output of new energy sources cannot change abruptly, and the power balance boundary. The anomaly confidence score is 0–100. A score greater than or equal to the first confidence threshold is considered a high-confidence anomaly, a score between the first and second confidence thresholds is considered a low-confidence anomaly, and a score less than the second confidence threshold is considered a normal state.

6. The method according to claim 5, characterized in that, The fused standardized supply and demand factor dataset and anomaly report include: By using a pre-trained LLM model, semantic alignment and context fusion are performed on the standardized supply and demand factor dataset and anomaly reports to obtain standard inference input.

7. The method according to claim 6, characterized in that, The three-level progressive tracing logic used to complete the causal analysis includes: Based on real-time factor changes and event logs, locate the direct triggering events of anomalies; By leveraging the causal reasoning capabilities of pre-trained LLM models, we can uncover the dynamic coupling relationships between multiple factors and obtain anomaly propagation paths. Retrieve semantically similar historical cases from the power industry knowledge base, compare evolutionary patterns, and generate a causal chain map.

8. A supply and demand factor anomaly identification and traceability device based on LLM, characterized in that, The device includes: Multi-source data acquisition module: used to acquire structured data of supply and demand factors from multiple channels, as well as unstructured data of factors affecting supply and demand from multiple dimensions; Unstructured data processing module: Used to perform semantic parsing on the unstructured data using a pre-trained model, extract structured related data, and construct a knowledge base linking unstructured text to structured factors; Data standardization module: used to perform data standardization processing on the structured data of the supply and demand factors and the structured related data to generate a standardized supply and demand factor dataset; Anomaly detection module: It is used to calculate the anomaly confidence score of each factor in the standardized supply and demand factor dataset by using an anomaly detection algorithm based on decision tree embedded with power system physical constraints, combined with the scenario adaptive dynamic prediction interval generated by the time series prediction model, and to distinguish between abnormal and normal states based on the anomaly confidence score of each factor. Anomaly report generation module: It is used to measure the nonlinear correlation strength between factors by mutual information method, and to determine the causal direction in time series by Granger causality test, and to construct a directional factor coupling network; it identifies typical anomaly patterns based on the directional factor coupling network, generates anomaly reports containing core information of anomalies based on the identification results, and prioritizes the anomalous factors based on the importance and scope of influence of the factors using the analytic hierarchy process. Source tracing module: Based on a professional knowledge base in the power field, it constructs reasoning templates through prompt word engineering, integrates standardized supply and demand factor datasets and anomaly reports, and completes cause analysis using a three-layer progressive source tracing logic. After verification by an expert rule base, it outputs interpretable anomaly source tracing results.

9. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by the main controller, implements each step of the supply and demand factor anomaly identification and tracing method based on LLM as described in any one of claims 1-7.