A multi-source data processing and causal correlation identification method for virtual power plant response level analysis
Patent Information
- Application Number
- CN202610895663.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2046-06-22
AI Technical Summary
[0011]解决现有技术中多源关联数据处理粗糙、时空映射关系不清、关键影响因素提取不足以及统一分析支撑能力不强等问题,为后续调节资源响应状态分析、能力评估及趋势预判提供稳定可靠的数据基础与特征基础
[0222] 1. This invention solves the technical problem of unifying the representation of multi-source heterogeneous data in the analysis of virtual power plant regulation resource response levels. Through unified time granularity mapping, semantic normalization, low-rank missing data repair, physical constraint projection, and Gaussian distribution normalization, this invention transforms system boundary information, environmental meteorological information, resource-side operational information, and time calendar information into a unified, complete, and computable data sample matrix, significantly improving the integrity, consistency, and modelability of the original data.
Smart Images

Figure CN122413371B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for multi-source data processing and causal correlation identification for virtual power plant response level analysis. Background Technology
[0002] With the continuous advancement of new power system construction, the scale of distributed and adjustable resources such as wind power, photovoltaics, energy storage, electric vehicles, and flexible loads is constantly expanding. The power system operation mode is shifting from the traditional "source follows load" to "multi-dimensional collaborative operation of source, grid, load, and storage." Against this backdrop, virtual power plants, as an important technological carrier for aggregating multiple types of distributed resources and achieving unified perception, unified analysis, and collaborative response, have gradually become a key link in improving the system's flexible adjustment and operational adaptability. Unlike traditional single-device monitoring or single-load control methods, virtual power plants do not only focus on the instantaneous operating status of a certain type of resource, but also need to continuously characterize the availability, response activity level, state change trends, and continuous adjustment potential of internal resources, thereby providing fundamental support for resource aggregation, state diagnosis, capacity assessment, response allocation, and collaborative control.
[0003] However, in the actual operation of a virtual power plant, the adjustment level of resource response is not determined solely by the resource's intrinsic state, but is influenced by a combination of factors, including external system boundary conditions, available renewable energy output, reserve capacity adequacy, environmental and meteorological changes, user-side behavioral characteristics, and time-cycle attributes. For example, fluctuations in renewable energy output alter system operating boundaries and resource demand; changes in meteorological conditions not only affect the availability of wind and solar resources but also indirectly alter user-side energy consumption behavior; and changes in reserve capacity affect the urgency and adjustability of resource response. These factors differ significantly in data sources, sampling frequencies, time granularity, spatial scale, variable dimensions, and field semantics, and their effects on resource response behavior often exhibit nonlinearity, time delay, coupling, and state-switching characteristics.
[0004] Therefore, the analysis of resource response levels in virtual power plants is no longer a simple discrimination problem that can be solved by relying on single equipment measurements or single historical sequences. Instead, it is a complex technical problem involving the unified normalization of multi-source heterogeneous data, spatiotemporal dynamic mapping, key feature extraction, and causal relationship identification. Without establishing a unified data processing and relationship identification mechanism at the front end, it is difficult to accurately depict the dynamic transmission relationship between "external environment changes—system boundary states—resource response behavior," and it is also difficult to provide a stable and reliable data foundation for subsequent response status analysis, capacity assessment, and trend prediction. Existing technologies typically have the following shortcomings when dealing with this type of problem:
[0005] 1) Static splicing of multi-source heterogeneous data is unsuitable for the highly dynamic operation of virtual power plants. Existing methods typically treat system boundary data, environmental meteorological data, resource operation data, and time calendar data as ordinary external variables, simply aggregating or horizontally splicing them without establishing a unified data normalization mechanism to address issues such as missing values, anomalous disturbances, duplicate records, dimensional differences, sampling frequency mismatches, and inconsistent field semantics between different data sources. For example, system boundary information is usually released at fixed time intervals, meteorological and environmental data have spatial distribution differences, and resource-side operation data has strong randomness and state jump characteristics. If these data are directly input into subsequent models side by side, it can easily lead to problems such as inconsistent physical meanings of variables at the same time, distortion of time alignment, and amplification of local anomalous data. Existing technologies lack a data quality closed-loop repair mechanism for regulating resource response behavior in virtual power plants, resulting in insufficient data sample completeness and poor scale consistency, making it difficult to provide a stable, reliable, and reusable input foundation for subsequent response level analysis.
[0006] 2) Low-dimensional time series modeling struggles to characterize the spatiotemporal coupling and transmission relationships of multiple factors on the response state of regulatory resources. The response level of regulatory resources within a virtual power plant is not determined solely by the state of a single resource, but rather by the combined effects of multiple dimensions, including external system boundaries, available renewable energy output, reserve adequacy, environmental and meteorological conditions, and user-side behavior. Existing technologies often construct historical features from a single time series perspective or perform static correlation analysis on only some variables, lacking a unified characterization of the dynamic mapping relationships between the time, space, and state dimensions. For example, changes in renewable energy output first alter the system boundary state, further affecting the availability and responsiveness of resources within the virtual power plant; changes in meteorological conditions not only affect available renewable energy output but may also indirectly alter user-side access behavior and resource operation status. Existing methods struggle to reveal such multi-level transmission chains between "environmental state—system boundary—resource response," resulting in insufficient explanatory and predictive capabilities for changes in the response state of regulatory resources under complex operating scenarios.
[0007] 3) The lagged effect structure of external driving variables relies on empirical settings, making it difficult to accurately extract key causal links. In actual operation, the impact of variables such as load level, renewable energy output, reserve capacity, meteorological environment, and time cycle on the response level of virtual power plant regulation resources is usually not instantaneous, but rather involves varying degrees of time delay and cumulative effects. Existing technologies often use fixed-length time windows, manually set lag orders, or simple historical sliding features to describe such relationships, lacking a data-driven causal lag identification mechanism. Since the effect delays of different variables are not consistent, an empirical window that is too short may miss key driving information, while an empirical window that is too long will introduce a large number of invalid lag features, causing feature redundancy and relationship confusion. Especially when the response state of regulation resources changes rapidly, existing methods have difficulty distinguishing between statistically relevant variables and true driving variables, resulting in insufficient extraction of key causal parent nodes, action paths, and lag orders, thus limiting the interpretability and credibility of the response level analysis results.
[0008] 4) Lack of a unified closed loop between high-dimensional feature construction, effective feature selection, and causal enhancement analysis. While existing methods construct temporal features, historical state features, statistical features, or some artificially composite features, these features typically remain at the empirical design level, lacking a unified technical framework that connects them with feature importance assessment, causal link identification, and hierarchical response analysis. On the one hand, multi-source heterogeneous data, after expansion, easily forms a high-dimensional feature space. Without an effective screening mechanism, low-contribution variables, duplicate variables, and noise variables will enter a large number of subsequent models, reducing the information density of the input features. On the other hand, simply relying on feature importance ranking makes it difficult to explain whether there is a clear lagged driving relationship between variables, easily mistaking short-term correlations for stable relationships. Existing technologies fail to organically combine multi-source feature generation, contribution screening, causal link enhancement, and hierarchical response level analysis, resulting in insufficient stability, generalization ability, and interpretability of the model under complex state transitions, local anomalies, and multivariate coupling scenarios.
[0009] In view of the shortcomings of existing technologies, it is necessary to propose a new solution to meet practical application needs. Summary of the Invention
[0010] The purpose of this invention is to provide a multi-source data processing and causal correlation identification method for virtual power plant response level analysis, which can solve the following problems raised in the prior art.
[0011] This invention addresses the shortcomings of existing technologies, such as coarse processing of multi-source correlated data, unclear spatiotemporal mapping relationships, insufficient extraction of key influencing factors, and weak unified analysis support capabilities. It provides a stable and reliable data and feature foundation for subsequent resource response status analysis, capability assessment, and trend prediction. Specifically, this invention addresses the following technical problems, including:
[0012] 1. To address the issues of numerous missing values, inconsistent units, and insufficient direct modeling capability in the original data, a data preprocessing mechanism suitable for analyzing the response level of regulation resources in virtual power plants is constructed.
[0013] 2. To solve the problem that data from different sources are difficult to form a unified association, so that various types of related data can establish a consistent spatiotemporal alignment relationship and dynamic mapping relationship around the resource regulation response behavior of virtual power plants, thereby providing a unified data organization foundation for subsequent feature extraction and relationship identification;
[0014] 3. To address the issues of insufficient identification of time-lag relationships among multi-source heterogeneous variables, incomplete causal link construction, and inadequate extraction of key driving factors, this study aims to achieve causal association identification and causal lag feature construction for the analysis of the response level of regulating resources, thereby improving the interpretability and effectiveness of feature links.
[0015] 4. To address the issues of excessive redundant features and lack of emphasis on key driving factors in existing feature construction, a unified feature generation and filtering mechanism will be constructed to improve the quality and interpretability of the final input features.
[0016] To achieve the above objectives, this invention proposes a multi-source data processing and causal correlation identification method for virtual power plant response level analysis, providing the following technical solution:
[0017] The technical solution adopted by this invention to solve its technical problem is: a method for multi-source heterogeneous data processing and causal correlation identification for analyzing the response level of virtual power plant regulation resources, comprising the following steps:
[0018] S1: Construct a unified access and semantic regularization model for multi-source heterogeneous related data;
[0019] S2: Construct a data missing repair and normalization model based on low-rank structure constraints;
[0020] S3: Construct a spatiotemporal multidimensional fusion and dynamic mapping model for regulating resource response behavior;
[0021] S4: Construct a causal feature generation model based on contribution screening and time-delay causal identification;
[0022] S5: Construct an analysis model for regulating resource response levels based on a dual-headed hierarchical structure, and form a unified data foundation.
[0023] In the above technical solution, the specific method for S1 to construct a unified access and semantic regularization model for multi-source heterogeneous related data is as follows:
[0024] The S1 construction of a unified access and semantic regularization model for multi-source heterogeneous related data includes:
[0025] S1.1 Construct a multi-source heterogeneous data access set;
[0026] S1.2 Unified analysis of time granularity mapping;
[0027] S1.3 Construct a unified semantic tag model.
[0028] Specifically, the S1 construction of a unified access and semantic regularization model for multi-source heterogeneous related data is as follows:
[0029] S1.1 Construct a multi-source heterogeneous data access set;
[0030] To analyze the resource response level of a virtual power plant, a multi-source heterogeneous data access set is formed by integrating system boundary data, environmental meteorological data, resource-side operational data, and time calendar data. Specifically, the system boundary data includes at least predicted load demand, available renewable energy output, reserve capacity, and conventional power output arrangement information; the environmental meteorological data includes at least temperature, humidity, dew point, precipitation, wind speed, air pressure, and cloud cover information; the resource-side operational data includes at least charging equipment status, real-time power, and state of charge information; and the time calendar data includes at least time, day of the week, month, holidays, and weekday markings.
[0031] Let the set of data sources accessed during the virtual power plant operation analysis be .
[0032] ;
[0033] In the formula: This represents system boundary data, including information on predicted load demand, renewable energy output, reserve capacity, and conventional power output arrangements. It represents environmental meteorological data, including information on temperature, humidity, dew point, precipitation, wind speed, air pressure, and cloud cover; This indicates the operational data on the resource side, including information on charging equipment status, real-time power, and state of charge. It represents time calendar data, including information such as time, day of the week, month, holidays, and weekday markers.
[0034] For the m-th type of data source, its original observation sequence is represented as:
[0035] ;
[0036] In the formula: This represents the j-th variable in the m-th data source at the original time. Observed values; This indicates the number of variables contained in the m-th data source; This represents the set of original time indexes corresponding to the data source.
[0037] S1.2 Unified analysis of time granularity mapping;
[0038] To address the issue of inconsistent sampling frequencies across different data sources, all data is uniformly mapped to a fixed analysis granularity. The preferred interval is 15 minutes, and a unified time index set is constructed. :
[0039] ;
[0040] For any original variable Through time aggregation operator Complete unified granularity mapping:
[0041] ;
[0042] In the formula: This represents the mapping value of the j-th variable in the m-th data source at a uniform time t; This represents the time index of the original data; Select the mean, maximum, minimum, last, cumulative, or state-held value based on the data type. N is 96 because there are 24 hours in a day, and the granularity of analysis is fixed. If the time is 15 minutes, then there are 96 points in a day.
[0043] S1.3 Construct a unified semantic tag model;
[0044] To avoid inconsistencies in the meaning of fields with the same name or inconsistencies in the names of synonymous fields across different systems, each data variable is defined as a semantic quintuple. :
[0045] ;
[0046] In the formula: Indicates the data source attribute; Indicates the time attribute; Indicates spatial attributes, Indicates dimensional units; The variable represents the physical semantics; m represents the m-th type of data source; j represents the j-th data variable.
[0047] Through semantic mapping function This maps variables from different sources to a unified standard variable space.
[0048] ;
[0049] In the formula: This represents the system boundary data after unified time mapping; This represents environmental meteorological data after unified time mapping; This represents the resource-side runtime data after unified time mapping; This represents time calendar data after unified time mapping; t represents the time period.
[0050] This forms an initial unified sample matrix for analyzing the response level of regulation resources in virtual power plants:
[0051] ;
[0052] In the formula: the superscript T is the transpose symbol; This represents the time period at the Nth time scale.
[0053] The core of this step lies in establishing a unified multi-source data representation mechanism for analyzing the response level of virtual power plant regulation resources. Because system boundary data, environmental meteorological data, resource-side operational data, and time calendar data exhibit significant differences in sampling periods, spatial scales, units of measurement, and field semantics, directly inputting them into subsequent analysis models can easily lead to problems such as variable time mismatch, inconsistent physical meanings, and imbalanced feature scales. Therefore, this invention first maps data from different sources to the same analysis timeline and constructs a unified semantic labeling system, providing a foundation for subsequent data repair, spatiotemporal fusion, and causal relationship identification.
[0054] In the above technical solution, the S2 construction of a data missing repair and normalization model based on low-rank structure constraints includes:
[0055] S2.1 Construct an observation mask for the missing locations;
[0056] S2.2 Missing data repair based on low-rank matrix factorization;
[0057] S2.3 Physical constraint projection and anomaly correction;
[0058] S2.4 Gaussian distribution normalization processing.
[0059] Specifically, the S2 model for data missing data repair and normalization based on low-rank structure constraints is as follows:
[0060] S2.1 Construct an observation mask for the missing locations;
[0061] Let a uniform sample matrix be defined. Where N represents the number of time periods for the sample, and P represents the number of feature variables; define the observation mask matrix. :
[0062] ;
[0063] Then the effective observation set Represented as:
[0064] ;
[0065] In the formula: subscripts i and j represent the matrix Positioning of elements within a given context.
[0066] S2.2 Missing data repair based on low-rank matrix factorization;
[0067] The associated variables preferably include system boundary variables, environmental meteorological variables, resource-side operational variables, and continuous composite features constructed from them; for time calendar markers, categorical state variables, and equipment number variables, rule-based imputation, mode imputation, or independent encoding methods are used, and they do not directly participate in the low-rank repair of continuous variables; considering that multi-source associated variables are constrained by common operational states, their sample matrix can be decomposed into a low-rank principal structure matrix, a sparse anomaly matrix, and a random perturbation matrix:
[0068] ;
[0069] In the formula: This represents a low-rank principal structure matrix that reflects the main operating principles of the system. This represents the sparse anomaly perturbation matrix caused by local anomalies, mutation points, or erroneous acquisition. This represents a random disturbance term.
[0070] The low-rank repair model is represented as:
[0071] ;
[0072] In the formula: , and The superscript T represents the low-rank factorization factor; This indicates the observation position projection operator, which calculates the reconstruction error only at valid observation positions; and These are the low-rank regularization coefficient and the sparse anomaly constraint coefficient, respectively. This represents the sum of squares of all elements in the matrix.
[0073] S2.3 Physical constraint projection and anomaly correction;
[0074] For variables with a defined physical range, further apply a projection onto the physically feasible region:
[0075] ;
[0076] In the formula: and Let represent the lower and upper physical limits of the j-th variable, respectively; this processing is used to prevent the missing data repair result from deviating from the actual feasible range. This represents the feature variables after missing data repair in low-order matrix factorization, with subscripts i and j indicating the position of elements in the matrix.
[0077] S2.4 Gaussian distribution normalization;
[0078] To eliminate differences in the dimensions and orders of magnitude of different variables, this invention employs Gaussian distribution normalization:
[0079] ;
[0080] In the formula: and Let represent the mean and standard deviation of the j-th variable in the sample set, respectively; To prevent extremely small positive numbers with a denominator of zero; These are the feature variables projected from the physical constraints. These are the standardized feature variables.
[0081] After the above processing, the normalized sample matrix is obtained:
[0082] ;
[0083] Where N represents the number of sample times and P represents the number of feature variables.
[0084] The core of this step lies in establishing a multi-source heterogeneous data quality repair mechanism. Actual operational data often contains issues such as local missing data, anomalous jumps, duplicate records, and dimensional differences. Without systematic processing, these issues can easily lead to unstable results in subsequent causal identification and response level analysis. Therefore, by leveraging the potential low-dimensional structural relationships existing between multi-source data, low-rank repair is performed on missing data. This is combined with sparse anomaly separation and Gaussian distribution normalization to form a complete, stable, and scale-consistent normalized data matrix.
[0085] Furthermore, step S2 also includes:
[0086] To further improve the stability of repairing multi-source heterogeneous variables, the sample matrix is divided into several sub-matrices based on the physical properties of the variables and their data sources:
[0087]
[0088] In the formula: Represents the submatrix of system boundary variables; Represents a submatrix of environmental meteorological variables; This represents a submatrix of runtime variables on the resource side; This represents a submatrix of composite running state variables. For different submatrices, low-rank repair is performed separately, and then they are recombined using a unified time index to form a complete sample matrix. This block repair method can avoid the interference of weakly correlated variables on the low-rank structure and improve the stability and physical rationality of the missing data repair results.
[0089] By solving the above model, the repaired complete sample matrix is obtained. :
[0090] ;
[0091] In the formula: and These represent the low-rank decomposition factors obtained from the optimization solution; the superscript T is the transpose symbol.
[0092] In the above technical solution, the S3 construction of a spatiotemporal multidimensional fusion and dynamic mapping model for regulating resource response behavior includes:
[0093] S3.1 Spatial aggregation of regional environmental data;
[0094] S3.2 Time-period variable cyclic encoding;
[0095] S3.3 Construct composite operating state characteristics;
[0096] S3.4 Form a spatiotemporal fusion feature set.
[0097] Specifically, the S3 constructs a spatiotemporal multidimensional fusion and dynamic mapping model oriented towards regulating resource response behavior as follows:
[0098] S3.1 Spatial aggregation of regional environmental data;
[0099] For environmental variables collected from multiple spatial nodes, a statistical aggregation method is used to form regional environmental characteristics; let the k-th type of environmental variable of the r-th spatial node at time t be... The regional environmental characteristics Represented as:
[0100] ;
[0101] In the formula: R represents the number of spatial nodes; in this way, scattered environmental information can be transformed into a unified regional state description. This indicates the calculation of the average value; This indicates the calculation to obtain the maximum value; This indicates the calculation of taking the minimum value; This indicates the calculation of the standard deviation.
[0102] S3.2 Time-period variable cyclic encoding;
[0103] To avoid periodic breaks caused by directly using time numbers, a cyclic encoding method is used to represent time period variables:
[0104] ;
[0105] ;
[0106] In the formula: The current time represents the sequence number within a complete cycle; N represents the total number of time periods in the sample, and when the analysis granularity is 15 minutes, N=96; t represents the time period.
[0107] S3.3 Construct composite operating state characteristics;
[0108] To describe the coupling relationship between system boundary states and the behavior of regulating resource responses, this invention constructs a net demand intensity. New energy penetration rate , reserve capacity and climbing intensity Composite characteristics:
[0109] ;
[0110] ;
[0111] ;
[0112] ;
[0113] In the formula: Indicates the load demand of the external system; and These represent the predicted available wind and solar power output, respectively. Indicates the positive reserve capacity; It is a very small positive number; t represents the time period.
[0114] S3.4 Forms a spatiotemporal fusion feature set;
[0115] Finally, by integrating standardized data, regional environmental characteristics, time cycle characteristics, and composite operational status characteristics, a basic fusion feature vector is formed. :
[0116] ;
[0117] In the formula: P represents the number of characteristic variables; This represents the value of the Pth variable in time period t.
[0118] The core of this step lies in characterizing the dynamic coupling relationship between the external environment, system boundaries, and resource response behavior. Since the resource response state of a virtual power plant is simultaneously affected by time period, spatial distribution, and operating status, this invention constructs spatiotemporal fusion features, periodic coding features, and composite operating status features based on standardized data to form a dynamic mapping feature set that can characterize the mechanism of changes in the resource response level.
[0119] In the above technical solution, the S4 method for constructing a causal enhancement feature generation model based on contribution screening and time-delay causal identification includes:
[0120] S4.1 Effective Feature Filtering Based on SHAP Contribution;
[0121] S4.2 Construct a set of candidate time-delay causal edges;
[0122] S4.3 Conditional independence test based on LPCMCI.
[0123] Specifically, step S4 constructs a causal enhancement feature generation model based on contribution screening and time-delay causal identification as follows:
[0124] S4.1 Effective Feature Filtering Based on SHAP Contribution;
[0125] The SHAP method is used to calculate the marginal contribution of each candidate feature to the target response level output; let the SHAP contribution value of the j-th feature in the i-th sample be . Then the global contribution of the j-th feature Defined as:
[0126] ;
[0127] In the formula: This represents the global contribution strength of the j-th feature to the target response level; Indicates the number of training samples
[0128] All candidate features according to Sort the features from largest to smallest to obtain a contribution ranking set, and then select the top K key features based on their contribution rates to obtain the effective feature set. :
[0129] ;
[0130] In the formula: This represents the feature that ranks Kth in the contribution ranking.
[0131] This process is used to filter out low-contribution and weak explanatory variables from high-dimensional fusion features, while retaining candidate driving variables that contribute significantly to the target response state, thereby reducing dimensionality and noise interference for subsequent causal link identification.
[0132] S4.2 Construct a set of candidate time-delay causal edges;
[0133] Considering that the effects of different external variables on the adjustment of resource response levels do not all occur at the same time, but may have varying degrees of time delay, in the effective feature set Construct a set of multi-lag candidate variables based on :
[0134] ;
[0135] In the formula: Indicates the lag order; This represents the maximum lag order, determined based on the analysis granularity and resource response time constant; when the analysis granularity is 15 minutes... This indicates a delay of 1 hour; This represents a candidate valid feature variable.
[0136] any candidate variable Record Construct its target response level Candidate time-delay causal edges:
[0137] ;
[0138] The candidate edge indicates that the j-th variable is in lag. During a given time period, it may have a driving effect on the target response level at the current moment.
[0139] S4.3 Conditional independence test based on LPCMCI;
[0140] To avoid misjudging statistical correlations caused by common periodic factors, common environmental disturbances, resource status linkages, or the transmission of intermediate variables as direct driving relationships, the Parcorr partial correlation conditional independence test method is introduced into the LPCMCI time-delay causal search framework to screen candidate time-delay causal edges.
[0141] First, the LPCMCI algorithm constructs a set of condition variables based on historical dependencies and candidate parent node search results:
[0142] ;
[0143] In the formula: This represents the set of conditional variables used to eliminate confounding effects, including the historical terms of the target variable itself, lagged terms of other candidate driving variables, time period terms, and associated variables that have a common driving relationship with the candidate variables; This represents the number of candidate condition variables.
[0144] Using the set of conditional control variables respectively For candidate driving variables and target response variable By performing regression fitting, we can obtain its conditional predicted values:
[0145] ;
[0146] ;
[0147] In the formula: and represents the fitted functions obtained by conditional regression on the candidate driving variable and the target response variable, respectively; t represents the time period; and These are the predicted values after regression fitting.
[0148] Furthermore, step S4.3 also includes:
[0149] Define the residual variable after removing the influence of the condition variable as follows:
[0150] ;
[0151] ;
[0152] In the formula: and The residual variable after removing the influence of conditional variables; t represents the time period; The actual value of the candidate variable; The actual value of the target response variable; and These are the predicted values after regression fitting.
[0153] Based on this, the partial correlation coefficient between the two residual sequences is calculated:
[0154] ;
[0155] In the formula: Indicates the set of control condition variables After that, candidate driving variables With the target response variable The net correlation strength that remains between them; and ...
[0156] when When the value is close to 0, it indicates that the linear residual correlation between the candidate driving variables and the target response variable is weak after controlling the set of condition variables; when... A larger value indicates a stronger net linear relationship between the two variables even after controlling for conditional variables. According to... The significance level can be calculated. If the following conditions are met:
[0157] ;
[0158] In the formula: Indicates the confidence level.
[0159] Then it is considered that in the set of control condition variables After that, candidate driving variables With the target response variable Significant conditional dependencies still exist between them, and candidate time-delay causal edges are preserved. .
[0160] Furthermore, step S4 also includes:
[0161] S4.4 Causal Influence Estimation and Feature Selection Based on Optimal Adjustment Set;
[0162] Furthermore, an optimal adjustment set mechanism is introduced to estimate the causal effect strength of candidate driving variables at corresponding lag orders. The purpose of this step is to avoid relying solely on the residual correlation strength to judge the magnitude of the variable's effect, but rather to quantitatively characterize the net effect of changes in candidate variables on changes in the target response level while minimizing the influence of confounding variables. Through this process, candidate edges that are statistically significant but have a weak actual effect are eliminated, thereby improving the effectiveness and interpretability of the input features of the subsequent tree model.
[0163] First, to reduce the interference of common driving factors, the historical state of the target variable itself, and other related variables on the estimation of causal effects, an optimal set of adjustment variables is constructed for each retained time-delay causal edge:
[0164] ;
[0165] In the formula: Indicates the causal edge The selected optimal adjustment set; represents the m-th adjustment variable; t represents the time period. The optimal adjustment set can be composed of the candidate parent node set obtained by LPCMCI search, the historical terms of the target variable, other significantly lagged driving variables, and state variables that have a common influence relationship with the candidate variables.
[0166] Then, a random forest nonlinear estimator is used to fit the causal effect function based on the optimal adjustment set. :
[0167] ;
[0168] In the formula: Indicates the candidate edge A random forest causal effect estimator is constructed; x represents the input driving variable value; t represents the time period. The random forest causal effect estimator fits the total effect relationship of candidate driving variables on the target response level under the constraint of the optimal adjustment set, and is used to describe the net effect of changes in candidate variables on the target response level after removing confounding factors;
[0169] To facilitate comparison of the strength of causal effects among different variables, two symmetrical intervention levels are set in the standardized variable space. and Let represent the states where the candidate driving variable is above or below the mean by one standard unit, respectively. Based on the fitted total effect function, the estimated target response values at the two intervention levels can be calculated separately:
[0170] ;
[0171] ;
[0172] In the formula: This indicates that the estimated target response value is higher than one standard unit of intervention level for the candidate driving variable; This indicates that the estimated target response value is lower than one standard unit of intervention level for the candidate driving variable;
[0173] Therefore, the causal effect strength of candidate time-delay causal edges is defined. for:
[0174] ;
[0175] when This indicates that the candidate variable has a positive driving effect on the target response level at the corresponding lag order; when This indicates that the candidate variable has a negative inhibitory effect on the target response level at the corresponding lag order; The larger the value, the stronger the net causal effect of the candidate variable on the target response level at the corresponding lag order.
[0176] Furthermore, step S4 also includes:
[0177] S4.5 Constructs a set of causal enhanced features:
[0178] To further eliminate candidate edges with weak actual effects, a threshold for causal effect strength is set. Only when the causal strength of the candidate edge is greater than Only when the candidate time-delay causal edge is selected will it be retained, and its corresponding lagged variable will be used as an effective causal lag feature. The final set of causal lag features is formed after passing the LPCMCI conditional independence test, the optimal adjustment set identifiability test, and the causal effect strength screening. for:
[0179] ;
[0180] In the formula: Indicates the first q The first effective causal lag feature, namely the first... Candidate variables in lag The corresponding historical values It has a significant causal effect on the change of the target variable at the current time t;
[0181] The core of this step lies in resolving the problems of redundancy in high-dimensional features and unclear delayed driving relationships. This invention first utilizes a tree model combined with SHAP contribution to effectively filter high-dimensional features, eliminating low-contribution and weakly correlated variables; then, it employs the LPCMCI causal identification method to identify the time-delayed action paths of different variables on the target response state, and constructs causal features based on the strength of causal effects.
[0182] In the above technical solution, the S5 method for constructing an analysis model for adjusting resource response levels based on a dual-headed hierarchical structure and forming a unified data foundation includes:
[0183] S5.1 Constructing unified core input features;
[0184] S5.2 Construct a baseline response state identification head;
[0185] S5.3 Construct a continuous analysis head for non-benchmark response levels;
[0186] S5.4 Constructs a probabilistic fusion and hard-switching output mechanism;
[0187] S5.5 forms a unified data foundation.
[0188] Specifically, S5 constructs an analysis model for adjusting resource response levels based on a dual-headed hierarchical structure, and forms a unified data foundation as follows:
[0189] S5.1 Constructing unified core input features;
[0190] The spatiotemporal fusion features obtained in S3, the effective filtering features obtained in S4, and the causal enhancement features are combined to form the core input vector:
[0191] ;
[0192] In the formula: This is the feature set after filtering effective features based on SHAP contribution. This is the set of causal lag features after filtering.
[0193] S5.2 Construct a baseline response state identification head;
[0194] This invention first constructs a baseline response state recognition head to determine whether the current sample is near the baseline response capability state. Let the target response level be... The baseline response level is The threshold of the baseline state neighborhood is Then the baseline response status label is defined as:
[0195] ;
[0196] In the formula: This indicates that the current sample is near the baseline response state; This indicates that the current sample is in the non-benchmark response interval; t represents the time period.
[0197] The baseline response state recognition head is represented by a tree-structured classification model as follows:
[0198] ;
[0199] In the formula: This represents the baseline response state identification model; The feature input vector representing time period t; This represents the probability that the current sample belongs to the baseline response state. This probability is used to characterize the degree to which the current state of the adjustment resource is close to the baseline response capability state, and serves as the state gating signal for the subsequent hierarchical fusion output.
[0200] S5.3 Construct a continuous analysis head for non-benchmark response levels;
[0201] For samples that deviate from the baseline response state, this invention further constructs a continuous analysis head for non-baseline response levels to characterize the continuous variation of the adjustment resource response level within the non-baseline range.
[0202] Let the non-benchmark sample set be... for:
[0203] ;
[0204] In the formula: The label indicates the baseline response status; t indicates the time period.
[0205] The continuous analysis model for non-benchmark response levels is then expressed as:
[0206] ;
[0207] In the formula: This represents a continuous analysis model for non-benchmark response levels. The feature input vector representing time period t; This represents the estimated response level under non-baseline conditions.
[0208] S5.4 Constructs a probabilistic fusion and hard-switching output mechanism;
[0209] After completing the training of the two analysis heads, this invention will use the baseline response state probability. Compared with non-benchmark response level estimates The results of the fusion were obtained to obtain the final response level analysis results:
[0210] ;
[0211] In the formula: As the baseline response level; Indicates the hard handover threshold for the baseline response state; This represents the final response level analysis result; t represents the time period. This mechanism can reduce unreasonable offsets near the baseline state, improve the model's ability to identify the baseline state concentration interval, and enhance output stability. When the baseline response state probability... Not less than the threshold When the system directly outputs the baseline response level; when Less than the threshold At that time, the system outputs the continuous response level after probability fusion.
[0212] S5.5 forms a unified data foundation;
[0213] Finally, the multi-source input data, spatiotemporal fusion features, LPCMCI time-delay causal features, SHAP contribution screening results, dual-head hierarchical model output results and evaluation results are all written into the data foundation base.
[0214] The data infrastructure includes the following levels:
[0215] First, the raw associated data layer is used to store system boundary data, environmental data, resource-side status data, and time calendar data;
[0216] Second, the feature construction layer is used to store time period features, historical state features, artificial composite state features, and LPCMCI time-delay causal features.
[0217] Third, the feature selection layer is used to store the SHAP contribution ranking results, the selected core feature set, and feature explanation information;
[0218] Fourth, the state recognition layer, used to store the baseline response state probabilities. State switching threshold and baseline state identification results;
[0219] Fifth, the response analysis layer, used to ultimately predict the response level output. and actual observed response level .
[0220] The core of this step lies in uniformly inputting the multi-source heterogeneous data processing results, spatiotemporal fusion characteristics, LPCMCI time-delay causal characteristics, and key features after contribution screening into the response level analysis model. This enables hierarchical analysis of the response status of virtual power plant regulation resources, the proximity of the baseline response status, the continuous response level in non-baseline intervals, and subsequent trends. Since virtual power plant regulation resources typically exhibit significant state differentiation during actual operation—that is, remaining near the baseline response capability state for extended periods while deviating from the baseline state and exhibiting continuous fluctuations under the influence of external boundary conditions, environmental conditions, and resource-side behavior changes—analyzing using only a single continuous model is easily affected by the concentration of baseline state samples, resulting in insufficient characterization of the variation amplitude in non-baseline response intervals. Therefore, this invention constructs a dual-head hierarchical structure consisting of a baseline response state identification head and a non-baseline response level analysis head.
[0221] By adopting the above technical solution, the present invention can produce the following beneficial effects:
[0222] 1. This invention solves the technical problem of unifying the representation of multi-source heterogeneous data in the analysis of virtual power plant regulation resource response levels. Through unified time granularity mapping, semantic normalization, low-rank missing data repair, physical constraint projection, and Gaussian distribution normalization, this invention transforms system boundary information, environmental meteorological information, resource-side operational information, and time calendar information into a unified, complete, and computable data sample matrix, significantly improving the integrity, consistency, and modelability of the original data.
[0223] 2. A spatiotemporal dynamic mapping and causal lag feature extraction mechanism for regulating resource response behavior was constructed. By constructing composite features such as regional environmental feature aggregation, time period encoding, net demand intensity, reserve adequacy, new energy penetration rate, and ramp-up intensity, this invention can characterize the dynamic coupling relationship between the external environment, system boundary, and resource response behavior. Furthermore, by utilizing the LPCMCI conditional independence test and the optimal adjustment set causal effect strength estimation, causal features with significant lag driving effects are screened out, reducing the arbitrariness caused by manually setting the lag window based on experience.
[0224] 3. Improved effective information density and interpretability of high-dimensional feature inputs. This invention utilizes a tree model combined with SHAP contribution analysis to screen candidate features, eliminating low-contribution variables, duplicate variables, and weak explanatory variables, ensuring that the final feature set input to the model possesses both statistical contribution and causal interpretability. This method can reduce the impact of high-dimensional feature redundancy on model stability and provide explanations for key driving factors in response level changes.
[0225] 4. A dual-head hierarchical response analysis structure suitable for scenarios where both baseline states are concentrated and non-baseline states continuously change. This invention uses a baseline response state identification head to determine whether a sample is near the baseline response state, and a non-baseline response level analysis head to characterize the continuous change pattern after deviating from the baseline state. Combined with a state switching threshold, hierarchical fusion output is achieved. This structure can reduce unreasonable fluctuations near the baseline state and improve the accuracy and stability of response level analysis in non-baseline intervals.
[0226] 5. A traceable and reusable unified data foundation has been established. This invention unifies the original correlation data, standardized data, spatiotemporal fusion features, LPCMCI time-delay causal features, SHAP contribution screening results, baseline response state probabilities, and final response level outputs into a unified data foundation, enabling full-process traceability from multi-source data input to response level analysis result output, and providing continuous support for virtual power plant resource status assessment, response allocation, and collaborative control. Attached Figure Description
[0227] Figure 1 This is a flowchart illustrating the overall process of the method involved in this patent.
[0228] Figure 2 This is a time-delay causal analysis diagram (time delay is 0) related to this patent.
[0229] Figure 3 This is a time-delay causal analysis diagram related to this patent (time delay is 15h);
[0230] Figure 4 This is a time-delay causal analysis diagram (time delay is 24h) related to this patent.
[0231] Figure 5 A graph showing the ranking of the contribution of the key driving factors involved in this patent.
[0232] Figure 6 This is a probability curve of the baseline response state involved in this patent;
[0233] Figure 7 This is a graph showing the overall test cycle response level analysis results related to this patent.
[0234] Figure 8This is a schematic diagram of the five layers of the data foundation base involved in this patent. Detailed Implementation
[0235] The following will refer to the appendices in the embodiments of the present invention. Figure 1-8 The technical solutions in the embodiments of the present invention are clearly and completely described herein. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0236] A method for processing multi-source heterogeneous data and identifying causal relationships for analyzing the regulation resource response level of virtual power plants is as follows:
[0237] S1, Construct a unified access and semantic regularization model for multi-source heterogeneous related data;
[0238] S2, Construct a data missing repair and normalization model based on low-rank structure constraints;
[0239] S3, construct a spatiotemporal multidimensional fusion and dynamic mapping model for regulating resource response behavior;
[0240] S4, Construct a causal feature generation model based on contribution screening and time-delay causal identification;
[0241] S5 constructs an analysis model for regulating resource response levels based on a dual-headed hierarchical structure, and forms a unified data foundation.
[0242] Furthermore, S1 includes:
[0243] S1.1 Constructing a multi-source heterogeneous data access set:
[0244] To analyze the resource response level of a virtual power plant, a multi-source heterogeneous data access set is formed by integrating system boundary data, environmental meteorological data, resource-side operational data, and time calendar data. The system boundary data includes predicted load demand, available renewable energy output, reserve capacity, and conventional power output scheduling information. The environmental meteorological data includes temperature, humidity, dew point, precipitation, wind speed, air pressure, and cloud cover information. The resource-side operational data includes charging equipment status, real-time power, and state of charge information. The time calendar data includes time, day of the week, month, holidays, and weekday markings.
[0245] S1.2 Unified analysis of time granularity mapping:
[0246] To address the issue of inconsistent sampling frequencies across different data sources, all types of data are uniformly mapped to a fixed analysis time granularity of 15 minutes, and a unified time index set is constructed. For any original variable, time aggregation operators such as mean, maximum, minimum, last, cumulative, or state-maintained value are used according to its data type to obtain the mapped value of the variable at a unified time.
[0247] S1.3 Constructing a unified semantic tagging model:
[0248] To address the issue of inconsistent meanings or names of fields with the same name in different data systems, each data variable is defined as a semantic label containing data source attributes, time attributes, spatial attributes, units of measurement, and the physical semantics of the variable. Then, a semantic mapping function is used to uniformly map variables from different sources to a standard variable space, thereby forming an initial unified sample matrix for analyzing the response level of regulation resources in virtual power plants.
[0249] Furthermore, S2 includes:
[0250] S2.1 Constructing an observation mask for missing locations:
[0251] Based on the initial unified sample matrix obtained from S1, an observation mask matrix is defined to mark whether there are valid observation values for each variable at each time step, and the valid observation set is determined accordingly.
[0252] S2.2 Missing data repair based on low-rank matrix decomposition:
[0253] The sample matrix corresponding to the multi-source correlation variables is decomposed into a low-rank principal structure matrix, a sparse anomaly perturbation matrix, and a random perturbation matrix. The low-rank principal structure matrix is used to reflect the main operating rules of the system, the sparse anomaly perturbation matrix is used to characterize the perturbations caused by local anomalies, mutation points, or erroneous acquisition, and the random perturbation matrix is used to characterize random noise. By minimizing the reconstruction error of the effective observation position and introducing a low-rank regularization term and a sparse anomaly constraint term, the repaired complete sample matrix is obtained.
[0254] S2.3 Constructing a block-based low-rank repair mechanism:
[0255] Based on the physical properties of the variables and the data sources, the sample matrix is divided into a system boundary variable submatrix, an environmental meteorological variable submatrix, a resource-side operational variable submatrix, and a composite operational state variable submatrix. After low-rank repair, the submatrix is recombined using a unified time index to form a complete sample matrix, thereby reducing the interference of weakly correlated variables on the low-rank structure.
[0256] S2.4 Performs physical constraint projection and anomaly correction:
[0257] For variables with a clear physical range, apply a physical feasible region projection so that the repaired variable value is between the physical lower limit and physical upper limit of the corresponding variable, preventing the missing data repair result from deviating from the actual feasible range;
[0258] S2.5 is normalized to a Gaussian distribution:
[0259] Gaussian distribution normalization is performed on the continuous variables after missing data repair and physical constraint projection to eliminate differences in the dimensions and orders of magnitude of different variables, resulting in a normalized sample matrix.
[0260] Furthermore, S3 includes:
[0261] S3.1 Regional Environmental Data Spatial Aggregation:
[0262] For environmental variables collected from multiple spatial nodes, a statistical aggregation method is used to form regional environmental characteristics. The statistical aggregation method includes the calculation of mean, maximum, minimum and standard deviation, thereby transforming the scattered environmental information into a unified regional state description.
[0263] S3.2 Time Period Variable Cyclic Encoding:
[0264] For periodic variables such as time, day of the week, and month, sine and cosine functions are used for cyclic encoding to avoid period breakage caused by directly using time numbers; when the analysis time granularity is 15 minutes, the total number of time periods in a complete daily cycle is taken as 96.
[0265] S3.3 Constructing Composite Operating State Characteristics:
[0266] Based on the coupling relationship between system boundary states and resource response behavior, composite operational state characteristics such as net demand intensity, renewable energy penetration rate, reserve adequacy, and ramp-up intensity are constructed. Net demand intensity is obtained by subtracting available wind power and solar power output from external system load demand; renewable energy penetration rate is obtained by the ratio of the sum of available wind power and solar power output to external system load demand; reserve adequacy is obtained by the ratio of positive reserve capacity to external system load demand; and ramp-up intensity is obtained by the difference in net demand intensity between adjacent time points.
[0267] S3.4 Formation of a spatiotemporal fusion feature set:
[0268] By integrating the standardized sample matrix, regional environmental characteristics, time cycle characteristics, and composite operational status characteristics, a basic fusion feature vector is formed for analyzing the response level of regulation resources in virtual power plants.
[0269] Furthermore, S4 includes:
[0270] S4.1 Effective feature selection based on SHAP contribution:
[0271] A tree model is used to fit the mapping relationship between candidate features and target response levels, and the SHAP method is used to calculate the marginal contribution of each candidate feature to the target response level output. The candidate features are sorted from largest to smallest according to their global contribution, and key features are selected according to a preset contribution rate or a preset number of features to obtain an effective feature set.
[0272] S4.2 Constructing a set of candidate time-delay causal edges:
[0273] Based on the effective feature set, a set of multi-lag candidate variables is constructed for each candidate feature, and the lag order is limited by the maximum lag order; the relationship between the candidate variables and the current target response level under different lag orders is used as candidate time-lag causal edges to characterize the delayed driving effect of external variables on the adjustment of resource response level.
[0274] S4.3 Perform conditional independence test based on LPCMCI:
[0275] The LPCMCI time-delay causal search method, combined with partial correlation conditional independence test, is used to screen candidate time-delay causal edges. During the test, a set of conditional variables is constructed to eliminate confounding effects. This set of conditional variables includes the target variable's own historical terms, lagged terms and time period terms of other candidate driving variables, and associated variables that have a common driving relationship with the candidate variables. When a candidate driving variable and the target response variable still have a significant conditional dependency after controlling for the set of conditional variables, the corresponding candidate time-delay causal edge is retained.
[0276] S4.4 Estimation of causal effect strength based on the optimal adjustment set:
[0277] For each retained time-delay causal edge, an optimal set of adjustment variables is constructed to reduce the interference of common driving factors, the historical state of the target variable itself, and other related variables on the estimation of causal effects; a random forest nonlinear estimator is used to fit the total effect relationship of candidate driving variables on the target response level under the constraint of the optimal adjustment set.
[0278] S4.5 Constructs a set of causal enhanced features:
[0279] Two symmetric intervention levels are set in the standardized variable space. The target response estimates corresponding to the candidate variables at different intervention levels are calculated respectively, and the difference between the two is used as the causal effect strength of the candidate time-lag causal edge. When the causal effect strength is greater than the preset threshold, the candidate edge is retained, and its corresponding lagged variable is used as the effective causal lagged feature, thus forming a set of causal enhancement features.
[0280] Furthermore, S5 includes:
[0281] S5.1 Constructing a unified core input feature:
[0282] The spatiotemporal fusion features obtained from S3, the effective screening features obtained from S4, and the causal enhancement features are combined to form the core input vector of the analysis model for regulating resource response level.
[0283] S5.2 Constructing a baseline response state identification head:
[0284] Define the target response level, the baseline response level, and the baseline state neighborhood threshold, and define the baseline response state label. When the target response level is within the neighborhood of the baseline response level, the sample is marked as the baseline response state. When the target response level deviates from the neighborhood of the baseline response level, the sample is marked as a non-baseline response state. A tree-structured classification model is used to construct a baseline response state recognition head, which outputs the probability that the current sample belongs to the baseline response state.
[0285] S5.3 Constructing a continuous analysis head for non-benchmark response levels:
[0286] For samples that deviate from the baseline response state, a non-baseline sample set is constructed, and a continuous analysis model of the non-baseline response level is trained based on the core input vector to characterize the continuous change pattern of the regulatory resource response level in the non-baseline range.
[0287] S5.4 constructs a probabilistic fusion and hard-switching output mechanism:
[0288] The baseline response state probability output by the baseline response state recognition head is fused with the response level estimate output by the non-baseline response level continuous analysis head to obtain the final response level analysis result. When the baseline response state probability is not less than the preset hard handover threshold, the baseline response level is directly output. When the baseline response state probability is less than the preset hard handover threshold, the continuous response level after probability fusion is output.
[0289] S5.5 forms a unified data foundation:
[0290] The multi-source input data, normalized sample matrix, spatiotemporal fusion features, SHAP contribution screening results, LPCMCI time-delay causal features, baseline response state probability, state switching threshold, final response level analysis results, and actual observed response levels are all written into the data foundation to achieve full-process traceability from multi-source data access to response level analysis result output.
[0291] Example:
[0292] The process of processing and response level analysis of multi-source heterogeneous correlation data based on a virtual power plant scenario in a certain province;
[0293] 1. Access and semantic regularization of multi-source heterogeneous correlated data;
[0294] The system first accesses multi-source heterogeneous data from a virtual power plant scenario in a certain province. This data includes system boundary data, environmental meteorological data, resource-side operational data, and time calendar data. Specifically, the system boundary data includes information such as load forecasting, renewable energy output forecasting, positive reserve capacity, negative reserve capacity, and conventional power output arrangements; the environmental meteorological data includes information such as temperature, humidity, dew point, precipitation, wind speed, air pressure, and cloud cover; the resource-side operational data includes information such as charging equipment status and real-time power; and the time calendar data includes information such as time of day, weekday, month, holiday markers, and weekday markers.
[0295] To address the differences in sampling periods, time granularity, variable units, and field semantics among various data sources, the system uniformly maps all raw data to a 15-minute analysis granularity and establishes a unified time index. Simultaneously, semantic labeling is standardized for variables from different sources, unifying the data source attributes, time attributes, spatial attributes, units of measurement, and physical semantics of the variables to form an initial unified sample matrix for analyzing the response level of virtual power plant regulation resources.
[0296] 2. Data missing data repair and standardization;
[0297] After obtaining the initial unified sample matrix, the system further constructs a data missing data repair and normalization model based on low-rank structure constraints. First, an observation mask for missing locations is constructed based on the observation status of each variable in the sample matrix to mark valid observations and missing values. Then, the sample matrix composed of continuous multi-source correlated variables is decomposed into a low-rank principal structure matrix, a sparse anomaly perturbation matrix, and a random perturbation matrix, and the missing data is repaired using the low-rank matrix decomposition method.
[0298] To improve the stability and physical rationality of the repair results, the system divides the sample matrix into sub-matrices for system boundary variables, environmental and meteorological variables, resource-side operational variables, and composite operational status variables, based on the source and physical attributes of the variables, and performs low-rank repair on each sub-matrix. For variables with a clear physical range, the system further applies physical feasible region constraints to prevent the repaired values from exceeding the actual feasible range. After completing missing value repair and anomaly correction, the Gaussian distribution normalization method is used to standardize each continuous variable, eliminating differences in the dimensions and orders of magnitude of different variables, ultimately obtaining a complete, stable, and scale-consistent normalized sample matrix.
[0299] 3. Spatiotemporal multidimensional fusion and dynamic mapping;
[0300] The system constructs a spatiotemporal multidimensional fusion and dynamic mapping model based on a normalized sample matrix, oriented towards regulating resource response behavior. For environmental meteorological data collected from multiple representative cities or spatial nodes, the system uses statistical methods such as mean, maximum, minimum, and standard deviation for spatial aggregation, forming regional-level environmental state characteristics to characterize the overall environmental changes in the area where the virtual power plant is located. For time-periodic information such as time, day of the week, month, and holidays, the system uses one-hot encoding or cyclic encoding for feature representation. Specifically, for time variables with obvious periodicity, sine and cosine cyclic encoding is used to avoid periodic breaks caused by directly using time numbering. At a 15-minute analysis granularity, a complete daily cycle includes 96 time periods.
[0301] Furthermore, based on the coupling relationship between the system boundary state and the response behavior of regulatory resources, the system constructs composite operational state characteristics such as net demand intensity, renewable energy penetration rate, reserve adequacy, and ramp-up intensity. Net demand intensity characterizes the equivalent regulatory demand after deducting available renewable energy output from external demand; renewable energy penetration rate characterizes the proportion of renewable energy output in the system boundary state; reserve adequacy characterizes the system's available regulatory margin; and ramp-up intensity describes the magnitude of net demand change between adjacent time points. Through the construction of these features, a basic fusion feature set is formed that reflects the dynamic coupling relationship between the external environment, system boundary state, and the response behavior of regulatory resources.
[0302] 4. Effective feature selection and time-delay causal link identification;
[0303] Building upon the basic fusion feature set, the system further constructs a causal enhancement feature generation model based on contribution-based selection and time-delay causal identification. First, a tree-structured model is used to fit the mapping relationship between candidate features and the target response level, and the SHAP method is used to calculate the marginal contribution of each candidate feature to the analysis results of the target response level. Features are then ranked according to their global contribution, selecting key features with high contributions and eliminating low-contribution variables, duplicate variables, and weak explanatory variables to obtain an effective feature set.
[0304] Subsequently, the system constructs a multi-lag candidate variable set based on the effective feature set, and uses the relationships between different variables at different lag orders pointing to the target response level as candidate time-lag causal edges. For these candidate time-lag causal edges, the LPCMCI causal identification method is used to perform conditional independence testing. Under the condition of controlling for the target variable's own historical terms, the lagged terms of other candidate driving variables, the time period terms, and common driving variables, the system determines whether a significant conditional dependency still exists between the candidate variables and the target response level.
[0305] For candidate causal edges that pass the conditional independence test, the system further introduces an optimal adjustment set mechanism and employs a random forest nonlinear estimator to estimate the causal strength of candidate driving variables on the target response level. By setting a symmetric intervention level, the difference between the estimated target response level under different intervention states is calculated, and this difference is used as an indicator of the causal strength. When the causal strength is greater than a preset threshold, the corresponding time-delayed causal edge is retained, and its lagged variables are used as effective causal enhancement features. Thus, the system forms a core input feature set that combines statistical contribution, time-delayed driving force, and causal explanatory power.
[0306] 5. Unified data foundation and dual-headed hierarchical response level analysis;
[0307] The system writes the core input features, after preprocessing, spatiotemporal fusion, contribution filtering, and causal enhancement, into a unified data foundation. This data foundation includes a raw correlated data layer, a normalized data layer, a feature construction layer, a feature filtering layer, a causal feature layer, a state identification layer, and a response analysis layer, enabling full-process traceability from multi-source data access to the output of response-level analysis results.
[0308] In the response level analysis phase, the system constructs a resource response level analysis model based on a dual-head hierarchical structure. The first head is the baseline response state identification head, which uses an XGBoost classification model or other tree-structured classification models to determine whether the current sample is near the baseline response capability state and outputs the baseline response state probability. The second head is the non-baseline response level continuous analysis head, which uses an XGBoost regression model or other regression models to perform continuous response level analysis on samples deviating from the baseline response state. When the model outputs, the system combines the baseline response state probability and the non-baseline response level estimation results. When the baseline response state probability is not less than a preset hard-switching threshold, the system directly outputs the baseline response level to reduce unreasonable fluctuations near the baseline state; when the baseline response state probability is less than the preset hard-switching threshold, the system outputs the continuous response level result after probability fusion. Through this dual-head hierarchical structure, the system can simultaneously adapt to two types of operating scenarios: concentrated baseline response states and continuous changes in non-baseline states.
[0309] To demonstrate the effectiveness and superiority of the multi-source heterogeneous data processing and causal relationship identification method proposed in this invention for analyzing the regulation resource response level of virtual power plants, a case study analysis can be conducted based on actual multi-source operating data, and the results can be illustrated with figures in conjunction with the model output.
[0310] first, Figures 2 to 4The results of time-delay causal analysis of LPCMCI under different lag scales are presented. As shown in the figures, causal connections with the average, minimum, and maximum responses exist in different directions and intensities between variables such as load forecasting, bidding space, reserve space, positive reserve space, wind power, gas turbines, and daily cycles, and the causal paths differ under different lag orders. For example, when the time lag is 0, there is a relatively obvious synchronous transmission relationship between load forecasting, reserve space, and positive reserve space; when the time lag is 15 hours, wind power shows a strong influence on the minimum response; and when the time lag is 24 hours, daily cycles, wind power, gas turbines, and reserve-related variables all participate in the lag transmission of response levels. These results demonstrate that this invention can identify the differentiated time-delay effects of multi-source variables on the regulation of resource response levels, rather than simply using a fixed historical window for empirical modeling.
[0311] Secondly Figure 5 The results show the ranking of the contributions of key driving factors. As can be seen from the figure, the average SHAP contribution value of renewable energy penetration rate is the highest, reaching 22.0657, significantly higher than other variables, indicating that changes in the proportion of renewable energy output are the main factor affecting the response level analysis results. Variables such as bidding space, squared renewable energy penetration rate, dew point temperature, response level at the same time yesterday, and response level at the same time the day before yesterday also have high contributions, indicating that system boundary conditions, environmental meteorological conditions, and historical response levels jointly influence the regulation of resource response behavior. These results demonstrate that this invention can screen out the main driving factors from high-dimensional multi-source features, improving the information density and interpretability of the input features.
[0312] again, Figure 6 The baseline response state probability curve is presented. The state switching threshold in the figure is set to 0.6. When the probability of the baseline response state is higher than this threshold, the model determines that the current sample is closer to the baseline response state; when the probability is lower than the threshold, it enters the continuous analysis process of the non-baseline response level. As can be seen from the figure, a significant high-probability peak only appears in certain periods during the test period, indicating that the adjustment resource is not always in the baseline response state, but rather experiences phased deviations and state switching processes. This result verifies the necessity of the dual-headed hierarchical structure for partitioning and modeling the baseline and non-baseline states.
[0313] at last, Figure 7The comparison results of the observed response level index and the analyzed response level index over the overall testing period are presented. As shown in the figure, the analyzed response level curve generally follows the changing trend of the observed response level, demonstrating a good ability to characterize major peak and valley changes and periodic fluctuations. Near the baseline state, the model output is relatively stable, and it can also reflect the rise and fall of the response level in non-baseline fluctuation ranges. These results indicate that this invention, through multi-source heterogeneous data processing, spatiotemporal fusion feature construction, SHAP key feature screening, LPCMCI time-delay causality identification, and dual-head hierarchical response analysis, can effectively support the response status analysis, capacity assessment, and trend prediction of virtual power plant regulation resources. Furthermore, Figure 8 This invention demonstrates the five-level structure of the unified data infrastructure built by this invention. Through this structure, data is processed in a complete closed loop, from initial acquisition to feature construction, filtering, state identification, and response analysis. This ensures the availability, stability, and interpretability of multi-source heterogeneous data in the analysis of virtual power plant response levels, providing a solid data foundation for subsequent regulation resource state analysis, capacity assessment, and trend prediction.
Claims
1. A method for multi-source data processing and causal correlation identification for virtual power plant response level analysis, characterized in that, Includes the following steps: S1: Construct a unified access and semantic regularization model for multi-source heterogeneous related data; S2: Construct a data missing repair and normalization model based on low-rank structure constraints; S3: Construct a spatiotemporal multidimensional fusion and dynamic mapping model for regulating resource response behavior; S4: Construct a causal feature generation model based on contribution screening and time-delay causal identification; S5: Construct an analysis model for regulating resource response levels based on a dual-headed hierarchical structure, and form a unified data foundation; The S1 construction of a unified access and semantic regularization model for multi-source heterogeneous related data includes: S1.1 Construct a multi-source heterogeneous data access set; S1.2 Unified analysis of time granularity mapping; S1.3 Construct a unified semantic tag model; S1.1 Construct a multi-source heterogeneous data access set; Let the set of data sources accessed during the virtual power plant operation analysis be . ; In the formula: This represents system boundary data, including information on predicted load demand, renewable energy output, reserve capacity, and conventional power output arrangements. It represents environmental meteorological data, including information on temperature, humidity, dew point, precipitation, wind speed, air pressure, and cloud cover; This indicates the operational data on the resource side, including information on charging equipment status, real-time power, and state of charge. It represents time and calendar data, including information such as time of day, day of the week, month, holidays, and weekday markers; For the m-th type of data source, its original observation sequence is represented as: ; In the formula: This represents the j-th variable in the m-th data source at the original time. Observed values; This indicates the number of variables contained in the m-th data source; This represents the set of original time indexes corresponding to the data source; S1.2 Unified analysis of time granularity mapping; To address the issue of inconsistent sampling frequencies across different data sources, all data is uniformly mapped to a fixed analysis granularity. Fixed analytical particle size To create a unified time index set for 15-minute intervals: ; For any original variable Through time aggregation operator Complete unified granularity mapping: ; In the formula: This represents the mapping value of the j-th variable in the m-th data source at a uniform time t; This represents the time index of the original data; Select the mean, maximum, minimum, last, cumulative, or state-held value based on the data type. N is 96 because there are 24 hours in a day, and the granularity of analysis is fixed. If the time is 15 minutes, then there are 96 points in a day; S1.3 Construct a unified semantic tag model; To avoid inconsistencies in the meaning of fields with the same name or inconsistencies in the names of synonymous fields across different systems, each data variable is defined as a semantic quintuple: ; In the formula: Indicates the data source attribute; Indicates the time attribute; Indicates spatial attributes, Indicates dimensional units; The variables represent their physical semantics, where m represents the m-th type of data source and j represents the j-th data variable. Through semantic mapping function This maps variables from different sources to a unified standard variable space. ; In the formula: This represents the system boundary data after unified time mapping; This represents environmental meteorological data after unified time mapping; This represents the resource-side runtime data after unified time mapping; This represents time calendar data after unified time mapping; t represents a time period; This forms an initial unified sample matrix for analyzing the response level of regulation resources in virtual power plants: ; In the formula: This represents the time period at the Nth time scale.
2. The multi-source data processing and causal correlation identification method for virtual power plant response level analysis according to claim 1, characterized in that, The S2 model for data missing data repair and normalization based on low-rank structure constraints includes: S2.1 Construct an observation mask for the missing locations; S2.2 Missing data repair based on low-rank matrix factorization; S2.3 Physical constraint projection and anomaly correction; S2.4 Gaussian distribution normalization processing.
3. The multi-source data processing and causal correlation identification method for virtual power plant response level analysis according to claim 2, characterized in that, S2.1 Construct an observation mask for the missing locations; Let a uniform sample matrix be defined. Where N represents the number of sample time points and P represents the number of variables; define the observation mask matrix. : ; Then the effective observation set Represented as: ; In the formula: subscripts i and j represent the matrix Positioning of elements within a text file; S2.2 Missing data repair based on low-rank matrix factorization; Correlated variables include system boundary variables, environmental meteorological variables, resource-side operational variables, and continuous composite features constructed from them; for time calendar markers, categorical state variables, and equipment number variables, rule imputation, mode imputation, or independent coding methods are used, and they do not directly participate in the low-rank repair of continuous variables; considering that multi-source correlated variables are constrained by common operational states, their sample matrix is decomposed into a low-rank principal structure matrix, a sparse anomaly matrix, and a random perturbation matrix: ; In the formula: This represents a low-rank principal structure matrix that reflects the main operating principles of the system. This represents the sparse anomaly perturbation matrix caused by local anomalies, mutation points, or erroneous acquisition. Represents a random disturbance term; The low-rank repair model is represented as: ; In the formula: , and It is a low-rank decomposition factor; This indicates the observation position projection operator, which calculates the reconstruction error only at valid observation positions; and These are the low-rank regularization coefficient and the sparse anomaly constraint coefficient, respectively. This represents the sum of squares of all elements in the matrix; S2.3 Physical constraint projection and anomaly correction; For variables with a defined physical range, further apply a projection onto the physically feasible region: ; In the formula: and Let these represent the lower and upper physical limits of the j-th variable, respectively. This process is used to prevent the missing data repair results from deviating from the practically feasible range. Represents the feature variables after missing data repair in low-order matrix factorization; S2.4 Gaussian distribution normalization; To eliminate differences in the dimensions and orders of magnitude of different variables, Gaussian distribution normalization is used: ; In the formula: and Let represent the mean and standard deviation of the j-th variable in the sample set, respectively; To prevent extremely small positive numbers with a denominator of zero; These are the feature variables projected from the physical constraints. These are the standardized feature variables; After the above processing, the normalized sample matrix is obtained: ; Where N represents the number of sample times and P represents the number of feature variables.
4. The multi-source data processing and causal correlation identification method for virtual power plant response level analysis according to claim 3, characterized in that, Step S2.2 also includes: To further improve the stability of repairing multi-source heterogeneous variables, the sample matrix is divided into several sub-matrices based on the physical properties of the variables and their data sources: ; In the formula: Represents the submatrix of system boundary variables; Represents a submatrix of environmental meteorological variables; This represents a submatrix of runtime variables on the resource side; This represents a submatrix of composite running state variables. For different submatrices, low-rank repair is performed separately, and then they are recombined using a unified time index to form a complete sample matrix. This block repair method can avoid the interference of weakly correlated variables on the low-rank structure and improve the stability and physical rationality of the missing data repair results. By solving the above model, the repaired complete sample matrix is obtained. : ; In the formula: and and represent the low-rank decomposition factors obtained from the optimization solution.
5. The multi-source data processing and causal correlation identification method for virtual power plant response level analysis according to claim 1, characterized in that, The S3 model for constructing a spatiotemporal multidimensional fusion and dynamic mapping model for regulating resource response behavior includes: S3.1 Spatial aggregation of regional environmental data; S3.2 Time-period variable cyclic encoding; S3.3 Construct composite operating state characteristics; S3.4 Form a spatiotemporal fusion feature set.
6. The multi-source data processing and causal correlation identification method for virtual power plant response level analysis according to claim 5, characterized in that, S3.1 Spatial aggregation of regional environmental data; For environmental variables collected from multiple spatial nodes, a statistical aggregation method is used to form regional environmental characteristics. Let the k-th type of environment variable of the r-th spatial node at time t be... The regional environmental characteristics Represented as: ; In the formula: R represents the number of spatial nodes; this method can transform scattered environmental information into a unified regional state description. This indicates the calculation of the average value; This indicates the calculation to obtain the maximum value; This indicates the calculation of taking the minimum value; This indicates the calculation of the standard deviation; S3.2 Time-period variable cyclic encoding; To avoid periodic breaks caused by directly using time numbers, a cyclic encoding method is used to represent time period variables: ; ; In the formula: This represents the sequence number of the current moment within a complete cycle; when the analysis granularity is 15 minutes, N=96; t represents the time period; S3.3 Construct composite operating state characteristics; To describe the coupling relationship between the system boundary state and the response behavior of regulating resources, a net demand intensity is constructed. New energy penetration rate , reserve capacity and climbing intensity Composite characteristics: ; ; ; ; In the formula: Indicates the load demand of the external system; and These represent the predicted available wind and solar power output, respectively. Indicates the positive reserve capacity; It is a very small positive number; t represents the time period; S3.4 Forms a spatiotemporal fusion feature set; Finally, by integrating standardized data, regional environmental characteristics, time cycle characteristics, and composite operational status characteristics, a basic fusion feature vector is formed. : ; In the formula: P represents the number of characteristic variables; This represents the value of the Pth variable in time period t.
7. The multi-source data processing and causal correlation identification method for virtual power plant response level analysis according to claim 1, characterized in that, The S4 construction of the causal enhancement feature generation model based on contribution screening and time-delay causal identification includes: S4.1 Effective Feature Filtering Based on SHAP Contribution; S4.2 Construct a set of candidate time-delay causal edges; S4.3 Conditional independence test based on LPCMCI.
8. The multi-source data processing and causal correlation identification method for virtual power plant response level analysis according to claim 7, characterized in that, S4.1 Effective Feature Filtering Based on SHAP Contribution; The SHAP method is used to calculate the marginal contribution of each candidate feature to the target response level output. Let the SHAP contribution value of the j-th feature in the i-th sample be... Then the global contribution of the j-th feature Defined as: ; In the formula: This represents the global contribution strength of the j-th feature to the target response level; Indicates the number of training samples All candidate features according to Sort the features from largest to smallest to obtain a contribution ranking set, and then select the top K key features based on their contribution rates to obtain the effective feature set. : ; In the formula: This represents the feature that ranks the Kth position in the contribution ranking; This process is used to filter out low-contribution and weak explanatory variables from high-dimensional fusion features, retaining candidate driving variables that contribute significantly to the target response state, thereby reducing dimensionality and noise interference for subsequent causal link identification. S4.2 Construct a set of candidate time-delay causal edges; Considering that the effects of different external variables on the adjustment of resource response levels do not all occur at the same time, but may have varying degrees of time delay, in the effective feature set Construct a set of multi-lag candidate variables based on : ; In the formula: Indicates the lag order; This represents the maximum lag order, determined based on the analysis granularity and resource response time constant; when the analysis granularity is 15 minutes... This indicates a delay of 1 hour; Indicates candidate valid feature variables; any candidate variable Record Construct its target response level Candidate time-delay causal edges: ; The candidate edge indicates that the j-th variable is in lag. During a given time period, there is a possibility that the target response level at the current moment may be influenced by that time period. S4.3 Conditional independence test based on LPCMCI; Within the LPCMCI time-delay causal search framework, the Parcorr partial correlation conditional independence test method is introduced to screen candidate time-delay causal edges; First, the LPCMCI algorithm constructs a set of condition variables based on historical dependencies and candidate parent node search results: ; In the formula: This represents the set of conditional variables used to eliminate confounding effects, including the historical terms of the target variable itself, lagged terms and time-period terms of candidate driving variables, and associated variables that have a common driving relationship with the candidate variables; The number of candidate condition variables; Using the set of conditional control variables respectively For candidate driving variables and target response variable By performing regression fitting, we can obtain its conditional predicted values: ; ; In the formula: and represents the fitted functions obtained by conditional regression on the candidate driving variable and the target response variable, respectively; t represents the time period; and These are the predicted values after regression fitting.
9. The multi-source data processing and causal correlation identification method for virtual power plant response level analysis according to claim 8, characterized in that, Step S4.3 also includes: Define the residual variable after removing the influence of the condition variable as follows: ; ; In the formula: and The residual variable after removing the influence of condition variables; The actual value of the candidate variable; The actual value of the target response variable; and These are the predicted values after regression fitting; Based on this, the partial correlation coefficient between the two residual sequences is calculated: ; In the formula: Indicates the set of control condition variables After that, candidate driving variables With the target response variable The net correlation strength that remains between them; and These represent the sample means of the corresponding residual sequences; when When the value is close to 0, it indicates that the linear residual correlation between the candidate driving variables and the target response variable is weak after controlling the set of condition variables; when... A larger value indicates a stronger net linear relationship between the two variables even after controlling for condition variables; according to The significance level can be calculated. If the following conditions are met: ; In the formula: Indicates the confidence level; Then it is considered that in the set of control condition variables After that, candidate driving variables With the target response variable Significant conditional dependencies still exist between them, and candidate time-delay causal edges are preserved. .
10. The multi-source data processing and causal correlation identification method for virtual power plant response level analysis according to claim 7, characterized in that, Step S4 also includes: S4.4 Causal Influence Estimation and Feature Selection Based on Optimal Adjustment Set; Furthermore, an optimal adjustment set mechanism is introduced to estimate the causal effect strength of candidate driving variables at corresponding lag orders. The purpose of this step is to avoid relying solely on the residual correlation strength to judge the magnitude of the variable's effect, but rather to quantitatively characterize the net effect of changes in candidate variables on changes in the target response level while minimizing the influence of confounding variables. Through this process, candidate edges that are statistically significant but have a weak actual effect are eliminated, thereby improving the effectiveness and interpretability of the input features of the subsequent tree model. First, to reduce the interference of common driving factors, the historical state of the target variable itself, and other related variables on the estimation of causal effects, an optimal set of adjustment variables is constructed for each retained time-delay causal edge: ; In the formula: Indicates the causal edge The selected optimal adjustment set; This represents the m-th adjustment variable; The optimal adjustment set can be composed of the candidate parent node set obtained by LPCMCI search, the historical term of the target variable, the significantly lagged driving variable, and the state variable that has a common influence relationship with the candidate variable. Then, a random forest nonlinear estimator is used to fit the causal effect function based on the optimal adjustment set. : ; In the formula: Indicates the candidate edge A random forest causal effect estimator is constructed; x represents the input driving variable value; t represents the time period; the random forest causal effect estimator fits the total effect relationship of candidate driving variables on the target response level under the constraint of the optimal adjustment set, and is used to describe the net effect of changes in candidate variables on the target response level after removing confounding factors; To facilitate comparison of the strength of causal effects among different variables, two symmetrical intervention levels are set in the standardized variable space. and Let represent the states where the candidate driving variable is above or below the mean by one standard unit, respectively. Based on the fitted total effect function, the estimated target response values at the two intervention levels can be calculated separately: ; ; In the formula: This indicates that the estimated target response value is higher than one standard unit of intervention level for the candidate driving variable; This indicates that the estimated target response value is lower than one standard unit of intervention level for the candidate driving variable; Therefore, the causal effect strength of candidate time-delay causal edges is defined. for: ; when This indicates that the candidate variable has a positive driving effect on the target response level at the corresponding lag order; when This indicates that the candidate variable has a negative inhibitory effect on the target response level at the corresponding lag order; The larger the value, the stronger the net causal effect of the candidate variable on the target response level at the corresponding lag order.
11. The multi-source data processing and causal correlation identification method for virtual power plant response level analysis according to claim 7, characterized in that, Step S4 also includes: S4.5 Construct a set of causal enhanced features; To further eliminate candidate edges with weak actual effects, a threshold for causal effect strength is set. Only when the causal strength of the candidate edge is greater than Only when the candidate time-delay causal edge is selected will it be retained, and its corresponding lagged variable will be used as an effective causal lag feature. The final set of causal lag features is formed after passing the LPCMCI conditional independence test, the optimal adjustment set identifiability test, and the causal effect strength screening. for: ; In the formula: Indicates the first q The first effective causal lag feature, namely the first... Candidate variables in lag The corresponding historical values It has a significant causal effect on the change of the target variable at the current time t.
12. The multi-source data processing and causal correlation identification method for virtual power plant response level analysis according to claim 1, characterized in that, The S5 constructs an analysis model for regulating resource response levels based on a dual-headed hierarchical structure, and forms a unified data foundation, including: S5.1 Constructing unified core input features; S5.2 Construct a baseline response state identification head; S5.3 Construct a continuous analysis head for non-benchmark response levels; S5.4 Constructs a probabilistic fusion and hard-switching output mechanism; S5.5 forms a unified data foundation.
13. The multi-source data processing and causal correlation identification method for virtual power plant response level analysis according to claim 12, characterized in that, S5.1 Constructing unified core input features; The spatiotemporal fusion features obtained in S3, the effective filtering features obtained in S4, and the causal enhancement features are combined to form the core input vector: ; In the formula: This is the feature set after filtering effective features based on SHAP contribution. This is the set of causal lag features after filtering. S5.2 Construct a baseline response state identification head; First, a baseline response state recognition head is constructed to determine whether the current sample is near the baseline response capability state; set up The target response level is The baseline response level is The threshold of the baseline state neighborhood is Then the baseline response status label is defined as: ; In the formula: This indicates that the current sample is near the baseline response state; This indicates that the current sample is in the non-benchmark response range; The baseline response state recognition head is represented by a tree-structured classification model as follows: ; In the formula: This represents the baseline response state identification model; The feature input vector representing time period t; This indicates the probability that the current sample belongs to the baseline response state; This probability is used to characterize the degree to which the current state of the adjustment resource approaches the baseline response capability state, and serves as the state gating signal for the subsequent hierarchical fusion output; S5.3 Construct a continuous analysis head for non-benchmark response levels; For samples that deviate from the baseline response state, a continuous analysis head for non-baseline response levels is further constructed to characterize the continuous variation of the regulatory resource response level within the non-baseline range. Let the non-benchmark sample set be... for: ; In the formula: Indicates the baseline response status label; The continuous analysis model for non-benchmark response levels is then expressed as: ; In the formula: This represents a continuous analysis model for non-benchmark response levels. The feature input vector representing time period t; This represents the estimated response level under non-baseline conditions; S5.4 Constructs a probabilistic fusion and hard-switching output mechanism; After completing the training of the two analysis heads, the baseline response state probabilities will be... Compared with non-benchmark response level estimates The results of the fusion were obtained to obtain the final response level analysis results: ; In the formula: As the baseline response level; Indicates the hard handover threshold of the baseline response state. This is the final response level analysis result; this mechanism can reduce unreasonable offsets near the baseline state, improve the model's ability to identify the baseline state concentration range and the stability of the output; When the baseline response state probability Not less than the threshold When the system directly outputs the baseline response level; when Less than the threshold At that time, the system outputs the continuous response level after probability fusion; S5.5 forms a unified data foundation; Finally, the multi-source input data, spatiotemporal fusion features, LPCMCI time-delay causal features, SHAP contribution screening results, dual-head hierarchical model output results and evaluation results are all written into the data foundation base. The data infrastructure includes the following levels: First, the raw associated data layer is used to store system boundary data, environmental data, resource-side status data, and time calendar data; Second, the feature construction layer is used to store time period features, historical state features, artificial composite state features, and LPCMCI time-delay causal features. Third, the feature selection layer is used to store the SHAP contribution ranking results, the selected core feature set, and feature explanation information; Fourth, the state recognition layer, used to store the baseline response state probabilities. State switching threshold and baseline state identification results; Fifth, the response analysis layer, used to ultimately predict the response level output. and actual observed response level .
Citation Information
Patent Citations
Optimal regulation and control method and system based on multi-source heterogeneous virtual load
CN111509728A
Digitization system and method for energy internet deployment management
CN113077101A