Distributed station daily clearing data checking method and system based on multi-dimensional feature fusion

By using a quantum genetically optimized random forest regression algorithm and a multi-level data verification and labeling system, the randomness and anomaly identification problems of daily data clearing at distributed stations were solved, achieving data accuracy and consistency, and ensuring data stability and reliability.

CN121412908APending Publication Date: 2026-01-27STATE GRID SHANDONG ELECTRIC POWER CO MARKETING SERVICE CENT (MEASURING CENT)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511514916.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

The power generation values ​​in the daily data of distributed power stations are highly random and greatly affected by weather. Traditional interpolation methods are difficult to adapt to, resulting in poor data verification accuracy and consistency, and difficulty in identifying abnormal patterns.

Method used

We employ a random forest regression algorithm based on quantum genetic optimization to fill in missing values, and combine EMD smoothing and Z-score standardization to process the data, constructing a multi-level data verification and labeling system. We then identify abnormal patterns through multi-source information fusion and cluster analysis.

Benefits of technology

It improves the accuracy, completeness and consistency of daily data from distributed stations, enables high-resolution anomaly identification and verification, and ensures data stability and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121412908A_ABST
    Figure CN121412908A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of distributed station daily clearing data checking, provides a distributed station daily clearing data checking method and system based on multi-dimensional feature fusion, establishes a multi-level data processing system oriented to a distributed station, and provides a random forest regression algorithm based on quantum genetic optimization to solve the problem of data missing. By analyzing the influence of multi-dimensional factors on station data quality, a station resource configuration-oriented data checking tag system, a multi-source information fusion-based checking tag system and a clustering analysis-based data exception mode tag system are constructed, and finally, a multi-dimensional and comprehensive daily clearing data checking system is formed. And the accuracy, integrity and consistency of the data are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of distributed facility daily clearing data verification technology, and particularly relates to a distributed facility daily clearing data verification method and system based on multi-dimensional feature fusion. Background Technology

[0002] Daily data verification, through a comprehensive review of key data such as power generation and equipment status, can promptly identify data anomalies. In particular, for distributed photovoltaic power generation, verification through power indicator curves ensures that the collected data is complete and conforms to the theoretical data range, preventing data anomalies from affecting grid dispatch. Regular verification can effectively identify potential equipment problems, such as improper overvoltage protection settings or abnormal anti-islanding functions, thereby preventing grid accidents. Daily data verification provides a foundation for the long-term supervision of distributed renewable energy.

[0003] However, when verifying daily data from distributed power stations, the power generation values ​​of the stations are random and greatly affected by weather conditions. Sudden changes are common, making it difficult for general interpolation methods to adapt. They also cannot obtain category labels that reflect the spatiotemporal differences of abnormal station patterns, thus affecting the verification resolution of anomaly identification. The accuracy, completeness, and consistency of data are poor in the traditional verification process. Summary of the Invention

[0004] To address the aforementioned issues, this invention proposes a distributed daily data verification method and system for power stations based on multi-dimensional feature fusion. This invention establishes a multi-level data processing system for distributed power stations and proposes a random forest regression algorithm based on quantum genetic optimization to solve the data missing problem. By analyzing the impact of multi-dimensional factors on power station data quality, it constructs a data verification label system for power station resource allocation, a verification label system based on multi-source information fusion, and a data anomaly pattern label system based on cluster analysis. Ultimately, a multi-dimensional and comprehensive daily data verification system is formed, ensuring the accuracy, completeness, and consistency of the data.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solution: In a first aspect, the present invention provides a distributed field station daily data verification method based on multi-dimensional feature fusion, comprising: Obtain daily clearing data from distributed data stations; The daily power generation data of distributed power stations is cleaned, features are extracted, and standardized. In the process of cleaning the power generation data, missing values ​​are filled in by the random forest algorithm, and the quantum genetic algorithm is used to iteratively optimize the number of generated trees and the maximum depth of the trees in the random forest algorithm. Based on the standardized distributed daily clearing data of the stations, a multi-level data verification label system is constructed from two dimensions: resource allocation characteristics and multi-source information characteristics. When constructing the multi-level data verification label system from the multi-source information characteristics dimension, variance analysis is used to screen feature factors that have a significant impact on data quality. A local distortion penalty factor is introduced on the basis of dynamic time warping path to combine the global alignment capability of dynamic time warping with the sensitivity of local Euclidean distance. A path-aware weighting mechanism is used to balance the flexibility of time warping with the need to maintain the shape. The daily clearing data curves of the stations are classified by time series clustering to identify normal and abnormal patterns. Seasonal differences and weather condition differences are also considered. Hierarchical clustering is used to refine the data anomaly types and obtain separate category labels for data anomalies. The verification module is configured to perform distributed daily data verification of the stations based on the constructed multi-level data verification label system.

[0006] Secondly, the present invention also provides a distributed daily data verification system for airfields based on multi-dimensional feature fusion, comprising: The data acquisition module is configured to acquire daily clearing data from distributed data stations. The data processing module is configured to perform power generation data cleaning, feature extraction, and standardization on the daily data of distributed power stations. Specifically, when cleaning the power generation data, missing values ​​are filled in using the random forest algorithm, and the quantum genetic algorithm is used to iteratively optimize the number of generated trees and the maximum depth of the trees in the random forest algorithm. The tag creation module is configured to: construct a multi-level data verification tag system based on the standardized distributed daily clearing data of the stations from two dimensions: resource configuration characteristics and multi-source information characteristics; when constructing the multi-level data verification tag system from the multi-source information characteristics dimension, the module uses variance analysis to screen feature factors that have a significant impact on data quality, introduces a local distortion penalty factor on the basis of dynamic time warping path, combines the global alignment capability of dynamic time warping with the sensitivity of local Euclidean distance, balances the flexibility of time warping with the need to maintain the shape through a path-aware weighting mechanism, classifies the daily clearing data curves of the stations through time series clustering, identifies normal and abnormal patterns, and considers seasonal and weather condition differences, refines the data anomaly types through hierarchical clustering, and obtains separate category tags for data anomalies; The verification module is configured to perform distributed daily data verification of the stations based on the constructed multi-level data verification label system.

[0007] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the distributed field station daily clearing data verification method based on multi-dimensional feature fusion described in the first aspect.

[0008] Fourthly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the program to implement the steps of the distributed field station daily clearing data verification method based on multi-dimensional feature fusion described in the first aspect.

[0009] Fifthly, the present invention also provides a computer program product, the computer program product comprising a computer program, which, when executed by a processor, implements the steps of the distributed field station daily clearing data verification method based on multi-dimensional feature fusion described in the first aspect.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention establishes a multi-level data processing system for distributed data stations and proposes a random forest regression algorithm based on quantum genetic optimization to solve the problem of missing data. By analyzing the impact of multi-dimensional factors on the data quality of data stations, it constructs a data verification labeling system for data station resource allocation, a verification labeling system based on multi-source information fusion, and a data anomaly pattern labeling system based on cluster analysis. Finally, a multi-dimensional and comprehensive daily clearing data verification system is formed, ensuring the accuracy, integrity, and consistency of the data.

[0011] 2. This invention addresses the issues of integrity and stability in daily data processing at distributed data stations by proposing a combined preprocessing method of "quantum forest gap filling - EMD smoothing - Z-score standardization" to improve the reliability of data analysis. Based on this, an improved DTW distance and K-shape clustering method incorporating path-aware weighting are used to achieve preliminary classification of abnormal patterns in daily data processing, effectively identifying curve shapes and temporal distribution differences. Subsequently, coupled feature vectors are constructed by fusing seasonal information and weather conditions, and hierarchical clustering is used to deeply subdivide abnormal data patterns, obtaining refined category labels that reflect the spatiotemporal differences of abnormal patterns at data stations, providing high-resolution anomaly identification criteria for subsequent verification.

[0012] 3. This invention constructs a multi-level data verification and labeling system from two dimensions: resource allocation characteristics and multi-source information characteristics. Regarding resource allocation labels, based on the combination of various energy forms within the power station, such as photovoltaic, wind power, and energy storage, multiple types of classification labels are established, and key operational indicators for each type of resource are extracted. Regarding multi-source information labels, multi-source data, including data from the power station monitoring system, meteorological observation data, GIS regional attributes, and operation and maintenance records, are integrated to form a comprehensive labeling system covering basic attributes, power generation behavior, energy allocation, and interactive responses. Furthermore, the ANOVA method is used to screen feature factors that significantly impact data quality, providing a multi-dimensional reference benchmark for anomaly detection and data correction.

[0013] 4. Based on the established normal data pattern library and refined anomaly pattern classification results, this invention conducts a full-process verification of daily data from distributed data stations. Newly collected data is matched with the corresponding station's normal patterns for deviation detection, quickly identifying and locating anomaly types. For missing data, a random forest regression algorithm based on improved quantum genetic optimization is used for completion. For abnormal fluctuations, data offsets, and equipment failure-related anomalies, targeted corrections are made based on station resource characteristics, real-time weather, and historical operating curves. During the verification process, the anomaly pattern library and feature labels are dynamically updated to continuously optimize the discrimination threshold and repair strategy, thereby efficiently ensuring the accuracy, integrity, and consistency of data in fully automated operation. Attached Figure Description

[0014] The accompanying drawings, which form part of this embodiment, are used to provide a further understanding of this embodiment. The illustrative embodiments and their descriptions are used to explain this embodiment and do not constitute an improper limitation of this embodiment.

[0015] Figure 1 This is a flowchart of the missing value imputation method based on the improved quantum genetics and random forest regression algorithm in Embodiment 1 of the present invention; Figure 2 This is a flowchart illustrating the construction process of the distributed daily clearing data reference base value for embodiment 1 of the present invention. Detailed Implementation

[0016] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0017] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0018] Example 1: This embodiment provides a distributed facility daily clearing data verification method based on multi-dimensional feature fusion. Starting from the cleaning and standardization process of facility daily clearing data, it proposes a multi-level distributed facility data verification method by analyzing the multi-dimensional influencing factors of facility data quality. It systematically constructs a complete verification label system for distributed facility daily clearing data from three aspects: resource allocation of distributed facilities, multi-source information fusion, and clustering of abnormal patterns in daily clearing data. The method includes: S1. Distributed site daily clearing data processing method: Original, reliable, and high-quality distributed daily data from farms is a prerequisite for accurate verification of target data. This section fills in the missing parts of the original distributed daily data from farms by integrating the Improved Quantum-Inspired Genetic Algorithm (IQGA) and Random Forest Regression methods. It also combines Empirical Mode Decomposition (EMD) smoothing and Z-score normalization methods to improve the integrity of the original data and achieve preprocessing of the distributed daily data from farms.

[0019] S1.1 Distributed power generation data cleanup: In distributed power plant daily data, the power generation values ​​of the power plants exhibit randomness, are greatly affected by weather conditions, and are prone to sudden changes, making conventional interpolation methods difficult to adapt. However, daily data also exhibits a certain periodicity, with the power generation behavior of the power plants often showing a periodic pattern on a daily basis. To meet the need for rapidly processing massive amounts of daily time-series data and filling in missing values, this embodiment employs an improved quantum genetic optimization-based random forest regression algorithm to complete the missing values ​​in the daily data.

[0020] Random forest regression (RFR) is an ensemble learning method that uses regression decision trees as base estimators. This algorithm has good tolerance for noise and outliers, and the random forest averages the predictions of the base estimators to determine the result of the ensemble estimator, thus effectively avoiding overfitting and having more significant generalization performance and accuracy than a single regression decision tree.

[0021] Assume X and Y are the input and output, respectively, and Y is a continuous variable. The input set is: (1) Suppose a regression tree has M leaves, meaning the regression tree divides the input space X into M units. A regression tree can have at most M distinct output values. In the partitioning of the regression tree, the j-th feature is selected. and its values As the splitting variable and splitting point, the splitting variable and splitting point divide the input space of the parent node into two, which can be represented as: (2) Each time a splitting variable and splitting point are selected for partitioning, all splitting variables are traversed. For a fixed splitting variable q, the splitting point s is scanned, and the one that minimizes the following expression is selected. : (3) in, , The subsets to the left of the split point s are respectively and the right subset In this context, the prediction constant that minimizes the squared error is the predicted mean value of the two sample groups.

[0022] Use the selected Divide the region and determine the corresponding output: (4) Repeat the above steps for the two sub-regions until the termination condition is met, thus obtaining a regression tree.

[0023] The generation rules for each regression tree in a random forest are as follows: S1.1.1 Perform bootstrap sampling from the training data, randomly selecting N samples with replacement to form the training set for a regression tree.

[0024] S1.1.2 If a sample has M features, randomly select m features from them. When splitting the regression tree, select the best of these m features as the splitting feature.

[0025] S1.1.3 During the regression tree generation process, each node is split according to step S1.1.2 until the termination condition is met. At the same time, the regression tree is pruned to prevent overfitting.

[0026] S1.1.4. Following steps S1.1.1 to S1.1.3, establish a large number of decision trees to form a random forest. Take the average value of these regression tree results as the final result to improve prediction accuracy and control overfitting.

[0027] The specific process for imputing missing battery values ​​based on random forest regression is as follows: The distributed power generation time-series data is divided into 15-minute segments on a daily basis to obtain daily cleared time-series data with a length of 96.

[0028] The known variables in the time series are used as features, and the variables with missing values ​​are used as labels. The data with numerical values ​​in the label variables are used as the training set, and the missing values ​​are used as the test set.

[0029] Missing values ​​were filled in using a random forest algorithm. Finally, the data was restored to hourly electricity consumption time-series data of the power station with timestamps.

[0030] S1.2 Iterative optimization of the number of spanning trees and the maximum depth of the trees: Two important parameters that affect the performance of the random forest algorithm are the number of spanning trees ( n_ estimator ) and the maximum depth of the tree ( max_depthThe number of spanning trees represents the number of base estimators, i.e., the number of regression trees built. Increasing the number of subtrees improves model stability and prediction accuracy, but setting it too high increases model complexity and computation time. A higher maximum tree depth allows the model to learn more specific relationships between samples, but it also increases the probability of overfitting. To obtain the optimal parameter settings, an improved quantum genetic algorithm is used to iteratively optimize the number of spanning trees and the maximum tree depth to determine the optimal parameter values.

[0031] The Quantum Genetic Algorithm (QGA) combines quantum computing with genetic algorithms. Addressing the problems of premature convergence and slow convergence speed in traditional genetic algorithms, QGA introduces quantum encoding and quantum rotation gates, enabling a single chromosome to represent multiple states simultaneously. This prevents premature convergence and accelerates convergence. In quantum computing, the information carrier is no longer the classical bit, but the qubit. Compared to the two states of 0 or 1 in traditional genetic algorithms, qubit-based chromosomes can exist in a linear superposition of 0 and 1 eigenstates. Therefore, qubit-based chromosomes contain more information, maximizing population diversity. A qubit-based chromosome can be represented as: (5) Among them, complex numbers α and β This represents the probability of a corresponding state occurring. In a quantum genetic algorithm, a criterion has... m A chromosome with 1 qubit can be represented as: (6) in, Indicates the first t The generation i One chromosome; m Indicates the number of genes on a chromosome; n Indicates population size.

[0032] In genetic algorithms, mutation is typically a random change. However, in quantum theory, state transitions are achieved through quantum rotation gates. Therefore, genes can be updated through mutation using quantum rotation gates. A quantum rotation gate can be represented as: (7) in, Let be the rotation angle. The angle of rotation variation is usually obtained using a lookup table.

[0033] After n rotation gate transformations, a superposition of 2^n states can be generated, thus its efficiency is much higher than that of traditional genetic algorithms. However, the rotation angle obtained by the lookup table method cannot be dynamically adjusted during the update process. Therefore, an adaptive adjustment method is considered, which makes the rotation angle change with the magnitude of the difference between each individual and the current best individual. The adaptively adjusted rotation angle can be expressed as: (8) in, and This represents the minimum or maximum rotation angle. and These are the minimum and maximum fitness of an individual in the current population, respectively; f It is the fitness value of the current individual.

[0034] As can be seen from the above formula, if the current individual's fitness is low, quantum mutation uses a larger rotation angle to accelerate the mutation speed; if the current individual's fitness is high, quantum mutation uses a smaller rotation angle to converge to the optimal solution. For example, if the current individual's fitness is less than a preset value, quantum mutation uses a first preset rotation angle; if the current individual's fitness is greater than or equal to the preset value, quantum mutation uses a second preset rotation angle. The first preset rotation angle is greater than the second preset rotation angle. The first and second preset rotation angles are not fixed values, but are used to express the difference between larger and smaller rotation angles. In the quantum genetic algorithm adjustment strategy, all individuals in the population must update their chromosomes towards the direction of the current maximum fitness. However, if the direction of the current maximum fitness is not the global optimum but only a local optimum, it may cause the entire population to fall into a local optimum. To address this, the idea of ​​probability partitioning is introduced to initialize the probability amplitude of the population's qubits, which can effectively expand population diversity and avoid the population falling into a local optimum.

[0035] (9) in, N Indicates population size; i Indicates the first i Individual populations.

[0036] In summary, the specific process of the missing charge value imputation method based on the improved quantum genetics and random forest regression algorithm is as follows: S1.2.1 Filter the current and charge time series data to be filled based on the missing rate.

[0037] S1.2.2 The quantum genetic algorithm sets the population size to 200 and the quantum encoding bit width to 8. It initializes the probability amplitude for each individual and generates binary codes for the key parameters of the random forest regression.

[0038] S1.2.3. Perform random forest regression on the non-missing data. Define the fitness function based on the square root error obtained from the random forest regression model, calculate the fitness value of all individuals in the population. The larger the fitness value, the better the model prediction effect. Take the individual with the largest fitness as the evolution direction of the next generation.

[0039] S1.2.4 Determine whether the termination condition is met. If it is met, terminate the quantum genetic algorithm, obtain the optimal random forest regression parameters, and perform random forest regression on the current-electricity time series to be filled. Otherwise, calculate the quantum rotation gate rotation angle of each individual, update the probability amplitude of all individuals, and repeat step S1.2.2.

[0040] S1.3, Daily Clearing Data Trend Feature Extraction: To extract the temporal characteristics of the data fluctuation process, the Empirical Mode Decomposition (EMD) method is used to decompose the original time series. This method decomposes the series based on the time scale characteristics of the data itself, without the need to pre-set basis functions as in wavelet packet decomposition. Therefore, it is suitable for analyzing nonlinear and non-stationary series with high signal-to-noise ratios. Identifying the extreme points of the series is crucial for subsequent calculation of the envelope using cubic spline interpolation. The upper envelope is obtained by interpolating local maxima, while the lower envelope is obtained by interpolating local minima. It is particularly important that both the upper and lower envelopes must contain all data points. Subsequently, the average envelope needs to be subtracted from the original series to obtain the intermediate series. The specific expression is as follows: (10) in, This represents the initial load sequence of the input for the first decomposition. The first-order characteristic mode function is obtained by averaging the upper and lower envelopes of the first decomposition. This represents the intermediate load sequence after the first decomposition. Subsequent decompositions should be performed by repeating the above process.

[0041] (11) in, Indicates the first The load sequence of the next decomposition; Indicates the first First-order characteristic mode function; Indicates the first The intermediate load sequence after the sub-decomposition.

[0042] The decomposition process is repeated iteratively until the intermediate load sequence after decomposition is obtained. It becomes a monotonic function or approaches zero. At this point, the decomposition stops, and the residual sequence is determined as the final output. Therefore, the original sequence can be represented as a sum of multiple characteristic mode functions: (12) in, Indicates the initial load sequence; Indicates the first First-order characteristic mode function, This represents the residual of the load sequence.

[0043] S1.4 Standardization of daily clearing data feature values: Distributed power station daily data includes time series data such as power output and electricity generation / consumption. These data have different dimensions and significantly different absolute values. Directly analyzing the raw indicator values ​​would emphasize the role of indicators with higher values ​​in the overall analysis, relatively weakening the role of indicators with lower values. Therefore, to ensure the reliability of the results, the raw data can be standardized. Standardization methods include Min-Max standardization and z-score standard deviation standardization. Min-Max normalization of the input data does not change the distribution of the original time series data; it only relates to the maximum and minimum values. If the maximum and minimum values ​​are outliers, most of the data will be compressed into a very small range, thus affecting the effectiveness of subsequent feature extraction. Z-score standard deviation standardization has high robustness to outliers. Considering that the distributed power station daily data has already undergone outlier removal and missing value imputation, the Z-score normalization method is used to perform feature preprocessing of the power station data, as detailed below.

[0044] Z-score standardization standardizes the time series data for each station, mapping the data to a time series with a mean of 0 and a variance of 1. Its expression is: (10) in, Indicates the first position in each sequence i Individual amplitude data, For the first in the sequence i The average of the sequences, For the first in the sequence i The variance of each sequence.

[0045] S2. Distributed site daily data classification and processing method: Following the systematic standardization of daily data from distributed power stations, a data foundation capable of providing analytical capabilities has been established. This section further mines the characteristics and patterns of power station generation and consumption behavior through daily data mining. Starting with multi-dimensional factors, it explores the mechanisms by which each factor influences power station behavior and constructs a classification model based on clustering algorithms to scientifically categorize the power generation and consumption behavior of distributed power stations, thereby revealing the intrinsic relationship between data characteristics and behavioral patterns.

[0046] S2.1 Distributed site tagging system oriented towards equipment configuration: The distributed power station tagging system for resource allocation aims to systematically classify power stations with different energy resource combinations based on in-depth analysis of the types of energy resources and the operational characteristics of different resource types within the power station. Secondly, based on the resource characteristics configured in the power station, key power generation and consumption behavior indicators are extracted to form refined power station tags.

[0047] S2.1.1 Distributed site tags based on internal resource configuration: Distributed power plant energy systems exhibit diverse forms due to the integration of various power generation resources. To clearly understand the energy production modes and grid interaction characteristics of these systems, a scientific classification based on the types of resources actually possessed by the power plant is necessary. While plant power load is common, the configuration of distributed photovoltaic (PV), distributed wind power (WT), and distributed energy storage (DES) varies across power plants. These differences in resource configuration directly affect the power plant's power generation behavior, energy output capacity, and interaction with the grid. Classifying power plants according to their internal resource types allows for precise analysis of the energy characteristics of different power plants, providing a basis for optimizing grid dispatch and improving energy utilization efficiency.

[0048] In the specific classification process, the concept of Cartesian product from set theory is adopted, using distributed photovoltaic (existing / non-existent), distributed wind power (existing / non-existent), and distributed energy storage (existing / non-existent) as three variables, forming a combination space with the inevitable plant power load. In this way, all possible configuration forms of distributed power station energy resources are comprehensively covered, realizing a systematic and logical division of the power station group, and further studying the energy production laws of the power station under each resource combination and its impact on grid operation.

[0049] 1) Tagging method: A hierarchical labeling system is adopted to identify power plants, with "plant power load" as the core label, supplemented by the presence indicators of photovoltaic (P), wind power (W), and energy storage (S) as feature labels. The specific labeling rules are as follows: ① Establish binary coding rules for resource existence: photovoltaic power is recorded as P=1 if it exists, otherwise P=0; wind power is recorded as W=1 if it exists, otherwise W=0; energy storage is recorded as S=1 if it exists, otherwise S=0.

[0050] ② Adopt a four-segment label structure of "plant power - photovoltaic label - wind power label - energy storage label", for example, the label for plant power + photovoltaic + wind power + energy storage site is L-P1-W1-S1.

[0051] ③ To enhance the semantic readability of the labels, typical combinations are given academic names, such as defining L-P0-W0-S0 as "pure electricity-consuming power station" and L-P1-W1-S1 as "wind-solar-storage synergistic power station".

[0052] Based on the above principles and tagging methods, resources can be divided into four basic categories, and the resource configuration combinations and tags for each category are shown in Table 1: Table 1. Classification of Station Resources

[0053] 2) Analysis of each category: Pure plant power supply type substation (L-P0-W0-S0): These types of power stations only have auxiliary power loads, and their energy consumption pattern is characterized by unidirectional power draw from the grid. Their load curves are relatively stable, and they are primarily used to maintain the operation of station equipment and monitoring systems. Their impact on the power grid is mainly reflected in basic load demand; their electricity consumption is relatively small, but a continuous power supply is required.

[0054] Photovoltaic power station (L-P1-W0-S0): The power station simultaneously possesses both plant power load and distributed photovoltaic (PV) systems, exhibiting an energy characteristic of "self-consumption with surplus power fed into the grid." PV output shows significant daily periodicity and seasonal variation; during peak daytime power generation, it can meet the plant's power needs and supply a large amount of electricity to the grid; at night, it relies entirely on grid power. The key indicators for this type of power station are PV power generation efficiency and grid-connected power, whose values ​​are significantly affected by installed capacity and sunlight conditions.

[0055] Wind power type station (L-P0-W1-S0): With plant power load and distributed wind power as the core configuration, wind power output is random and intermittent, and the power generation curve is highly correlated with wind speed changes. While there is no obvious daily periodicity, seasonal characteristics may exist. The power supplied to the grid by this type of wind farm fluctuates significantly, while its own plant power consumption is relatively stable but accounts for a small proportion. Key performance indicators for wind power farms include turbine utilization hours and power generation efficiency.

[0056] Energy storage station (L-P0-W0-S1): Equipped with plant power loads and distributed energy storage systems, their operation modes can be divided into two categories: "peak-valley arbitrage" and "ancillary services." The former maximizes economic benefits by charging during off-peak hours and discharging during peak hours; the latter provides ancillary services such as frequency regulation and voltage regulation to the power grid. The energy storage capacity configuration and charging / discharging strategies of energy storage stations directly affect their economic efficiency and system value.

[0057] Wind and solar type station (L-P1-W1-S0): This type of power station integrates plant power load, distributed photovoltaic power, and decentralized wind power to achieve multi-energy complementary power generation. The wind-solar complementary characteristics make the power generation curve smoother compared to a single energy source: photovoltaic power generation is the main source during the day, while wind power supplements it at night and during cloudy or rainy weather. The power generation characteristics of this type of station exhibit "double randomness," but the overall reliability is improved compared to a single energy source.

[0058] Photovoltaic-storage type power station (L-P1-W0-S1): Equipped with plant power loads, distributed photovoltaic (PV) power, and energy storage systems, this system forms an integrated "generation-storage-use" energy unit. The energy storage system effectively mitigates the volatility of PV power generation, smoothing the power generation curve. Its typical operating mode is as follows: PV power generation prioritizes meeting the plant's power needs, with surplus electricity stored in energy storage or fed into the grid; when PV output is insufficient, the energy storage system discharges to supplement it. This type of power station can provide relatively stable power output, reducing grid regulation pressure, and can flexibly adjust charging and discharging strategies based on electricity price signals or dispatch instructions.

[0059] Wind storage type wind power station (L-P0-W1-S1): By combining plant power load, distributed wind power, and distributed energy storage, a localized solution for wind power absorption is formed. The energy storage system can effectively suppress random fluctuations in wind power, improving power quality and wind power absorption capacity. Its operational characteristics are: when wind power output fluctuates, the energy storage system responds quickly to smooth the power output; during periods of surplus wind power, the stored energy is charged, and during periods of insufficient wind power, it is discharged, achieving peak shaving and valley filling for wind power.

[0060] Wind-solar-storage power station (L-P1-W1-S1): Integrating plant power load, distributed photovoltaic power, decentralized wind power, and distributed energy storage, this system possesses the strongest energy self-regulation capability. Wind-solar hybrid power generation combined with energy storage regulation forms a multi-layered energy synergy mechanism: wind-solar hybridization reduces power generation fluctuations, while the energy storage system further smooths out residual fluctuations and optimizes energy allocation. The energy interaction of this type of power station exhibits multi-timescale characteristics: on a short timescale, energy storage smooths out wind-solar fluctuations; on a long timescale, energy optimization is achieved through joint wind-solar-storage scheduling.

[0061] S2.1.2 Distributed site tags based on resource characteristics: The resource-characteristic-based power plant labeling system extracts key characteristic indicators of power plant generation and consumption behavior to construct a set of physical characteristic labels reflecting the power generation mode, equipment operating status, and energy utilization efficiency of the power plant. Depending on the internal resource configuration of the power plant, different types of power plants may include labels such as plant power load characteristics, distributed photovoltaic operation characteristics, distributed wind power operation characteristics, and distributed energy storage operation characteristics, comprehensively characterizing the power plant's generation and consumption behavior.

[0062] 1) Plant power load characteristic label: The power load characteristic tags for plant use mainly reflect the power consumption behavior patterns and load change patterns of the plant. Through statistical analysis of historical power consumption data, the following (but not limited to) key indicators are extracted: ① Peak power label for plant power consumption (PEAK_LOAD_POWER / kW): This reflects the maximum power demand of the plant by statistically analyzing the average daily maximum power consumption of the plant over a longer period of time.

[0063] ②VALLEY_LOAD_POWER / kW: This reflects the minimum power demand of a plant by statistically analyzing the average daily minimum power consumption over a longer period.

[0064] ③ Average Daily Electricity Consumption (AVERAGE_DAILY_ELECTRICITY_CONSUMPTION / kWh): This indicator reflects the scale of electricity consumption of a power plant by statistically analyzing its average daily electricity consumption over a relatively long period. This metric is directly related to the plant's operation and maintenance costs and energy consumption level.

[0065] ④ Plant Power Consumption Ratio (AUXILIARY_POWER_RATIO / %): Defined as the ratio of plant power consumption to total power generation, reflecting the plant's energy self-sufficiency level and operational efficiency. The lower the ratio, the higher the plant's energy utilization efficiency.

[0066] 2) Distributed photovoltaic feature tags: For power plants equipped with distributed photovoltaic systems, it is necessary to extract operational characteristic indicators of the photovoltaic system and construct a photovoltaic characteristic labeling system: ① Photovoltaic grid connection mode label (PV_GRID_CONNECTION_TYPE): includes two modes: "self-consumption with surplus power fed into the grid" (M1) and "full power fed into the grid" (M2).

[0067] ② Photovoltaic installed capacity label (PV_INSTALLED_CAPACITY / kW): Reflects the power generation capacity of the photovoltaic system at the site.

[0068] ③ Photovoltaic daily average power generation label (PV_AVERAGE_DAILY_ELECTRICITY / kWh): Reflects the actual power generation efficiency and operating status of the photovoltaic system, which is affected by many factors such as irradiance conditions, equipment performance, and maintenance status.

[0069] ④ Peak Value Label for Photovoltaic Power Generation (PEAK_PHOTOVOLTAIC_POWER / kW): This label reflects the maximum daily power generation level of the photovoltaic devices installed at the power station by statistically analyzing the average daily peak value of photovoltaic power generation over a longer period of time.

[0070] ⑤ Photovoltaic Utilization Hours Label (PV_UTILIZATION_HOURS / h): Defined as the ratio of annual power generation to installed capacity, reflecting the annual utilization efficiency of the photovoltaic system.

[0071] 3) Distributed wind power operation characteristic tags: For wind farms equipped with distributed wind power, it is necessary to extract operational characteristic indicators of the wind power system and construct a wind power characteristic labeling system: ① Wind power grid connection mode label (WT_GRID_CONNECTION_TYPE): Includes two modes: "self-generation and self-consumption with surplus power fed into the grid" (M1) and "full power fed into the grid" (M2).

[0072] ② Wind power installed capacity label (WT_INSTALLED_CAPACITY / kW): Reflects the power generation capacity of the wind power system at the site.

[0073] ③ Wind power daily average power generation label (WT_AVERAGE_DAILY_ELECTRICITY / kWh): Reflects the actual power generation efficiency and operating status of the wind power system, which is affected by many factors such as wind conditions, equipment performance, and maintenance status.

[0074] ④ Peak Wind Power Generation Average Value Label (PEAK_WIND_POWER / kW): This label reflects the maximum daily power generation level of the wind power equipment installed at the wind farm by statistically analyzing the average daily peak wind power generation of the wind farm over a longer period of time.

[0075] ⑤ Wind power utilization hours label (WT_UTILIZATION_HOURS / h): Defined as the ratio of annual power generation to installed capacity, reflecting the annual utilization efficiency of the wind power system.

[0076] ⑥ Wind Power Volatility Rate (WT_FLUCTUATION_RATE / %): Defined as the ratio of the standard deviation to the average value of daily wind power output, reflecting the stability of wind power output. The larger this value, the greater the fluctuation in wind power output, and the higher the requirements for grid regulation.

[0077] 4) Distributed energy storage operation characteristic tags: For power stations equipped with energy storage systems, it is necessary to extract the operational characteristic indicators of the energy storage systems and construct an energy storage characteristic labeling system: ① Energy storage capacity label (ENERGY_STORAGE_CAPACITY / kWh): Reflects the energy storage capacity of the site's energy storage system.

[0078] ②Charge and Discharge Power Label (CHARGE_AND_DISCHARGE_POWER / kW): The maximum charging and discharging power of the energy storage system, reflecting the power regulation capability of the energy storage system.

[0079] ③ Charge and Discharge Time Preference Labels (CHARGE_AND_DISCHARGE_TIME_PREFERENCE): By analyzing the charge and discharge time distribution of the energy storage system, the energy storage usage mode of the site is identified. Systems that can be set to charge during off-peak hours and discharge during peak hours are classified as "Peak-Valley Arbitrage" (T1); systems that charge during peak hours and discharge during off-peak hours are classified as "Renewable Energy Consumption" (T2); systems primarily used to smooth power fluctuations are classified as "Fluctuation Smoothing" (T3); and systems with no obvious time pattern are classified as "Comprehensive Regulation" (T4), reflecting the energy storage operation strategy of the site.

[0080] ④ Energy Storage Response Characteristic Tags (ENERGY_STORAGE_RESPONSE_CHARACTERISTICS): Evaluates the energy storage system's responsiveness and speed to external dispatch signals. Based on the energy storage system's response characteristics to grid dispatch commands, electricity price signals, or system demand, it can be categorized as follows: response time <100 milliseconds is "ultra-fast response" (R1); response time 100 milliseconds-1 second is "fast response" (R2); response time 1-5 seconds is "medium-speed response" (R3); and response time >5 seconds is "slow response" (R4), etc., providing an evaluation basis for energy storage's participation in grid ancillary services and system regulation.

[0081] ⑤ Energy Storage Cycle Count Tag (ENERGY_STORAGE_CYCLE_COUNT / time): This tag counts the number of charge-discharge cycles of the energy storage system within a specific time period, reflecting the usage frequency and lifespan of the energy storage system.

[0082] S2.2 Distributed site tagging system based on multi-source information: As smart grid and energy internet technologies continue to develop, the precise characterization of distributed power generation station energy production behavior can rely on more diversified data to drive this process. Currently, power generation behavior patterns, energy configuration structures, and operational characteristics exhibit significant diversity and heterogeneity, and the dimensions of data acquisition have broken through the constraints of traditional single channels. Information data encompasses the station monitoring system, equipment operation logs, and meteorological monitoring data at the direct objective monitoring layer; station operation and maintenance records and equipment maintenance archives at the management layer; and can further integrate meteorological forecast system data, geographic information system (GIS) regional attribute data, and power dispatch system operation data at the ecological linkage layer. The fusion and collaborative analysis of multi-source data can comprehensively cover dimensions such as the station's basic attributes, energy interaction patterns, and operational decision-making rules, providing a solid multi-dimensional data foundation for the construction of a precise energy management system, the optimization design of electricity pricing strategies, and the implementation of power grid dispatch mechanisms.

[0083] The power plant monitoring system and equipment operation logs can accurately capture the real-time power generation behavior trajectory of the power plant, presenting an objective and continuous dynamic of energy production. Meteorological monitoring data and power dispatching system operation data can supplement the power plant's energy production capacity and system response characteristics from the perspectives of environmental conditions and system interaction. Meanwhile, the power plant operation and maintenance records and equipment maintenance archives, with the advantage of directly recording the power plant's operating status, can deeply mine information that is difficult to obtain completely through real-time monitoring, such as changes in equipment performance and fluctuations in operating efficiency. To clearly demonstrate the power plant classification path under multi-source data collaboration, this embodiment integrates multi-source data such as the power plant monitoring system, meteorological monitoring system, geographic information system, and operation and maintenance archives. By constructing a tag system that integrates the characteristics of multi-dimensional data sources, it collaboratively refines the characterization of power plant features, supporting the practice of precise energy management and daily data verification.

[0084] S2.2.1 Tagging system for multi-source data fusion: Based on the multi-dimensional attributes of multi-source data such as station monitoring systems, meteorological monitoring data, GIS regional attributes, and operation and maintenance records, a four-layer tagging system covering basic attributes, energy configuration, operational characteristics, and interactive responses is constructed. This system, through collaborative coding and feature integration of multi-source data, achieves a comprehensive characterization of the energy production characteristics of distributed power stations, providing multi-dimensional reference for daily data verification.

[0085] The tagging system is built upon standardized coding rules, using differentiated coding strategies to standardize the representation of multi-source heterogeneous information such as data from power plant monitoring systems, meteorological monitoring records, GIS regional attribute information, and operation and maintenance archives. Combining the existence characteristics of remote monitoring and operation and maintenance records from power plant monitoring systems (dual verification) and grid dispatch response participation marking, binary coding is used to map this information to discrete 0-1 values. For quantifiable attributes such as power generation, equipment runtime, and installed capacity collected by the monitoring system, numerical coding is used to preserve the original quantitative characteristics. For discrete classification data such as power plant type, grid connection method, and climate zone code, a semantic-numerical mapping relationship is constructed through independent coding. Through differentiated coding processing of multi-source data, a structured tagging system encompassing basic attributes, energy production, and interactive responses is constructed, laying the data foundation for subsequent daily data verification of distributed power plants.

[0086] S2.2.2, Variable significance analysis based on analysis of variance (ANOVA): After completing the construction of the multi-source data labeling system, it is necessary to further verify the differences in the contribution of each dimension of labels to the classification of distributed power stations. To analyze the nonlinear correlation characteristics between different labels and power station power generation behavior, analysis of variance (ANOVA) can be used to assess the significant differences in the correlation between different labels and power station power generation behavior. Based on ANOVA, by calculating the ratio of between-group variance to within-group variance, the system evaluates the significant impact of label variables derived from different data sources, such as power station monitoring systems and meteorological monitoring data, on the classification results of power station groups. The core idea of ​​ANOVA is to decompose the total variation of observed data into components attributable to different sources and compare the relative sizes of these sources to determine whether there are statistically significant differences in the means between different treatment groups or conditions. Traditional statistical methods for statistically significant differences suffer from false positive inflation when repeatedly using two-sample t-tests to compare the means of three or more independent groups. ANOVA considers the treatment effect to be real and effective if the variation caused by between different groups is significantly greater than the variation caused by within-group random error.

[0087] Specifically, the ANOVA analysis begins with the calculation of the Total Sum of Squares (SST). The SST reflects the total sum of squares (SST) around its total mean. The sum of the dispersion of the data represents the total variation of the entire data set. Its calculation formula is: (11) in, Number of representative groups (number of treatment levels); Representing the The sample size of the group; Representing the Group 1 One observation value; It is the average of all observed values.

[0088] ANOVA divides the total sum of squares (SST) into two main parts: the between-groups sum of squares (SSB or SSTR) and the within-groups sum of squares (SSW or SSE).

[0089] (12) Between-group sum of squares (SSB) measures the difference between the means of different groups. ) and the overall mean ( The sum of squares of the differences between the groups reflects the variation caused by different grouping factors: (13) Within-group sum of squares (SSW) is a measure of the sum of squares (SSW) of each observation within a group relative to its own group mean. The sum of squared differences (SQP) represents the variation caused by random error or individual differences, and is usually regarded as the background noise of the sequence. (14) However, the size of the sum of squares (SSB and SSW) is affected by the number of groups ( ) and sample size ( The impact of this is such that directly assessing the sample based on the sum of squares is unreasonable. To make fair comparisons, the mean square between groups (MSB) is used to assess the sample: (15) Among them, the degrees of freedom between groups The formula for calculating the within-group mean square (MSW) is: (16) Among them, the degrees of freedom within the group (N is the total sample size). Assuming the null hypothesis holds and the ANOVA assumptions are satisfied, MSW is essentially the pooled within-group variance, which is a measure of the common population variance. An estimate.

[0090] The core test statistic for ANOVA is the F-statistic, which is defined as the ratio of the between-group mean square (MSB) to the within-group mean square (MSW): (17) Under the null hypothesis ( Under the condition that the following holds true and the basic assumptions of ANOVA (independence of observations, within-group normality, homogeneity of variance) are satisfied, the F-statistic follows a sequence with numerator degrees of freedom of . The denominator has degrees of freedom. The F-distribution. If there are indeed differences in the means between different groups, i.e., the alternative hypothesis... If at least two means are unequal, then MSB will tend to be greater than MSW, resulting in an F-value greater than 1. The calculated F-value is then compared with the critical value of the F-distribution at a selected significance level (α, typically 0.05). Compare. If There is sufficient statistical evidence to reject the null hypothesis and conclude that there are at least two groups with significantly different population means.

[0091] (18) (19) in, This intuitively explains the proportion of total variation contributed by differences between groups. That is to The adjusted estimate aims to reduce its positive bias and provide an estimate of the overall effect size that is closer to unbiased.

[0092] When applying ANOVA, it is crucial to carefully assess whether its key assumptions are met: the independence of observations is essential and is usually ensured through experimental design (such as randomization); the data within each group should approximately follow a normal distribution; ANOVA is robust to this assumption with large sample sizes, but severe skewness may affect the results; one of the most important assumptions is homogeneity of variance, meaning that the population variances of each group are equal. .

[0093] S2.2.3, Construction of Multi-Source Information Tags: Using ANOVA-based variable significance analysis, the influence of multi-source data fusion labels on electricity consumption behavior indicators such as peak-valley differences, average electricity consumption levels, and electricity consumption fluctuations was quantified. Labels with large, medium, and small influence were selected to construct multi-source information labels for daily clearing data of distributed power stations.

[0094] S2.3 Classification of power generation and consumption behavior of distributed power stations based on cluster analysis: The power generation behavior of distributed power plants is influenced by multiple factors, including installed capacity type, meteorological conditions, and equipment status, exhibiting significant heterogeneity. These differences significantly impact the verification of daily power plant data. Scientifically classifying distributed power plant power generation types and establishing classification standards is a crucial foundation for achieving effective daily data verification for distributed power plants.

[0095] S2.3.1 Selection of Clustering Method: In the daily data verification process, the combined application of K-shape clustering and Hierarchical clustering has significant advantages. K-shape clustering analyzes long-term daily data from distributed data stations, categorizing curves to capture the dynamic changes in data, accurately identifying data curve patterns from different days, and providing a normal operating baseline for data verification. This clustering method helps to discover the periodic and fluctuating characteristics of daily data from data stations, distinguishing between normal and abnormal patterns, thus laying the foundation for subsequent accurate verification.

[0096] Hierarchical clustering further categorizes data anomaly patterns into different types based on the varying curve proportions of massive data stations under different weather conditions and seasons, uncovering the seasonal and meteorological dependence patterns of data anomalies. This method considers the impact of time and environmental factors on data quality at data stations, identifying data anomaly types with similar characteristics by analyzing the proportion of various data anomaly patterns in different time periods, such as communication interruption, equipment failure, and data offset. The two clustering methods complement each other: K-shape clustering focuses on individual data curve characteristics, while Hierarchical clustering grasps the patterns of group data anomalies. Their combination allows for a comprehensive and in-depth construction of data verification models from both micro and macro perspectives, providing strong support for power companies to conduct precise data quality management and optimize data processing workflows. The specific methods are as follows.

[0097] (1) K-shape clustering: Distributed facility daily data is time series data. Due to its high dimensionality and low information density, traditional clustering methods are not very effective for processing time series data. K-shape clustering is an algorithm specifically designed for time series clustering. Its principle is similar to K-means clustering, but by improving the distance metric and cluster center selection method, K-shape clustering improves the verification capability of time series data while maintaining computational efficiency.

[0098] Traditional clustering methods mostly rely on Euclidean distance to calculate the similarity between samples. While Euclidean distance is simple and intuitive to calculate, it cannot accurately reflect the amplitude distortion and phase lead-lag relationships between sequences. To address this, some researchers have adopted distance metrics based on dynamic time warping (DTW), using local nonlinear alignment to reflect amplitude distortion and phase relationships between sequences. However, this computational complexity increases significantly with increasing data volume, raising computational costs. K-shape clustering proposes a novel distance metric—Shape-based distance (SBD)—which achieves global phase alignment between different sequences by calculating the normalized cross-correlation (NCC) between them.

[0099] The calculation process of SBD is as follows: In order to solve the amplitude distortion problem and achieve proportional invariance, each time series is first standardized by Z-score so that the mean of each series is 0 and the standard deviation is 1.

[0100] For two time series of equal length and In order to achieve shift invariance, S-step shift Represented as: (20) (twenty one) (twenty two) Among them, when hour, Indicates swipe to the right The position is filled with zeros on the left; when hour, Indicates swiping left Fill the space with zeros on the right; Let sequence Keep it static, and make exist Swipe up to calculate and Inner product of each shift Arrange them to obtain a cross-correlation coefficient sequence of length 2m-1. By normalizing its coefficients, we can obtain the NCC, which can be expressed as follows: (twenty three) (twenty four) (25) Find the position that maximizes NCC by comparing the size of the NCC sequence. This allows us to determine the maximum similarity between the two sequences and the total step size required for the translation. To facilitate distance measurement, the SBD between two sequences is calculated as follows. The NCC ranges from [-1, 1], and thus the SBD ranges from [0, 2]. The smaller the SBD, the more similar the two sequences are.

[0101] (26) To achieve efficient solution of SBD, it is necessary to require a sequence of cross-correlation coefficients. Fast computation. Since the convolution of two time series can be obtained through Fourier transform, we first perform discrete Fourier transforms on both sequences, then calculate the product of the two sequences, and finally perform inverse discrete Fourier transform to calculate the SBD. The specific calculation process can be expressed as follows: (27) in, This indicates taking the complex conjugate in the frequency domain. and Let represent the Discrete Fourier Transform and its inverse, respectively, and calculate them as follows: (28) (29) in, Indicates the length of the time series. Represents a sequence The value in the time domain, Represents a sequence The value in the frequency domain, .

[0102] The calculation speed of cross-correlation coefficients can be greatly improved by using the Fast Fourier Transform, thus enabling rapid calculation of SBD.

[0103] In verifying daily data from distributed data stations, attention is paid to both the similarity of the overall curve shape and the temporal distribution characteristics of outliers. While SBD distance demonstrates excellent performance in clustering similarly shaped sequences, its overemphasis on shape similarity leads to some outliers with significantly different temporal distributions being identified as belonging to the same anomaly type. This embodiment proposes a path-aware weighting mechanism that introduces a local distortion penalty factor on top of the DTW path. This distance metric innovatively combines the global alignment capability of DTW with the sensitivity of local Euclidean distance. The path-aware weighting mechanism balances the flexibility of temporal distortion with the need for shape preservation, providing a more refined similarity characterization for anomaly pattern recognition in daily data.

[0104] Dynamic Time Warping (DTW) can extract the nonlinear temporal similarity relationship between two time series, and can be used to match the delay similarity between two fluctuating sequences. (The last sentence appears to be incomplete and possibly refers to a separate topic: "Taking two time series...") and For example, let X represent the standard sequence and Y represent the sequence to be warped. Calculate the Euclidean distance between the two sequences. Construct an N×M Euclidean distance matrix. The goal of the DTW algorithm is to find the shortest path in the matrix. The objective function can be expressed as: (30) in, Representing the Each fluctuation process corresponds to a time pair; This represents the total number of time pairs. Its constraints include endpoint constraints, monotonicity constraints, and continuity constraints, expressed as: (31) (32) (33) Under these constraints, the DTW algorithm becomes a shortest path problem. Dynamic programming, Floyd's algorithm, and Dijkstra's algorithm can be used to solve this type of problem. However, the DTW algorithm only focuses on the similarity of values ​​between two sequences, without considering the distortion of the x-axis (timestamp axis). Therefore, a penalty term needs to be added for points with severe distortion to reduce the distance coefficient of sequences with large peak-valley time differences. Since the station load sequences are time-span sequences with consistent time resolution, the degree of distortion of different paths can be represented by the path forward direction, where the penalty coefficient of the k-th sequence pair can be expressed as: (34) in, This is the index of the element corresponding to the k-th time pair in the standard time series X; This is the index of the element corresponding to the (k-1)th time pair in the standard time series X; This is the element index corresponding to the k-th time pair in the time series Y to be distorted; This is the element index corresponding to the (k-1)th time pair in the time series Y to be distorted.

[0105] The above formula penalizes off-diagonal shifts (time axis distortion) in the path; the greater the stretch (i.e., the further the path deviates from the diagonal), the smaller the weight. The distance between path pairs is determined by... The DTW distance between the two sequences is further calculated using a penalty factor and Softmax weighting as follows: (35) in, This is the vector corresponding to the standard time series X; Let Y be the vector corresponding to the time series to be warped. For the standard sequence X, the index is The element value; For the index of the sequence Y to be distorted The element values ​​are then used. The DTW distance of this path-aware weighted mechanism is used as the sequence distance evaluation metric in subsequent K-shape clustering.

[0106] Traditional partition-based clustering algorithms often determine cluster centers by calculating the arithmetic mean of all elements in each sample, and use this mean as the representative of that class of samples. However, this calculation method also ignores the differences in amplitude and phase between time series when applied to time series clustering. Therefore, K-shape transforms the process of finding cluster centers into an optimization problem. Considering that cross-correlation coefficients can reflect the similarity between sequences, it selects the sequence that maximizes the sum of squared cross-correlation coefficients of all sequences within the cluster. As the cluster center, the optimization function can be expressed as: (36) in, Indicates the first k The set of all time series in a cluster.

[0107] Before and after two iterations of clustering calculations, the cluster center sequence generally does not show a significant shift. Therefore, it can be roughly assumed that the other sequences within the cluster are aligned with the cluster center sequence in phase. Thus, the denominator in the above formula can be omitted, simplifying it to: (37) To simplify the expression, the above formula is converted into vector form: (38) Since the cluster center sequence is a calculated value and not necessarily the actual time series existing within the cluster, therefore, in the above formula... Without Z-score standardization, and to simplify calculations, K-shape omits the standard deviation calculation in Z-score standardization and instead performs an approximate standard deviation calculation: (39) in, I It is the identity matrix. J It is a matrix of all ones. m This represents the length of the time series.

[0108] This can be transformed into the Rayleigh quotient form in linear algebra for solving, thus yielding the cluster centers: (40) in, Let be the cluster center sequence of the k-th cluster, used to represent the time series vector of that cluster. It is its transpose vector; The transformation matrix used for approximate standard deviation calculation is defined as follows: I is the identity matrix, J is a matrix of all ones, and m is the length of the time series. for The transpose of the matrix; Let P be a single time series vector within the k-th cluster, belonging to the cluster sequence set P. k , for The transpose of .

[0109] (2) Hierarchical clustering: Hierarchical clustering algorithms occupy an important position in the clustering method system. Unlike partitioning clustering algorithms such as K-means, which require a pre-defined number of clusters, hierarchical clustering constructs a tree-like hierarchical structure to cluster data in a progressive manner. It can not only explore the cluster structure of data, but also show the hierarchical relationship between data. It plays an irreplaceable role in many academic and practical applications, such as species kinship analysis in biological taxonomy and customer group hierarchy construction in market segmentation.

[0110] Hierarchical clustering algorithms can be clearly divided into two main categories based on their clustering direction: Agglomerative Hierarchical Clustering (AHC) and Divisive Hierarchical Clustering (DHC). Agglomerative Hierarchical Clustering (AHC) follows a bottom-up clustering logic. Initially, each data point in the dataset is treated as an independent cluster unit, meaning each data point forms its own cluster. Subsequently, based on a pre-defined inter-cluster similarity metric, the algorithm iteratively searches for the two clusters with the highest similarity (closest distance) and merges them into a new cluster. As the iteration continues, the number of clusters gradually decreases, while the cluster size continuously expands until specific termination conditions are met, such as reaching a preset number of clusters or the inter-cluster distance exceeding a set threshold, ultimately forming a hierarchical clustering result. Distributed power station daily data patterns are complex and the number of clusters is difficult to determine in advance. However, AHC eliminates the need to pre-set the number of clusters and can flexibly divide the data according to actual needs. Furthermore, distributed power station data exhibits hierarchical characteristics across different dimensions such as seasons and weather conditions, which can be clearly displayed through the tree-like hierarchical structure constructed by AHC. In addition, the various distance metrics in AHC are robust to noise and outliers in the data, ensuring stable verification results. Moreover, the power sector not only needs to classify data anomalies but also needs to understand the relationships between categories. The clustering hierarchy diagram generated by AHC can intuitively present the category merging process, providing a comprehensive basis for data verification strategy formulation and anomaly handling process optimization.

[0111] In applications of hierarchical clustering algorithms, the data representation and related assumptions are fundamental to the algorithm's operation. Let the dataset... , which includes n Each data object. All located in m In the dimensional feature space, it can be represented as a vector. ,Right now , here ( ) represents a data object In the j The values ​​can be taken in each feature dimension.

[0112] Because hierarchical clustering is sensitive to sample dimensionality, principal component analysis (PCA) is needed to reduce the dimensionality of samples when the dimensionality is large. PCA is a multivariate statistical analysis technique for data compression and feature extraction. It transforms multiple correlated variables into a few uncorrelated composite variables, which contain most of the information provided by the original variables. This method constructs a series of linear combinations of the original variables to form new variables, ensuring that these new variables reflect as much information as possible from the original variables while remaining uncorrelated with each other. Data information is mainly reflected in the variance of the data variables; the larger the variance, the more information it contains. The cumulative variance contribution rate is usually used to measure this. PCA involves obtaining the correlation matrix from the data matrix formed by the input variables of multiple samples, obtaining the cumulative variance contribution rate based on the eigenvalues ​​of the correlation matrix, and then determining the principal components based on the eigenvectors of the correlation matrix. The specific steps are as follows: 1) Standardization of raw data. To eliminate the impact of different units and excessively large numerical differences in the original variables, the original variables are standardized. Assume there are m indicators. X 1 ,X 2 ,…,X m Each characteristic of an object is represented by a matrix. If there are N objects, they can be represented by an N×m matrix, i.e.: (4-38) First, a central standardization process is performed to generate a standard matrix. ,Right now (4-39) in, ; ; , They are indicator variables The mean and variance of.

[0113] 2) Establish the correlation matrix R, and calculate its eigenvalues ​​and eigenvectors, i.e.: (4-40) In the formula Given the standardized data matrix, the eigenvalues ​​of the autocorrelation matrix R are obtained. and the corresponding feature vectors u 1 ,u 2 ,…,u m .

[0114] 3) Determine the number of principal components. The variance contribution rate and cumulative variance contribution rate are respectively: (4-41) (4-42) The number of principal components selected depends on the cumulative variance contribution rate. Usually, when the cumulative variance contribution rate is greater than 75% to 95%, the first p principal components contain most of the information that the m original variables can provide, and the number of principal components is p.

[0115] 4) The eigenvectors corresponding to the p principal components are Then the matrix formed by the p principal components of the N samples is: (4-43) Select the first p-th order principal components after transformation As input variables after dimensionality reduction.

[0116] To quantify the similarity between data objects, this embodiment introduces a distance metric function. This function maps two m-dimensional data vectors to a non-negative real number. The value of this real number reflects the distance between the two data objects; the smaller the distance, the higher the similarity between the two data objects. In practical applications, Euclidean distance is one of the most commonly used distance metrics, and its calculation formula is: (41) in, and These are the feature vectors of two data objects.

[0117] The specific process of agglomerative hierarchical clustering is as follows: 1) Initialization: When the agglomerative hierarchical clustering algorithm starts, it performs initialization operations. This involves processing each data point in the dataset. Initialize each as an independent cluster This forms the initial cluster set. At this point, the number of clusters in the cluster set is equal to the number of data points in the dataset, each cluster contains only one data point, and the entire clustering system is in its most dispersed state.

[0118] 2) Distance calculation: After initialization, the algorithm proceeds to the distance calculation phase. Based on the selected inter-cluster distance metric, the current cluster set is... All cluster pairs (in , Represents the current cluster set The number of clusters (cardinality) is used to calculate the distance between any two clusters. Different distance metrics measure the similarity between clusters from different perspectives, and their calculation results will directly affect subsequent cluster merging decisions. 3) Cluster merging: After obtaining the distances of all cluster pairs, the cluster pair with the smallest distance is found by comparing these distance values. That is, satisfying This means clustering C p and C q They have the highest similarity among all current cluster pairs, therefore they are merged into a new cluster. After the merge operation is completed, the cluster set is... Update and remove the existing clusters. C p and C q and the newly generated clusters C pq Adding a cluster reduces the size of the cluster set by one, resulting in a corresponding change in the cluster structure.

[0119] 4) Iteration Termination: After completing one cluster merge operation, the algorithm continues to repeat the distance calculation and cluster merge steps. During this iterative process, the number of clusters continuously decreases, while the size of the clusters gradually increases. The algorithm stops iterating when a pre-defined termination condition is met. The termination condition typically includes reaching a specified number of clusters k, i.e., when the cluster set... The process stops when the number of clusters decreases to k; or when the maximum inter-cluster distance exceeds a threshold. This means that continuing to merge clusters at this point may lead to an unreasonable cluster structure, thus stopping the algorithm. 5) Output of results: When the algorithm stops running and meets the termination condition, the cluster set at this point... This is the final clustering result. The clustering result exhibits a hierarchical structural feature. The dendritic clustering diagram, which is a tree structure, can intuitively show the merging process and hierarchical relationships of the clusters, providing data analysts with rich information and facilitating further understanding of the internal structure and patterns of the data.

[0120] This embodiment uses the Ward method to measure the distance between clusters. The core objective of the Ward method is to minimize the increase in the sum of squared errors (SSE) of all observations within a cluster during each merging process, thereby ensuring that the data points within each cluster are as close as possible and the differences between clusters are as obvious as possible in the clustering results.

[0121] For a cluster It contains n data points Clustering center of mass The calculation formula is: (42) Then the intra-cluster squared error of this cluster SSE C The sum of the squares of the distances from each data point to the centroid, i.e.: (43) in, Indicates that the data points are in The value at the kth feature dimension; Indicates the center of mass The value taken on the k-th feature dimension, where m is the feature dimension of the data point.

[0122] Suppose there are two clusters and Their respective intra-cluster squared errors are as follows: SSE i and SSE j The number of data points are respectively and When these two clusters are merged into a new cluster... At that time, the centroid of the new cluster for: (44) New clustering Intra-cluster squared error SSE ij for: (45) Ward Distance Defined as the increase in the total squared error resulting from merging two clusters, i.e.: (46) Through derivation, the Ward distance can be expressed in a more concise form: (47) in, It is clustering and The distance between the centers of mass is usually calculated using Euclidean distance: (48) in, and Representing clustering and The value of the centroid on the k-th feature dimension.

[0123] S2.3.2, Construction of Electricity Consumption Behavior Pattern Tags: Based on the above two-step clustering results, the first step is to obtain multi-cluster distributed station data patterns based on the daily clearing data curves of the stations; based on the clusters obtained in the first step, the second step further refines the proportion of data anomaly patterns in seasonal differences and weather condition differences, and obtains several types of distributed station data anomaly pattern labels at the annual time scale, providing accurate classification basis for daily clearing data verification.

[0124] S2.4, Distributed Facility Daily Clearing Data Baseline Classification Construction Process: First, resource characteristic tags for power plants are constructed. Starting from the power generation load and internal resources of the power plants, the resource configuration of the power plants is tagged, classifying them into different types such as pure power generation type, photovoltaic type, wind power type, energy storage type, wind-solar type, photovoltaic-storage type, wind-storage type, and wind-solar-storage type. Simultaneously, power generation characteristic tags for the power plants are extracted, including key indicators such as peak power generation, average daily power generation, and power generation ratio. If the power plant is equipped with distributed photovoltaic devices, photovoltaic power generation characteristic tags are extracted, including key indicators such as grid connection method, installed capacity, and photovoltaic utilization hours. If the power plant is equipped with distributed wind power devices, wind power characteristic tags are extracted, including key indicators such as grid connection method, installed capacity, and wind power volatility. If the power plant is equipped with distributed energy storage devices, energy storage characteristic tags are extracted, including key indicators such as energy storage capacity, maximum charging and discharging power, charging and discharging time preference, and response characteristics, to gain a deeper understanding of the power plant's behavior.

[0125] Secondly, a multi-source information tagging system is constructed. This integrates multi-source information, including the basic attributes of the power plant, power generation behavior, energy configuration, and interactive responses, to further enrich the data verification benchmark. ANOVA is used to evaluate the differences in the contribution of multi-source information to indicators such as power generation volatility and data anomaly rate, screening key characteristic factors affecting data quality and providing a more comprehensive assessment of the power plant's data characteristics.

[0126] Finally, based on the clustering analysis results, anomaly pattern labels were constructed for the daily clearing data. To ensure data integrity and accuracy, the daily clearing data from the stations underwent preprocessing, including random forest imputation, EMD smoothing, and Z-score standardization to improve data quality. Furthermore, an improved DTW distance and K-shape clustering algorithm were used to classify the daily clearing data curves from the stations, identifying normal and anomaly patterns. Then, considering seasonal and weather condition differences, a seasonal-weather coupled feature vector was constructed. Hierarchical clustering was then used to further refine the data anomaly types, obtaining individual category labels for data anomalies, providing accurate classification criteria for subsequent data verification.

[0127] S3. Distributed site daily data verification method: Based on the constructed normal pattern library and the abnormal pattern clustering results of path-aware weighted improved DTW+K-shape, as well as the formed multi-dimensional feature label system (covering resource configuration labels, multi-source information labels, and abnormal pattern labels), this embodiment designs a full-process daily clearing data intelligent verification mechanism to achieve closed-loop operation from pattern matching and quantification of deviations to anomaly repair and threshold self-optimization. The verification process logically inherits the results of the first two parts: using normal patterns as the comparison benchmark, providing physical, environmental, and equipment reference conditions with multi-dimensional labels, guiding the anomaly type determination through abnormal pattern labels, and using intelligent algorithms to achieve adaptive data correction, thereby ensuring the long-term stability and reliability of large-scale distributed site operation data.

[0128] Newly acquired daily workload sequence First, locate its comparison dimension in the multi-dimensional feature label space, that is, match labels with the same site and the same resource configuration. L r Same climate zone label L c Seasonal-weather coupling tags L stw Normal pattern set below The normal pattern set is represented by the corresponding cluster center curves from the aforementioned clustering results. Its statistical distribution composition, This is the index for the normal pattern set. To simultaneously quantify shape deviation and amplitude offset, the following dual metrics are defined: (49) (50) in, The path-aware weighted Shape-Based Distance function, proposed in the first part, introduces a shape-based distance with a local distortion penalty. It inherits the phase alignment advantage of K-shape clustering and suppresses anomalous mismatches due to temporal misalignment. It uses shape similarity as an index... Characterize the similarity of curve shapes. These reflect the overall amplitude ratio differences. These two quantities, along with the daily meteorological deviation index extracted from the multi-dimensional information tags, reflect the overall amplitude ratio differences. Input the binary Gaussian discriminant model together: (51) in, , and These are the mean and covariance matrices under normal conditions, respectively. This achieves a dynamic balance between physical consistency and statistical anomalies, preventing rigid judgments in the classification labels of the first and second parts in real-world applications.

[0129] When the distributed discriminant model provides an anomaly determination, it will directly classify the anomaly into missing, fluctuating, offset, or equipment failure anomalies based on the anomaly pattern label results in the second part. For missing anomalies (i.e.... The improved quantum genetic optimization random forest regressor proposed in Part 1 is invoked, taking multi-dimensional feature labels as input (e.g., resource allocation, installed capacity, energy storage response labels, daily weather, etc.) to perform conditional imputation on missing points. The prediction formula is as follows: (52) in, Indicates missing points The predicted value; This represents a random forest regression model with hyperparameters of . , which represents the optimal combination of parameters that minimizes the mean square error; This represents the feature vector of the missing point, which includes the correction value of the nearest time, historical data of the same day, meteorological features, label information, etc.; N is the number of training samples; This represents the true value of the j-th sample in the training set. and represent the number of decision trees and the maximum tree depth, respectively. The global optimum is given by the quantum genetic algorithm. The chromosome probability is initialized to inherit the adaptive rotation angle strategy of the improved QGA in the first part, thus inheriting its fast convergence characteristics in high-dimensional sparse feature space.

[0130] For fluctuating anomalies, this embodiment first obtains the set of intrinsic mode functions through EMD decomposition. Select the high-frequency component set As a source of fluctuations, real-time meteorological features from the second part of the multi-source information label are then introduced. Establish power response weighting function Correcting the fluctuation components: (53) in, The value of the h-th high-frequency mode function obtained from EMD decomposition at time k; This represents the corrected high-frequency modal components; Meteorological weighting coefficients representing high-frequency components; The power response function characterizes the mapping of meteorological features to the theoretical power change magnitude. This is the meteorological feature vector corresponding to time k; This is a very small constant. Finally, the corrected component is superimposed with the low-frequency component and processed through the mode library subspace. Perform shape-constrained projection to obtain the repaired version. .

[0131] For offset-type anomalies, the overall offset is estimated by jointly using the device status label from the multidimensional label and the historical offset distribution. ,implement: (54) in, This represents the corrected daily curve; This indicates the uniform amplitude offset throughout the day; Operators for statistical median; This indicates the low-output period of the power station. This offset estimation based on the low-output period follows physical constraints and is adjusted by correction coefficients based on the output conditions of photovoltaic, wind power, etc., making the offset correction equally effective for multi-energy power stations.

[0132] For equipment failure-type anomalies, the repair value is constrained by combining the resource configuration tag and the equipment available capacity tag: (55) in, This represents the repair value at time k; Indicates the percentage of available equipment capacity. ; The first in the pattern library k Maximum normal output at any given time; Indicates samples in the pattern library t In the k Efforts are made at all times. Simultaneously, the potential curves of external meteorological driving models are prioritized to replace model values, ensuring that the repaired data does not exceed physical limits while improving the realism of the recovery.

[0133] The entire verification process forms a dynamic self-learning closed loop: after each verification, new normal curve samples are added to the normal pattern library, which in turn affects the discrimination model parameters. and repair strategy parameters Perform Bayesian optimization to minimize the overall cost function as follows: (56) in, , , These are all weighting coefficients, set by the site owner or business scenario; Indicates the root mean square error; The goodness-of-fit coefficient; The percentage of inconsistencies between the corrected data and historical normal patterns primarily stems from the difference rate between the corrected data and historical daily patterns of the same type. Through this closed-loop optimization driven by multi-label, multi-pattern, and multi-source data, the daily data verification in Part Three not only continues the preprocessing and classification results of Parts One and Two but also elevates them to an online, evolving intelligent quality maintenance system. This system forms a tight coupling between pattern matching, anomaly detection, and physical correction, ensuring the long-term high reliability and high utilization value of the data under dynamic changes in multiple factors such as weather, seasons, and equipment aging.

[0134] Example 2: This embodiment provides a distributed daily data verification system for airfields based on multi-dimensional feature fusion, including: The data acquisition module is configured to acquire daily clearing data from distributed data stations. The data processing module is configured to perform power generation data cleaning, feature extraction, and standardization on the daily data of distributed power stations. Specifically, when cleaning the power generation data, missing values ​​are filled in using the random forest algorithm, and the quantum genetic algorithm is used to iteratively optimize the number of generated trees and the maximum depth of the trees in the random forest algorithm. The tag creation module is configured to: construct a multi-level data verification tag system based on the standardized distributed daily clearing data of the stations from two dimensions: resource configuration characteristics and multi-source information characteristics; when constructing the multi-level data verification tag system from the multi-source information characteristics dimension, the module uses variance analysis to screen feature factors that have a significant impact on data quality, introduces a local distortion penalty factor on the basis of dynamic time warping path, combines the global alignment capability of dynamic time warping with the sensitivity of local Euclidean distance, balances the flexibility of time warping with the need to maintain the shape through a path-aware weighting mechanism, classifies the daily clearing data curves of the stations through time series clustering, identifies normal and abnormal patterns, and considers seasonal and weather condition differences, refines the data anomaly types through hierarchical clustering, and obtains separate category tags for data anomalies; The verification module is configured to perform distributed daily data verification of the stations based on the constructed multi-level data verification label system.

[0135] The working method of the system is the same as that of the distributed station daily clearing data verification method based on multi-dimensional feature fusion in Example 1, and will not be repeated here.

[0136] Example 3: This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the distributed field station daily clearing data verification method based on multi-dimensional feature fusion described in Embodiment 1.

[0137] Example 4: This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the program, it implements the steps of the distributed field station daily clearing data verification method based on multi-dimensional feature fusion described in Embodiment 1.

[0138] Example 5: This embodiment provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the method described in Embodiment 1.

[0139] The above description is merely a preferred embodiment of this practice and is not intended to limit the scope of this practice. Various modifications and variations can be made to this practice by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this practice should be included within the protection scope of this practice.

Claims

1. A distributed field station daily data verification method based on multi-dimensional feature fusion, characterized in that, include: Obtain daily clearing data from distributed data stations; The daily power generation data of distributed power stations is cleaned, features are extracted, and standardized. In the process of cleaning the power generation data, missing values ​​are filled in by the random forest algorithm, and the quantum genetic algorithm is used to iteratively optimize the number of generated trees and the maximum depth of the trees in the random forest algorithm. Based on the standardized distributed daily clearing data of the stations, a multi-level data verification label system is constructed from two dimensions: resource allocation characteristics and multi-source information characteristics. When constructing the multi-level data verification label system from the multi-source information characteristics dimension, variance analysis is used to screen feature factors that have a significant impact on data quality. A local distortion penalty factor is introduced on the basis of dynamic time warping path to combine the global alignment capability of dynamic time warping with the sensitivity of local Euclidean distance. A path-aware weighting mechanism is used to balance the flexibility of time warping with the need to maintain the shape. The daily clearing data curves of the stations are classified by time series clustering to identify normal and abnormal patterns. Seasonal differences and weather condition differences are also considered. Hierarchical clustering is used to refine the data anomaly types and obtain separate category labels for data anomalies. Based on the constructed multi-level data verification and labeling system, distributed daily data verification of the stations is carried out.

2. The distributed station daily clearing data verification method based on multi-dimensional feature fusion as described in claim 1, characterized in that, Filling in missing values ​​using the random forest algorithm involves: dividing the electricity time-series data of the distributed power stations into daily segments for each preset time period to obtain daily clear time-series data of a preset length; using known variable data in the time series as features and variables with missing values ​​as labels; and filling in the missing values ​​using the random forest algorithm.

3. The distributed station daily clearing data verification method based on multi-dimensional feature fusion as described in claim 2, characterized in that, A quantum genetic algorithm is used to iteratively optimize the number of spanning trees and the maximum depth of trees in the random forest algorithm. This includes: filtering current and charge time series data to be filled based on the missing rate; setting the population size and quantum encoding bits for the quantum genetic algorithm; initializing the probability amplitude of each individual and generating binary codes for key parameters of random forest regression; performing random forest regression on non-missing data, defining a fitness function based on the square root error obtained from the random forest regression model, calculating the fitness values ​​of all individuals in the population, and selecting the individual with the highest fitness as the evolutionary direction for the next generation; determining whether the termination condition is met; if so, terminating the quantum genetic algorithm to obtain the optimal random forest regression parameters and performing random forest regression on the current and charge time series to be filled; otherwise, calculating the quantum rotation gate rotation angle for each individual and updating the probability amplitude of all individuals.

4. The distributed station daily clearing data verification method based on multi-dimensional feature fusion as described in claim 3, characterized in that, Adaptive adjustment of rotation angle: ; in, and This represents the minimum or maximum rotation angle. and These are the minimum and maximum fitness of an individual in the current population, respectively. f It is the fitness value of the current individual; if the fitness of the current individual is less than the preset value, the quantum mutation adopts the first preset rotation angle to accelerate the mutation speed of the individual; if the fitness of the current individual is greater than or equal to the preset value, the quantum mutation adopts the second preset rotation angle to make it converge to the optimal solution; the first preset rotation angle is greater than the second preset rotation angle.

5. The distributed station daily clearing data verification method based on multi-dimensional feature fusion as described in claim 1, characterized in that, Establish a distributed power station labeling system oriented towards resource allocation: First, based on the internal resource types of the power stations, systematically classify power stations with different energy resource combinations; second, based on the resource characteristics of the power station configuration, extract key power generation and consumption behavior indicators to form refined power station labels.

6. The distributed station daily clearing data verification method based on multi-dimensional feature fusion as described in claim 5, characterized in that, Distributed photovoltaic (PV), distributed wind power, and distributed energy storage are used as three variables to form a combination space with plant power load. The resource configuration combinations and tags for each category include pure plant power type, PV type, wind power type, energy storage type, wind-solar type, solar-storage type, wind-storage type, and wind-solar-storage type. According to the internal resource configuration of the plant, different types of plants include plant power load characteristic tags, distributed PV operation characteristic tags, distributed wind power operation characteristic tags, and distributed energy storage operation characteristic tags.

7. The distributed station daily clearing data verification method based on multi-dimensional feature fusion as described in claim 1, characterized in that, By integrating multi-source data from power plant monitoring systems, meteorological monitoring systems, geographic information systems, and operation and maintenance records, a tagging system that integrates the characteristics of multi-dimensional data sources is constructed to collaboratively refine the characterization of power plant features. This involves combining remote monitoring from the power plant monitoring system with dual verification of operation and maintenance records, and using the existence characteristics of grid dispatch response participation markers. Binary encoding is used to map this information to discrete 0-1 values. For quantifiable attributes such as power generation, equipment runtime, and installed capacity collected by the monitoring system, numerical encoding is used to preserve the original quantitative characteristics. For discrete classification data such as power plant type, grid connection method, and climate zone code, a semantic-numerical mapping relationship is constructed through independent encoding. Based on the analysis of variance method, the significance of label variables derived from different data sources on the classification results of the station population is evaluated by calculating the ratio between-group variance and within-group variance.

8. The distributed station daily clearing data verification method based on multi-dimensional feature fusion as described in claim 1, characterized in that, Based on the dynamic time-warped path, a local distortion penalty factor is introduced to combine the global alignment capability of dynamic time warping with the sensitivity of local Euclidean distance. A path-aware weighting mechanism balances the flexibility of time warping with the requirement of shape preservation, providing similarity characterization for daily data anomaly pattern recognition. The penalty factor is: ; in, This is the index of the element corresponding to the k-th time pair in the standard time series X; This is the index of the element corresponding to the (k-1)th time pair in the standard time series X; This is the element index corresponding to the k-th time pair in the time series Y to be distorted; This is the element index corresponding to the (k-1)th time pair in the time series Y to be distorted; The dynamic time bending distance is: ; Among them, among them, This is the vector corresponding to the standard time series X; Let Y be the vector corresponding to the time series to be warped. For the index of the standard sequence X The element value; For the index of the sequence Y to be distorted The element value; The sequence that maximizes the sum of squared cross-correlation coefficients of all sequences within the cluster is selected as the cluster center. : ; in, Let be the cluster center sequence of the k-th cluster; It is its transpose vector; The transformation matrix used for approximate standard deviation calculation is defined as follows: I is the identity matrix, and m is the length of the time series. for The transpose of the matrix; For a single time series vector within the k-th cluster, for The transpose of .

9. The distributed station daily clearing data verification method based on multi-dimensional feature fusion as described in claim 8, characterized in that, This paper employs agglomerative hierarchical clustering, a type of hierarchical clustering algorithm, following a bottom-up clustering logic. Initially, each data point in the dataset is treated as an independent clustering unit. Based on a pre-defined inter-cluster similarity metric, the algorithm iteratively searches for the two clusters with the highest similarity and merges them into a new cluster. The hierarchical clustering algorithm includes: initializing each data point in the dataset as an independent cluster to form an initial cluster set; calculating the distance between all cluster pairs in the current cluster set; after obtaining the distances of all cluster pairs, finding the cluster pair with the smallest distance by comparison, determining the cluster with the highest similarity among all current cluster pairs, merging the clusters with the highest similarity into a new cluster, and updating the cluster set.

10. The distributed station daily clearing data verification method based on multi-dimensional feature fusion as described in claim 1, characterized in that, The newly acquired data is matched with the normal patterns of the corresponding stations for deviation detection to identify and locate the anomaly types. For missing data, a random forest regression algorithm optimized by quantum genetics is used to complete it. For abnormal fluctuations, data offsets and equipment failures, targeted corrections are made in combination with station resource characteristics, real-time weather and historical operation curves. During the verification process, the anomaly pattern library and feature labels are dynamically updated to continuously optimize the discrimination threshold and repair strategy.

11. A distributed daily data verification system for depots based on multi-dimensional feature fusion, characterized in that, include: The data acquisition module is configured to acquire daily clearing data from distributed farms. The data processing module is configured to perform power generation data cleaning, feature extraction, and standardization on the daily data of distributed power stations. Specifically, when cleaning the power generation data, missing values ​​are filled in using the random forest algorithm, and the quantum genetic algorithm is used to iteratively optimize the number of generated trees and the maximum depth of the trees in the random forest algorithm. The tag creation module is configured to: construct a multi-level data verification tag system based on the standardized distributed daily clearing data of the stations from two dimensions: resource configuration characteristics and multi-source information characteristics; when constructing the multi-level data verification tag system from the multi-source information characteristics dimension, the module uses variance analysis to screen feature factors that have a significant impact on data quality, introduces a local distortion penalty factor on the basis of dynamic time warping path, combines the global alignment capability of dynamic time warping with the sensitivity of local Euclidean distance, balances the flexibility of time warping with the need to maintain the shape through a path-aware weighting mechanism, classifies the daily clearing data curves of the stations through time series clustering, identifies normal and abnormal patterns, and considers seasonal and weather condition differences, refines the data anomaly types through hierarchical clustering, and obtains separate category tags for data anomalies; The verification module is configured to perform distributed daily data verification of the stations based on the constructed multi-level data verification label system.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the distributed daily clearing data verification method for field stations based on multi-dimensional feature fusion as described in any one of claims 1-10.

13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that, When the processor executes the program, it implements the steps of the distributed daily clearing data verification method for field stations based on multi-dimensional feature fusion as described in any one of claims 1-10.

14. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the distributed daily clearing data verification method for field stations based on multi-dimensional feature fusion as described in any one of claims 1-10.