Physical constraint fused time sequence intelligent deviation correction method

By using adaptive evolutionary optimization of geospatial clustering and a hybrid feature-driven dynamic correction prediction model, the problem of poor wind speed data quality in mountainous areas was solved, achieving high-precision data correction and reliable meteorological data support.

CN121502175AInactive Publication Date: 2026-02-10NANJING UNIV OF INFORMATION SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610044744.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-14
Publication Date
2026-02-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies struggle to provide high-quality meteorological data, especially wind speed data, in mountainous environments. There are numerous missing measurements, jumps, drifts, and non-physical abrupt changes. Furthermore, traditional methods are ill-suited to the strong nonlinearity and time-varying nature of wind speed, resulting in inconsistent data quality that fails to meet the needs of power grid operations.

Method used

An adaptive evolutionary optimization geospatial clustering model is combined with a physical consistency anomaly factor model and a hybrid feature-driven adaptive dynamic correction prediction model. By constructing a physical-statistical joint constraint envelope and adaptive dynamic correction prediction, physically consistent and temporally self-consistent correction values ​​are generated.

Benefits of technology

It achieves high-precision correction of wind speed data in mountainous areas, provides stable and reliable basic data support, and improves the accuracy and reliability of wind power assessment and line micro-meteorological monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121502175A_ABST
    Figure CN121502175A_ABST
Patent Text Reader

Abstract

The invention discloses a physical constraint-fused time sequence intelligent correction method, which comprises the following steps of: carrying out double clustering based on a meteorological station geographic position and a meteorological element time sequence curve shape by adopting a self-adaptive evolutionary optimization geographic space clustering model; an optimal classification scheme is selected through a weighted evaluation index of the contour coefficient and the intra-cluster distance quadratic sum; on the basis of the classification result, noise identification and marking are carried out on each type of meteorological element data by jointly using a local abnormal factor algorithm and dynamic physical constraint conditions based on cut-in wind speed, cut-out wind speed and a theoretical power curve; constructing a mixed feature refining model, and forming a feature set with high expression ability; carrying out 15-minute single-step prediction by using a self-adaptive dynamic correction prediction model driven by mixed features; comprehensively evaluating the prediction result; the prediction precision is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of geographic correction technology, specifically to a temporal intelligent correction method that integrates physical constraints. Background Technology

[0002] In mountainous areas, due to complex terrain, limited communication conditions, and difficulties in constructing observation points, conventional meteorological stations are extremely sparse, making it difficult to form effective coverage of the regional meteorological field. In recent years, although power grid companies have deployed a number of miniature meteorological monitoring devices on transmission line towers to fill the spatial gaps in conventional observations, the high difficulty of inspections in mountainous areas, high equipment maintenance costs, and the tendency of extreme environments to degrade sensor performance have resulted in a large number of missing measurements, jumps, drifts, and non-physical abrupt changes in the meteorological data from these towers, leading to inconsistent raw data quality. Traditional statistical methods rely on manually set empirical thresholds, which are difficult to adapt to the strong nonlinearity and time-varying characteristics of meteorological elements such as wind speed in mountainous areas, and are prone to missed detections or misjudgments. While purely data-driven machine learning models can capture certain temporal patterns, they lack physical boundary constraints on wind speed changes; in mountainous environments where terrain has a significant impact, these models are prone to producing physically inconsistent correction results, such as predicting wind speed values ​​beyond the actual possible range, making the corrected data unusable for subsequent analysis and decision-making. Furthermore, the large spatial differences between meteorological stations in mountainous areas, high sampling noise, and short effective observation windows make it difficult for existing methods to guarantee stability in anomaly identification, feature representation, and short-term forecasting. Especially when tower meteorological data maintenance cycles are long and equipment failure rates are high, outliers often appear in consecutive segments. Traditional methods often cannot reliably fill in long-sequence anomalies and lack a reliable quantification mechanism to determine the reliability of correction results. In summary, existing technologies cannot meet the demand for high-quality meteorological data in actual mountain power grid operations. There is an urgent need for a meteorological data correction method that can simultaneously combine physical constraints, temporal structure, and intelligent forecasting capabilities to reliably correct raw data in mountainous environments characterized by sparse observations, difficult maintenance, and frequent anomalies. This would provide stable and reliable basic data support for applications such as wind power assessment, line micrometeorological monitoring, and mountain meteorological forecasting. Summary of the Invention

[0003] Purpose of the Invention: The purpose of this invention is to provide a time-series intelligent correction method that integrates physical constraints. It uses a physical consistency anomaly factor model to construct a physical-statistical joint constraint envelope for noise labeling, and calls a hybrid feature-driven adaptive dynamic correction prediction model to generate physically consistent and time-series self-consistent correction values ​​by backtracking outliers for 15 minutes. This provides stable and reliable basic data support for applications such as wind power assessment, line micro-meteorological monitoring, and mountain meteorological forecasting, and solves the problems existing in the background technology.

[0004] Technical Solution: The present invention provides a time-series intelligent error correction method that integrates physical constraints, applicable to wind speed data error correction, comprising the following steps:

[0005] (1) An adaptive evolutionary optimization geospatial clustering model is adopted, driven by the dual features of geographical structure and temporal morphology, and combined with composite statistical weighted evaluation indexes to achieve adaptive selection of the optimal clustering scheme;

[0006] (2) Based on the classification results, a physical consistency anomaly factor model is constructed for each type of meteorological data. The physical-statistical joint constraint envelope is constructed by using dynamic physical constraints and local data structures. Data outside the physical-statistical joint constraint envelope is marked as outliers. The physical-statistical joint constraint envelope is identified and constructed by the physical consistency anomaly factor model. The physical consistency anomaly factor model marks outliers for samples that meet the above dual inconsistencies by comprehensively judging whether the data simultaneously violates the dynamic physical constraints and presents statistically significant outlier characteristics.

[0007] (3) Construct a hybrid feature refinement model, extract multi-domain features by integrating time domain, frequency domain and statistical domain, and combine PCA dimensionality reduction and random forest feature importance screening to form a hybrid feature set with high expressive power and strong discriminative power;

[0008] (4) After the outlier is automatically retrieved and marked, the input sequence is constructed based on the high-confidence normal data within the previous 15 minutes, and the hybrid feature-driven adaptive dynamic correction prediction model is called to complete the 15-minute scale single-step back prediction of the moment, thereby generating a physically consistent and temporally self-consistent correction value.

[0009] (5) The prediction results are comprehensively evaluated based on the weighted integrated index of time series prediction error and uncertainty quantification credibility;

[0010] (6) When the prediction result meets both the accuracy constraint and the credibility threshold requirement, the model uses the prediction value as the final correction result to fill in missing or abnormal data, thereby realizing the spatiotemporal consistency and continuity reconstruction of the data sequence.

[0011] Furthermore, in step (1), the adaptive evolutionary optimization geospatial clustering model is as follows: the K-means clustering is optimized using a genetic algorithm that adaptively and dynamically adjusts the crossover probability and mutation probability, with an initial crossover probability of 0.8, an initial mutation probability of 0.2, and a decrease rate of 0.02 per generation; the optimal number of clusters for K-shape clustering is automatically determined using a particle swarm optimization algorithm combined with the silhouette coefficient, and the convergence efficiency is improved through a gradient inertia weight reduction strategy.

[0012] Furthermore, in step (2), the dynamic physical constraints are specifically: setting the cut-in wind speed and cut-out wind speed range, and setting a dynamic buffer around the theoretical wind speed curve, with the buffer width adaptively adjusted by the wind speed change rate.

[0013] Furthermore, the constraint expression for the dynamic buffer is:

[0014]

[0015] in, This is the actual wind speed value. This is the theoretical wind speed value. The rated wind speed value; when a point is identified as an outlier by both of the above outlier detection methods, the point is identified as an outlier and filled with a null value.

[0016] Further, in step (4), the hybrid feature refinement model is as follows: In the multi-domain feature extraction stage, the original wind speed sequence is reconstructed into a set of high-dimensional feature matrices describing the signal trend, period, amplitude and energy through cross-domain mapping; principal component analysis is introduced to reduce the dimensionality of the high-dimensional feature matrix, and finally, based on the residual feature screening mechanism, principal components with feature importance higher than the median are retained; among them, the time domain features are used to describe the local fluctuation features of wind speed changing with time, and the methods such as adjacent time step difference, rate of change and second derivative are used to characterize the response behavior of wind speed to changes in the real-time meteorological field; the frequency domain features are used to characterize the periodicity, energy distribution and scale change of the wind speed sequence as aggregated spectral features, power spectral density, wavelet band energy and Fourier entropy by performing fast Fourier transform and wavelet transform on the wind speed sequence, so as to realize the systematic characterization of the frequency domain of the wind speed curve; the statistical domain features are used to characterize the stability and variation of the wind speed sequence from the perspective of the overall distribution, and the measurement indicators include variance, standard deviation, extreme value difference and absolute energy index.

[0017] Furthermore, the feature selection mechanism for the residuals is as follows: the wind power is initially fitted using the baseline model, the predicted residuals are regarded as the target quantity, the nonlinear mapping between the residuals and the principal components is modeled using random forest regression, and the feature importance of each principal component to the residuals is calculated.

[0018] Furthermore, in step (5), the adaptive dynamic correction prediction model is as follows: the preliminary prediction results of the ARIMA model are used as the observation values ​​and input into the adaptive Kalman filter system, and dynamic correction is achieved by updating the state estimate and error covariance in real time; the formula is as follows:

[0019] ;

[0020] ;

[0021] ;

[0022] in, These are the real-time predictions from the ARIMA model. The original wind power sequence;

[0023] ;

[0024] in, For a moment The best estimate of the system state based on past observations and forecasts. For the observation matrix, Let be the state transition matrix.

[0025] Furthermore, in step (6), the confidence-weighted integration index is composed of the root mean square error, mean absolute error, and prediction interval coverage index. The weights for different application scenarios should be dynamically adapted according to the actual operation of the corresponding meteorological stations; the formula is as follows:

[0026] ;

[0027] in, The weights are non-negative and satisfy the following conditions: , This represents the maximum wind speed at the station.

[0028] Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: It proposes an adaptive evolutionary optimization geospatial clustering model, achieving efficient and accurate classification of wind turbines. Simultaneously, it employs a local anomaly factor algorithm combined with a self-developed physical constraint to jointly construct a physical-statistical joint constraint envelope, achieving high-accuracy noise identification. Before filling out outliers, this invention innovatively proposes and introduces a hybrid feature refinement framework to enrich the training set features, improving the model's ability to express complex dynamic patterns and facilitating subsequent predictions with higher accuracy. When filling out outliers, it creatively proposes a hybrid feature-driven adaptive dynamic correction prediction model for single-step 15-minute backtracking correction. Furthermore, it achieves comprehensive accuracy evaluation of the prediction results through a weighted integrated index of time-series prediction error and uncertainty quantification credibility, significantly improving the accuracy of outlier correction. Attached Figure Description

[0029] Figure 1 This is a flowchart of the present invention. Detailed Implementation

[0030] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0031] like Figure 1 As shown, this embodiment of the invention provides a time-series intelligent correction method that integrates physical constraints, including the following steps:

[0032] (1) An adaptive evolutionary optimization geospatial clustering model is used to classify wind turbines based on both geographical location and wind power curve shape, and the best classification result is selected. Specifically, a genetic algorithm improved by adaptive dynamic crossover and mutation probability is used in conjunction with the K-means clustering algorithm to cluster meteorological stations based on their geographical locations; and a K-shape algorithm is used in conjunction with the particle swarm optimization algorithm to cluster the measured wind speed curves of meteorological stations. Theoretically, it is believed that the wind speed curves measured by meteorological stations that are closer together are more similar, so the clustering results of the two methods will not differ much. By comparing the contour coefficients and sum of squared Euclidean distances of the various types of stations obtained by the two methods, the most suitable station classification method is selected as the best wind turbine category. The main improvement of the K-means algorithm is that a genetic algorithm is used to help update the cluster centers in each iteration, and gradient is used to dynamically change the crossover probability and mutation probability to achieve large crossover in the early stage and convergence in the later stage, and large mutation in the early stage and stability in the later stage, so that the model can converge faster. The main improvement of the K-shape algorithm is that particle swarm optimization is used in conjunction with contour coefficients to find the optimal number of clusters, and dynamic parameter adjustment is also used to make the model converge quickly, which helps to improve the clustering accuracy.

[0033] (2) Based on each type of site, noise is identified and denoised using the local anomaly factor algorithm and physical constraint method.

[0034] (3) Use a hybrid feature refinement framework to enrich the features of the training set.

[0035] (4) Based on the aforementioned data preprocessing and feature construction results, a hybrid feature-driven adaptive dynamic correction prediction model is used for ultra-short-term single-step prediction (15 min) and real-time correction. Finally, based on the credibility weighted integration index of time-series prediction error and uncertainty quantification, the model prediction results are comprehensively evaluated for accuracy. After the overall evaluation accuracy meets the standard, the predicted value is used as the repair value to replace the original outlier value. The implementation process of the hybrid feature-driven adaptive dynamic correction prediction model is as follows: using historical wind speed sequences and training set data refined by multi-domain features, an ARIMA(p,d,q) model is trained to obtain preliminary prediction results; the predicted value is used as the observation input to introduce an adaptive Kalman filter system, and the prediction error is corrected and the optimization result is output through state estimation and dynamic optimization.

[0036] (5) The prediction results are comprehensively evaluated based on the weighted integrated index of time series prediction error and uncertainty confidence.

[0037] In step (1), K-means is a typical unsupervised clustering algorithm based on Euclidean distance to measure sample similarity. Its basic idea is to obtain a stable clustering partition by iteratively optimizing the within-cluster sum of squares (WCSS) between samples within a cluster and their corresponding centroids. The algorithm recalculates the centroids and updates the sample affiliation in each iteration. If the WCSS change is lower than a preset threshold for two consecutive iterations, the algorithm is considered to have converged and outputs the final cluster structure. In wind speed prediction scenarios, K-means clustering can be performed on wind turbines based on their geographical location to provide structured grouping for subsequent power range modeling. To evaluate the site clustering effect, the silhouette coefficient is introduced as a performance indicator. Its profile coefficient is denoted as The overall clustering quality can be described as:

[0038] ;

[0039] Where N is the number of categories. The range of values ​​is The closer the value is to 1, the better the clustering effect.

[0040] Genetic Algorithms (GAs) are a class of global optimization methods that simulate the processes of natural selection and inheritance. Their basic process includes: randomly generating an initial population, calculating individual fitness, selecting individuals based on fitness, generating offspring according to a set crossover probability, perturbing some genes based on mutation probability, and retaining historically best individuals. The core of genetic algorithms is fitness-proportional selection; when using proportional selection, individuals... The probability of being selected is:

[0041] ;

[0042] in, Individual The probability of being selected. It is the first The fitness of an individual It is the first The fitness of an individual It refers to population size.

[0043] Traditional genetic algorithms typically employ fixed crossover and mutation probabilities, which can easily lead to low search efficiency or premature convergence. To improve search efficiency, this study introduces a linear annealing scheme, similar to simulated annealing, which gradually decreases the crossover and mutation probabilities with each iteration. This maintains strong exploration capabilities in the early stages of the algorithm and enhances convergence stability in later stages. The dynamic probabilities can be expressed as:

[0044] ;

[0045] ;

[0046] in, For the current algebra, For the maximum number of generations, For dynamic crossover probability, The initial crossover probability, This represents the dynamic mutation probability. This represents the initial mutation probability. In this embodiment, 1000 iterations are used. Take 0.7, Let's take 0.2. Therefore, the genetic operator is:

[0047] ;

[0048] in, Represented by probability Execute the given crossover operator, Represented by probability Execute the given mutation operator, Indicates the first generation, This indicates the evaluation strategy.

[0049] Because K-means is sensitive to initial centroids, random initialization can easily lead to the algorithm getting trapped in local optima and unstable results. This model introduces Global Algorithm (GA) into the centroid search stage. By globally optimizing multiple candidate centroid sets in continuous iterations, using the silhouette coefficient as the fitness function to guide the search direction, and combining it with a dynamic probability mechanism to improve the stability of the solution and global exploration capability, it significantly reduces the risk of getting trapped in local optima and improves the performance of the final clustering structure. The overall framework integrating GA and K-means can be formally expressed as:

[0050] ;

[0051] in, For the set of centroids, This indicates the last generation of the iteration. The objective function of the dynamic probabilistic genetic algorithm is defined in this framework, which effectively combines global initialization with local fine-grained search.

[0052] K-shape is an unsupervised clustering method specifically designed for time series data. Its core idea is to capture the shape similarity of time series data to identify objects with similar dynamic change characteristics. In this study, it is used to cluster observation stations based on the similarity characteristics between wind speed data, thereby identifying groups of stations with similar output fluctuation patterns. Finally, the results obtained from K-shape and K-means are fused to form a more stable clustering scheme with more consistent physical properties.

[0053] To avoid manually setting the number of clusters To mitigate the subjective bias introduced by K-shape and further enhance its ability to represent wind speed temporal characteristics, this study introduces Dynamic Inertial Weighted Particle Swarm Optimization (PSO) to automatically search for the number of clusters in K-shape. PSO considers the particle's "position" as the candidate cluster number and its "velocity" as the search direction. In each iteration, the particle's velocity and position are updated in the following way:

[0054] ;

[0055] ;

[0056] in, Represents particles In time speed, Represents particles In time Location, and These represent individual historical optimality and global optimality, cognitive factors and social factors, respectively. All are set to 1.5. yes Random numbers within the interval, inertia weight A phased dynamic strategy is adopted to achieve the goal of wide-area search in the early stage and accelerated convergence in the later stage:

[0057] ;

[0058] During the optimization process, the negative silhouette coefficient is used as the fitness function, i.e.:

[0059] ;

[0060] in, is the silhouette coefficient, k is the number of clusters, and the optimal position obtained by adjustment is rounded and directly used as the input number of clusters for K-shape.

[0061] In this way, the dynamic weighted average sorting (PSO) constitutes the "automatic clustering optimization module," while the K-shape serves as the "time series clustering module" to classify the morphological characteristics of wind speed. The combination of the two enhances the objectivity and stability of the clustering process, making the structural identification of multidimensional time-series data from meteorological stations more stable and reliable.

[0062] In step (2), the local anomaly factor algorithm is an unsupervised anomaly detection algorithm based on density for local outlier detection. It identifies anomalies by comparing the local density differences between a sample and other samples in its neighborhood. Its core idea is that if the local reachability density of a point is significantly lower than that of the surrounding points, it indicates that the point is in a sparse region and can be regarded as a potential anomaly.

[0063] Building upon the LOF algorithm constraints, this model further introduces physical constraints to improve the engineering rationality of anomaly identification. The main idea is to limit outliers caused by wind speeds decreasing to zero or being excessively high by setting cut-in and cut-out wind speed values. Using the theoretical wind speed curve, a buffer zone (c ranging from [0,1]) is established around the curve using a constraint exponent c. Data outside this range is considered outliers. The physical constraints are as follows:

[0064] ;

[0065] in, This is the actual power value. This is the theoretical power value. This represents the rated power value. When a point is identified as an outlier by both of the above outlier detection methods, that point is classified as an outlier and filled with a null value.

[0066] In step (3), the constructed hybrid feature refinement framework aims to extract multi-scale information from the wind speed time series curves of the stations and transform it into an effective feature set suitable for the prediction model to improve prediction accuracy. The framework mainly consists of three parts: multi-domain (time domain, frequency domain, statistical domain) feature extraction, feature dimensionality reduction, and important feature selection based on residuals.

[0067] First, in the multi-domain feature extraction stage, the main objective is to construct a multi-domain feature system based on the physical logic of meteorological element prediction. Specifically, time-domain features are used to describe the local fluctuations of wind speed over time, employing methods such as adjacent time step differences, rates of change, and second derivatives to characterize the response of wind speed to changes in the real-time meteorological field. Frequency-domain features, through Fast Fourier Transform (FFT) and wavelet transform, characterize the periodicity, energy distribution, and scale variations of the wind speed sequence as aggregated spectral features, power spectral density, wavelet band energy, and Fourier entropy, thus achieving a systematic characterization of the wind speed curve in the frequency domain. Statistical domain features characterize the stability and amplitude of the wind speed sequence from an overall distribution perspective, using indicators including variance, standard deviation, extreme value difference, and absolute energy. Through this cross-domain mapping, the original wind speed sequence is reconstructed into a set of high-dimensional statistics describing the signal trend, periodicity, amplitude, and energy.

[0068] However, multi-domain features are usually large in scale, which can easily lead to the curse of dimensionality or overfitting, resulting in a decrease in the accuracy of the prediction model. Therefore, the second stage introduces the Principal Component Analysis (PCA) method to perform dimensionality reduction on the feature matrix. By using linear projection, high-dimensional features are compressed into a low-dimensional space that retains the main variance contribution (95% variance is retained in the case study), thereby reducing redundant features and improving the compactness of feature representation.

[0069] In the reduced principal component space, this model further constructs a feature selection mechanism based on residuals. First, the baseline model is used to initially fit the wind power, treating the predicted residuals as the target quantity. Random forest regression is used to model the nonlinear mapping between the residuals and the principal components, and the feature importance of each principal component to the residuals is calculated. The residuals here reflect complex structures not yet captured by the baseline model; therefore, the importance metric of random forest can effectively identify features sensitive to prediction errors and crucial for improving model performance. Finally, principal components with feature importance higher than the median are retained as input features for subsequent prediction models. This step, through a three-stage refinement process of multi-domain information fusion, redundant feature compression, and highlighting key structural features, largely preserves the dominant variation patterns, implicit periodicity, and prediction-sensitive structural features in the wind power time series, thereby improving the accuracy and robustness of subsequent steps and making the wind speed model more closely reflect the dynamic behavior of real wind fields.

[0070] In step (4), an adaptive dynamic correction model is used to perform a 15-minute single-step prediction and the prediction results are corrected and updated in real time.

[0071] To achieve ultra-short-term (15-minute) wind speed prediction for meteorological fields, this invention constructs an adaptive dynamic correction prediction model based on an ARIMA model and an adaptive Kalman filter. This model takes historical wind speed sequences and features extracted through a feature refinement framework as input, and first performs time series modeling using an ARIMA model. The ARIMA model consists of three parts: autoregressive AR(p), difference I(d), and moving average MA(q), which are used to characterize correlation, stationarization trend, and noise and abrupt changes, respectively. Specifically, for any wind speed sequence... The ARIMA(p,d,q) model determines its structure by setting the autoregressive order p, differencing order d, and moving average order q based on the ACF-PACF test results, and performs single-step predictions based on historical sequences and feature data to obtain initial predicted values. The ARIMA model can capture the overall trend and periodic changes in wind power, but its response to rapid fluctuations and sudden changes in the wind field is limited. To further improve prediction accuracy and robustness, an adaptive dynamic Kalman filter is introduced to correct the ARIMA predictions.

[0072] The Kalman filter is a recursive optimal state estimation method that dynamically updates the state estimate and its uncertainties by fusing predicted and observed values. To address the irregular abrupt changes in meteorological fields and wind speeds, this model introduces an adaptive noise adjustment system based on the standard Kalman filter. Specifically, the observation noise is dynamically adjusted according to the mean and standard deviation of the ARIMA model prediction errors over the most recent m time steps. With process noise The specific method can be expressed as follows:

[0073] ;

[0074] ;

[0075] ;

[0076] in, These are the real-time predictions from the ARIMA model. This is the original wind power sequence.

[0077] By employing an adaptive mechanism that dynamically adjusts the noise covariance, the model can dynamically adjust its confidence level in observed or predicted values ​​based on wind speed fluctuations. When wind speed fluctuations intensify, the model automatically increases noise and strengthens its reliance on observed values, while when wind speeds are stable, it strengthens its reliance on its own trend for prediction, thereby effectively improving the adaptability and reliability of short-term wind speed forecasts. Based on the above mechanism, the final predicted value of this model is:

[0078] ;

[0079] in, For a moment The best estimate of the system state based on past observations and forecasts. For the observation matrix, The state transition matrix is ​​used. Finally, the result with the smallest corrected prediction error is output as the predicted value. This method preserves ARIMA's ability to capture time series trends and periodic changes while utilizing an adaptive Kalman filter to improve its response to rapid changes and noise. This significantly improves prediction accuracy and stability in ultra-short-term wind speed forecasting, providing a reliable basis for weather-dependent scheduling and operation.

[0080] In step (5), the Error-Uncertainty Quantification Credibility Weighted Integrated Index for Time-Series Prediction (EUQC-WII-TS) is a multi-dimensional collaborative evaluation model designed for meteorological field correction prediction. It integrates quantified prediction error, verification of uncertainty credibility, and engineering adaptability to construct a multi-dimensional collaborative evaluation system oriented towards the actual operational needs of meteorological fields. Its core indicators consist of Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Prediction Interval Coverage Probability (PICP), and meteorological field scenario strain evaluation indicators. By establishing a balance between error sensitivity and interval confidence, it achieves a systematic evaluation of the model's prediction robustness under different operating conditions. Its overall formula is:

[0081] ;

[0082] in, The weights are non-negative and satisfy the following conditions: , This represents the maximum wind speed at the station.

[0083] In the formula, the RMSE term is highly sensitive to large deviations and reflects the model's predictive stability under extreme conditions such as sudden wind speed jumps and cutting out of risk zones. It is an important quantitative method for measuring the model's ability to cope with strong non-stationary signals. The MAE term characterizes the model's average deviation level during normal operation, focusing more on measuring the overall smoothness and consistency of the prediction, reflecting its applicability in long-term continuous scheduling. The PICP term is used to verify the coverage reliability of the prediction interval and is a key indicator for measuring the model's uncertainty estimation ability, reflecting the effectiveness of the prediction system in risk control and decision security. The weighting of the three terms balances statistical accuracy and risk credibility, enabling it to characterize the prediction system's performance in actual wind farm environments from multiple perspectives. The closer the indicator is to 1, the better the model is in error control, interval coverage, and engineering adaptability, and the more reliable the wind speed prediction can be under typical conditions such as high wind shear, wind speed fluctuations, and load scheduling constraints. This indicator allows for a unified evaluation of different prediction models, providing quantitative basis for wind farm dispatch strategies, reserve capacity planning, and operational risk management. It also makes prediction results more interpretable and has greater engineering value, enhancing the effectiveness and credibility of the model under real wind farm operating conditions.

[0084] This invention selects measured data from a wind turbine factory in Pizhou City, Jiangsu Province, primarily using a wind speed sample dataset, which includes the operating status data of 33 wind turbine units from July 1, 2024 to February 1, 2025. The above data is used to predict wind speed using this wind speed prediction model. This model divides the original wind turbine data into five categories (the results of K-means clustering using a fusion of genetic algorithm and K-shape clustering using a fusion of particle swarm optimization algorithm are shown in Tables 1 and 2; the final classification result is based on the average profile coefficient comparison). Then, outlier identification and labeling are performed for each of the five categories to obtain processed data (the parameters for various types of cleanup are shown in Table 3). Subsequently, ACF and PACF judgments are performed on the five-step data to obtain the p, d, and q values ​​for each category (the tailing parameters for each category are shown in Table 4), and a single-step prediction is performed using an adaptive dynamic correction prediction model. The prediction results are shown in Table 5.

[0085] Table 1. K-means clustering results using a fusion genetic algorithm. ; Table 2. K-shape clustering results using the fused particle swarm optimization algorithm. ; Table 3. Setting Table for Various Impurity Removal Parameters ; Table 4. Test results for various ACF and PACF types ; Table 5. Prediction results of the hybrid feature-driven adaptive dynamic correction prediction model .

Claims

1. A time-series intelligent deviation correction method incorporating physical constraints, characterized in that, Includes the following steps: (1) An adaptive evolutionary optimization geospatial clustering model is adopted, driven by the dual features of geographical structure and temporal morphology, and combined with composite statistical weighted evaluation indexes to achieve adaptive selection of the optimal clustering scheme; (2) Based on the classification results, a physical consistency anomaly factor model is constructed for each type of meteorological data. The physical-statistical joint constraint envelope is constructed by using dynamic physical constraints and local data structures. Data outside the physical-statistical joint constraint envelope is marked as outliers. The physical-statistical joint constraint envelope is identified and constructed by the physical consistency anomaly factor model. The physical consistency anomaly factor model marks outliers for samples that meet the dual inconsistencies by comprehensively judging whether the data simultaneously violates the dynamic physical constraints and presents statistically significant outlier characteristics. (3) Construct a hybrid feature refinement model, extract multi-domain features by integrating time domain, frequency domain and statistical domain, and combine PCA dimensionality reduction and random forest feature importance screening to form a hybrid feature set with high expressive power and strong discriminative power; (4) After the outlier is automatically retrieved and marked, the input sequence is constructed based on the high-confidence normal data within the previous 15 minutes, and the hybrid feature-driven adaptive dynamic correction prediction model is called to complete the single-step back prediction on the 15-minute scale, thereby generating a physically consistent and temporally self-consistent correction value. (5) The prediction results are comprehensively evaluated based on the weighted integrated index of time series prediction error and uncertainty quantification credibility; (6) When the prediction result meets both the accuracy constraint and the credibility threshold requirement, the model will use the prediction value as the final correction result to fill in missing or abnormal data, thereby realizing the spatiotemporal consistency and continuity reconstruction of the data sequence.

2. The time-series intelligent correction method integrating physical constraints according to claim 1, characterized in that, In step (1), the adaptive evolution optimization geospatial clustering model is as follows: the genetic algorithm that adaptively and dynamically adjusts the crossover probability and mutation probability is used to optimize K-means clustering, with an initial crossover probability of 0.8, an initial mutation probability of 0.2, and a decrease rate of 0.02 per generation; the particle swarm optimization algorithm combined with the silhouette coefficient is used to automatically determine the optimal number of clusters for K-shape clustering, and the convergence efficiency is improved by the gradient inertia weight decrease strategy.

3. The time-series intelligent correction method integrating physical constraints according to claim 1, characterized in that, In step (2), the dynamic physical constraints are as follows: set the cut-in wind speed and cut-out wind speed range, and set a dynamic buffer around the theoretical wind speed curve. The width of the buffer is adaptively adjusted by the wind speed change rate.

4. The time-series intelligent correction method integrating physical constraints according to claim 3, characterized in that, The constraint expression for the dynamic buffer is: ; in, This is the actual wind speed value. This is the theoretical wind speed value. The value is the rated wind speed. When a point is identified as an outlier by two outlier detection methods, it is considered an outlier and is filled with null values.

5. The time-series intelligent deviation correction method incorporating physical constraints according to claim 1, characterized in that, In step (4), the hybrid feature refinement model is as follows: In the multi-domain feature extraction stage, the original wind speed sequence is reconstructed into a set of high-dimensional feature matrices describing the signal trend, period, amplitude and energy through cross-domain mapping; principal component analysis is introduced to reduce the dimensionality of the high-dimensional feature matrix; finally, based on the feature screening mechanism of residuals, principal components with feature importance higher than the median are retained; among them, the time domain features are used to describe the local fluctuation features of wind speed changing with time, and the methods such as adjacent time step difference, rate of change and second derivative are used to characterize the response behavior of wind speed to changes in real-time meteorological field. Frequency domain features are used to characterize the periodicity, energy distribution, and scale changes of wind speed sequences by performing fast Fourier transform and wavelet transform, as aggregated spectral features, power spectral density, wavelet band energy, and Fourier entropy, thus achieving a systematic characterization of the wind speed curve in the frequency domain. Statistical domain features are used to characterize the stability and variation of wind speed sequences from the perspective of overall distribution, with metrics including variance, standard deviation, extreme value difference, and absolute energy index.

6. The time-series intelligent deviation correction method incorporating physical constraints according to claim 5, characterized in that, The feature selection mechanism for the residuals is as follows: the baseline model is used to initially fit the wind power, the predicted residuals are regarded as the target quantity, random forest regression is used to model the nonlinear mapping between the residuals and the principal components, and the feature importance of each principal component to the residuals is calculated.

7. The time-series intelligent correction method integrating physical constraints according to claim 1, characterized in that, In step (5), the adaptive dynamic correction prediction model is as follows: the preliminary prediction results of the ARIMA model are used as the observation values ​​and input into the adaptive Kalman filter system, and dynamic correction is achieved by updating the state estimate and error covariance in real time; the formula is as follows: ; ; ; in, These are the real-time predictions from the ARIMA model. The original wind power sequence; ; in, For a moment The best estimate of the system state based on past observations and forecasts. For the observation matrix, Let be the state transition matrix.

8. The time-series intelligent deviation correction method incorporating physical constraints according to claim 1, characterized in that, In step (6), the confidence-weighted integration index is composed of the root mean square error, mean absolute error, and prediction interval coverage index. The weights for different application scenarios should be dynamically adapted according to the actual operation of the corresponding meteorological stations; the formula is as follows: ; in, For non-negative weights, satisfying , This represents the maximum wind speed at the station.