Method for identifying and intelligently controlling and cleaning data of automatic weather station in complex terrain wind farm

By constructing a multi-dimensional feature matrix and a hybrid classification model, combined with blockchain technology, the problem of anomaly identification and quality control of wind farm data in complex terrain was solved. This enabled high-precision, low-latency data processing and repair, supported real-time quality control and traceability, and improved the data integrity and business adaptability of wind farms.

CN122432933APending Publication Date: 2026-07-21ELECTRIC POWER RES INST STATE GRID SHANXI ELECTRIC POWER
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ELECTRIC POWER RES INST STATE GRID SHANXI ELECTRIC POWER
Filing Date
2026-05-11
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing quality control methods are difficult to adapt to wind farms with complex terrain, resulting in low accuracy in identifying abnormal data, high missed detection rate, high false detection rate, and low processing efficiency. This affects short-term power forecasting and power generation assessment of wind farms, and increases the wind curtailment rate and the number of unplanned outages.

Method used

A multi-dimensional feature matrix is ​​constructed, and a hybrid classification model combining LightGBM, bidirectional LSTM and attention mechanism is used for initial data inspection and classification. Blockchain technology is used to achieve full-process traceability of quality control, formulate differentiated repair strategies, and adapt to different terrain features and extreme weather scenarios.

Benefits of technology

It achieves an abnormal data identification accuracy of ≥95%, a false negative rate of ≤3%, a false positive rate of ≤5%, a daily data processing time of ≤10 minutes, and a reduction of NMAE of ≥20% when cleaned data is used for power prediction. It adapts to different wind farm terrain characteristics and supports real-time quality control and data traceability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432933A_ABST
    Figure CN122432933A_ABST
Patent Text Reader

Abstract

The application provides a complex terrain wind farm automatic weather station data anomaly identification and intelligent quality control cleaning method, and relates to the technical field of new energy meteorological data processing. The method comprises the following steps: acquiring multi-source data, performing standardization processing and physical quantity conversion, and constructing a multi-dimensional feature matrix; at least one of threshold and logic verification, time sequence consistency detection, distribution offset detection and extreme weather adaptation verification is executed for initial detection, and the data triggering any initial detection abnormal condition is marked as suspicious data; a hybrid classification model containing LightGBM, bidirectional LSTM and attention mechanism is used to classify and determine the suspicious data; and differential repair is respectively performed on different classification labels. The abnormal data identification accuracy of the application is greater than or equal to 95%, the missed detection rate is less than or equal to 3%, and the false detection rate is less than or equal to 5%, and the application can especially accurately distinguish between real extreme data and sensor fault data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of new energy meteorological data processing technology, and in particular to a method for identifying and intelligently cleaning data anomalies from automatic weather stations in wind farms with complex terrain; the automatic weather station includes commonly used meteorological observation equipment in wind farms such as wind measurement towers and multi-element weather stations. Background Technology

[0002] Wind farms in complex terrain, due to their undulating topography and complex airflow, exhibit a far greater variety of anomalies in their automatic weather station (including wind measurement towers and multi-element weather stations) observation data compared to wind farms in plains areas. Specifically: Sensor-level anomalies (strong correlation with wind farms). Wind speed sensors are prone to low readings (deviations can reach -30% to -50%) due to icing in winter, directly affecting the determination of the turbine start-up threshold (e.g., a 3MW turbine start-up wind speed of 3m / s may be misjudged as 2m / s after icing, causing delayed turbine start-up); wind direction sensors may experience periodic jamming due to turbine wake impact (e.g., being fixed at 350° in the turbine wake direction for a long time), interfering with the wind direction-power mapping logic of wind farm micro-site selection and power prediction; temperature and humidity sensors may experience local temperature deviations exceeding ±2℃ (air temperature) and ±5%RH (humidity) due to nacelle heat dissipation radiation, affecting the effectiveness of the turbine's heat dissipation control strategy; Transmission-level anomalies (wind farm terrain constraints). Complex terrain (such as canyons and mountains) obstructs communication links, resulting in daily data loss of 2-4 hours, with the loss periods concentrated during periods of strong winds (such as wind speeds ≥12m / s), directly affecting short-term power forecasting (requiring data integrity ≥95%). Electromagnetic interference (such as electromagnetic radiation from the unit's converter) causes random pulse values ​​(wind speeds jump to 99m / s instantaneously, air pressure drops by 20hPa), which can easily be misjudged as extreme weather, leading to erroneous unit shutdowns (single erroneous shutdown losses exceed 10,000 yuan).

[0003] Environmental anomalies (related to wind farm wind energy characteristics). High turbulence intensity in canyon terrain (turbulence intensity reaches 0.2~0.4), resulting in the superposition of high-frequency noise in wind speed, affecting the calculation of turbulence intensity in wind energy resource assessment (when the turbulence intensity exceeds 0.25, the rated power of the unit needs to be reduced); heavy precipitation and hail cover the sensor probe, causing short-term data distortion (such as humidity suddenly rising to 100% and then dropping to 0).

[0004] Existing quality control methods are difficult to adapt to complex terrain scenarios, and their core shortcomings include: Fixed threshold methods cannot distinguish between real extreme conditions and faults in wind farms: In the prior art, patent application CN202310567890.1 discloses a quality control method for meteorological data of plain wind farms, which uses a fixed threshold method combined with time series analysis, without considering the slope-aspect effect of complex terrain. When applied to the Shanxi canyon wind farm, the false detection rate exceeded 32%. The wind speed limit of 50 m / s recommended in "GB / T 37523—2019 Technical Specification for Topographic Measurement of Onshore Wind Farms" cannot distinguish between real gusts (such as instantaneous 55 m / s) caused by the canyon narrowing effect and sensor faults. The anomaly identification method based on a single LSTM model disclosed in patent application CN202210345678.2 has poor adaptability to the nonlinear diurnal variation of temperature in mountain wind farms, with a false detection rate exceeding 25%. The aforementioned existing technologies do not take into account the terrain characteristics of wind farms and equipment operation and maintenance parameters, resulting in poor quality control in complex terrain scenarios. In the Shanxi Canyon Wind Farm, the misjudgment rate exceeded 32%, which led to the exclusion of effective wind energy data and an 8% increase in the annual power generation assessment deviation.

[0005] Time series analysis methods are poorly adapted to the meteorological patterns of wind farms: The traditional 3σ principle is based on the assumption of normal distribution and is poorly adapted to the nonlinear diurnal variation of mountain air temperature in wind farms (such as the temperature difference between sunny and shady slopes reaching 8~10℃) and the periodic jump of wind direction (the valley wind reverses day and night, and the daily deflection angle exceeds 180°), with a false negative rate of over 25%, which directly leads to a 15% increase in short-term power prediction error.

[0006] The parameters related to the wind farm terrain and the turbine were not taken into account: the relative position of the wind measurement tower and the turbine (such as upwind / downwind direction), altitude difference (500~800m), terrain obstruction (wind speed reduction caused by mountain blocking) and other wind farm-specific parameters were not included. Different wind measurement towers in the same wind farm were not subject to the same quality control standard. For example, high wind speeds on steep windward slopes (real wind energy rich areas) were misjudged as abnormal, resulting in a waste of wind energy resources.

[0007] The processing efficiency cannot meet the real-time business needs of wind farms: The daily observation data of wind farms is ≥1 million (including wind speed and wind direction at the 1-minute level). Traditional serial processing takes more than 60 minutes, which is far from meeting the timeliness requirements of short-term power prediction (requiring the output of quality control results within 15 minutes) and real-time grid dispatch. This results in power prediction lag and an increase in wind curtailment rate of 3%~5%.

[0008] The above problems directly lead to a 15% to 25% increase in short-term power prediction error for wind farms (more than 30% under extreme weather conditions), an annual power generation assessment deviation of more than 5% (calculated for a 100MW wind farm, the annual revenue loss exceeds 2 million yuan), and also affect the safety of grid dispatch and increase the number of unplanned unit shutdowns (more than 5 times per year on average). Summary of the Invention

[0009] This invention provides a method for identifying and intelligently cleaning data anomalies from automatic weather stations in wind farms with complex terrain, in order to achieve at least one of the following objectives: 1) The accuracy rate of abnormal data identification is ≥95%, the false negative rate is ≤3%, and the false positive rate is ≤5%, especially able to accurately distinguish between real extreme data and sensor fault data; 2) Processing time for 1 million data entries per day is ≤10 minutes, with real-time quality control supported and a delay of ≤5 minutes; 3) When the cleaned data is used for power prediction, the normalized mean absolute error (NMAE) is reduced by ≥20%; 4) It can adapt to different wind farm terrain features (altitude, slope, aspect), without the need for repeated manual parameter adjustments, and is suitable for complex scenarios such as mountainous areas, canyons, and plateaus. It can also adapt to extreme weather scenarios and sensor aging scenarios.

[0010] To achieve the above objectives, this invention provides a method for anomaly identification and intelligent quality control cleaning of automatic weather station data in wind farms with complex terrain, the method comprising: Multi-source data is acquired and standardized by unifying time resolution, marking missing data, and initially screening invalid data. Physical quantity transformation is performed on the obtained standardized time-series data to construct a multi-dimensional feature matrix. The multi-source data includes wind farm automatic weather station observation data, geographic information data, equipment operation and maintenance data, and extreme weather and operation data. The multi-dimensional feature matrix includes basic statistical features, cross-parameter physical correlation features, terrain adaptation features, equipment health features, and time and extreme weather features. Based on the standardized time-series data and multi-dimensional feature matrix, at least one of the following preliminary checks is performed: threshold and logic verification, time-series consistency detection, distribution offset detection, and extreme weather adaptation verification. Data that triggers any preliminary check anomaly condition is marked as suspicious data, and the anomaly type is recorded by binary encoding. Using the standardized time-series data, multidimensional feature matrix, and suspicious data as input, a hybrid classification model including LightGBM, bidirectional LSTM, and attention mechanism is used to classify and determine the suspicious data. For high-confidence results with a maximum probability of classification label greater than or equal to a set threshold (preferably 0.9, with a value range of 0.85~0.95), the corresponding final classification label is directly output. For low-confidence results with all label probabilities less than the set threshold, they are marked as suspicious data in an intermediate state (label 3) and manual review is triggered. The final classification label includes normal data, erroneous data, missing data, and sensor aging data. Suspicious data (label 3) is not the final classification label and needs to be reclassified to the above final classification label after manual review. Differentiated repair is performed on different classification labels, invalid data that meets the removal rules is isolated and removed, and full-process traceability information containing original data information, quality control cleaning process information, operation subject and time information is generated for each data point in the whole process, so as to obtain a standardized meteorological dataset after quality control.

[0011] Furthermore, the physical quantity conversion includes wind direction angle conversion, barometric altitude correction, and humidity unit conversion; The wind direction angle conversion includes converting wind directions within the range of 0° to 360°. d The conversion formula to Cartesian coordinates is as follows: u=sin( d × π / 180), v =cos( d ×π / 180) Where u and v For Cartesian coordinate system components; The barometric altitude correction includes correcting the barometric pressure observed by each meteorological tower to sea-level pressure. The correction formula is:

[0012] in, To observe air pressure, h The altitude of the wind measurement tower, T 0 represents the sea-level reference temperature (288.15 K). c Let M be the vertical temperature lapse rate (0.0065 K / m) and M be the molar mass of dry air (0.0289644 kg / mol). R The universal gas constant is 8.31432 J / (mol·K). The acceleration due to gravity (9.80665 m / s²) 2 ); The humidity unit conversion includes converting relative humidity to specific humidity. q The conversion formula is: q =0.622× e / ( P -0.378 e ) in, P This refers to the ambient atmospheric pressure measured on-site by the wind measurement tower. e The vapor pressure is calculated using the following formula: e = e 0×exp[ a ×( T - T 0) / ( T -b )] in, e 0 represents the baseline saturated water vapor pressure (0.61078 hPa), and exp represents the natural exponential function. a The empirical fit coefficient is 17.27. b It is an empirical constant (35.86℃). T This is the actual measured temperature.

[0013] Furthermore, the terrain adaptation features include a slope-aspect correction function and a target-reference station wind speed difference rule, wherein: The slope-aspect correction function f ( s , a (), based on wind tunnel experimental data fitting from wind farms; targeting slope s For plateau terrain with a slope of ≤15°, a basic slope correction factor should be set. f s =1.0+0.008× s For 15° < s For mountainous terrain with a slope of ≤30°, a base slope correction factor should be set. f s =1.15+0.006×(s-15); For s For canyon terrain with a slope greater than 30°, a basic slope correction factor should be set. f s =1.45; the basic slope correction coefficient is dynamically adjusted in conjunction with the angle between the slope aspect and the prevailing wind direction. s When the altitude is ≤30°, the plateau terrain f ( s , a )= f s ×1.08, Mountainous terrain f ( s , a )= f s ×1.1, Canyon Topography f ( s , a )= f s ×1.12; 30°< s When <150°, f ( s , a )= f s When the included angle is ≥150°, the plateau terrain... f ( s , a )= f s×0.88, Mountainous terrain f ( s , a )= f s ×0.9, Canyon terrain f ( s , a )= f s ×0.85; The terrain-corrected wind speed formula is: v '= v / f ( s , a ),in v The measured wind speed is given by s, where s is the slope of the location of the wind measurement tower. a Slope direction, v 'This is the corrected wind speed; The target-reference station wind speed difference rule calculates the wind speed difference Δ between the target meteorological tower and reference meteorological towers within a designated area. v =| v target - v ref |, among which v target The measured wind speed values ​​of the target anemometer tower to be verified during the same period. v ref To reference the concurrently measured wind speed values ​​from the wind measurement tower, and in conjunction with the terrain barrier coefficient... k block Set the difference threshold Δ v max =2× k block ×(1+ s / 30), of which s The slope of the target anemometer tower location; when Δ v >Δ v max When the time is not specified, it is marked as a spatial correlation anomaly; when there is no effective reference wind tower in the surrounding set area, the search range is automatically expanded to 5km.

[0014] Furthermore, the cross-parameter physical correlation features include wind-pressure correlation features, air density features, and wind field stability features; wherein: Calculate the Pearson correlation coefficient between the pressure elevation gradient and wind speed of adjacent anemometer towers. r As a characteristic of wind-pressure correlation, the calculation formula is: r = Cov (Δ P / Δ h , v ) / [ s (Δ P / Δ h )× s ( v )] in, Cov (Δ P / Δ h , v ) represents the covariance between the pressure gradient at altitude and the wind speed. s (Δ P / Δ h ) represents the standard deviation of the barometric altitude gradient. s ( v ) represents the standard deviation of wind speed; R is the ratio of the standard deviation of wind direction to wind speed. dv As a characteristic of wind field stability; The air density characteristic includes air density. r And wind energy density, air density r The calculation formula is: r =1.293×(273.15 / (273.15+ T ))×( P sea / 1013.25)×(1-0.378× U / 100) in, T To measure the actual temperature, P sea To correct air pressure at sea level, U Relative humidity; Wind energy density is r × v 3 .

[0015] Furthermore, the temporal consistency detection includes: based on the standardized time-series data, setting the sliding window size, constructing a VAE model containing a 2-layer LSTM encoder and a 2-layer LSTM decoder, calculating the mean square error (MSE) of the real-time data sequence and the VAE reconstructed sequence; and setting an anomaly threshold (MSE). th =3×μ MSE , where μ MSE The mean MSE of the training set is given when MSE > MSE. th When this occurs, it is marked as a timing anomaly.

[0016] Furthermore, the distribution offset detection includes: Using historical normal data from wind farms as samples, a Gaussian kernel function was employed. K ( x Construct a KDE model for each meteorological element, with the kernel function formula as follows: K ( x ) =

[0017] in, The data consists of real-time meteorological observations of the wind farm to be verified. The bandwidth of the KDE model is determined using the following formula. :

[0018] Where σ is the sample standard deviation, IQR is the sample interquartile range, n is the sample size, and min is the minimum value function; Real-time observation data x The probability density of its historical normal distribution is calculated using the KDE model. f ( x ):

[0019] in, For the first i One historical normal sample; Set probability density threshold f th ,when f ( x )< f th When the weather is abnormal, it is marked as an anomaly; during periods of extreme weather, the probability density threshold is lowered. f th .

[0020] Furthermore, a hybrid classification model comprising a static feature extraction model, a temporal feature extraction model, and a dynamic weight fusion mechanism is employed to classify and determine suspicious data. The static feature extraction model is selected from LightGBM, XGBoost, or Random Forest; the temporal feature extraction model is selected from bidirectional LSTM, GRU, or Transformer; and the dynamic weight fusion mechanism is an attention mechanism. The first layer extracts static features using a static feature extraction model and outputs a static feature classification probability vector. ; The second layer extracts temporal features through a temporal feature extraction model and outputs a temporal feature classification probability vector. P Bi-LSTM ( k ); The third layer calculates dynamic weights and completes probability fusion through a dynamic weight fusion mechanism, specifically including: Calculate static feature attention score score static : score static =

[0021] in, The first learnable weight parameter, This represents the terrain complexity coefficient. =min(s / 30 ,1), where s is the slope of the wind measurement tower and H is the equipment health score; Calculate attention score for temporal features :

[0022] in, This is the second learnable weight parameter; The static feature attention scores and temporal feature attention scores are weighted and normalized using the Softmax function, and the final classification probability fusion formula is determined as follows:

[0023] in, For the final classification probability vector after fusion, take... As a preliminary classification label, The attention score is the normalized static feature. The attention score for the normalized temporal features; When the slope of the meteorological tower is greater than 30°, the tower should be raised. When the equipment health score H < 0.5, improve... .

[0024] Furthermore, differentiated repairs are performed on different category labels, including: For sensor aging data, first calculate the aging deviation coefficient. β The formula is: β =1-[0.3×min(service life / 8, 1)+0.4×min(calibration interval / 12, 1)+0.3×min(historical number of failures / 10, 1)] The service life is in years, the calibration interval is in months, and the historical failure count is the number of failures within the past year. β The value range is 0.5 to 1.0; The data is then corrected using a repair formula, which is: v 修复 = v 观测 × β +0.5×( v参考 - v 观测 ) in, v 修复 The wind speed after repair. v 观测 The actual observed wind speed, v 参考 The wind speed is the same as that at a nearby reference station. The temperature correction formula is: T 修复 = T 观测 +0.3×( T 参考 - T 观测 )×(1- β ) in, T 修复 For the restored temperature, T 观测 The actual observed temperature. T 参考 Temperatures were within the normal range for this time of year. For short-term missing data with a duration of ≤1 hour, linear interpolation and fusion with neighboring data are used for data restoration. The restoration formula is as follows: v 最终 = v 时序插值 ×(1- w )+ v 邻近参考站 × w in, v 最终 This is the final repair value for short-term missing data. v 时序插值 This is the result of linear interpolation of normal data before and after the period of missing data. w Dynamic weights for data from nearby reference meteorological towers. v 邻近参考站 This refers to the corresponding data from nearby reference stations during the same period; For transmission interference-related errors, multi-source weighted fusion interpolation is used for repair. The weight allocation rule is: data from different layers of the same tower (50%) > data from nearby reference wind towers (30%) > ERA5 reanalysis data (20%), and is dynamically adjusted according to the validity of the data source.

[0025] Furthermore, the full-process traceability information is stored using a data traceability chain in the form of a consortium blockchain constructed with blockchain technology. The traceability chain includes three types of consensus nodes: wind farm data center, operation and maintenance management platform, and third-party supervision agency. Each node synchronously stores a complete traceability copy containing the original data hash value, quality control cleaning operation log, operation entity ID, and timestamp. The node consensus mechanism ensures that the data is tamper-proof and tamper-proof. The full-process traceability information supports multi-dimensional combined queries by time range, wind measurement tower ID, anomaly type, cleaning status, and repair algorithm type. The invalid data removal process includes: generating an invalid data removal list for data whose error exceeds the standard after repair, long-term missing data that cannot be repaired, invalid data after aging repair, and worthless data that cannot be clearly classified. After manual review and confirmation, the invalid data is transferred to an invalid data isolation library for storage, while retaining the original data. At the same time, a removal mark is made at the corresponding position in the normal dataset to retain a complete traceability link.

[0026] Furthermore, after obtaining the standardized meteorological dataset after quality control, the method further includes: verifying the quality control effect based on the standardized meteorological dataset after quality control through four categories of wind farm indicators: data integrity, data accuracy, operational usability, and economic benefits; and performing dynamic iterations of incremental training, manual feedback optimization, and terrain and equipment adaptation updates based on the verification results, feedback data from the quality control process, and manual review results; wherein: The incremental training includes: updating the training data using a sliding time window every first cycle, retaining historical valid data, and performing incremental training on the hybrid classification model every second cycle; if extreme weather such as typhoons or thunderstorms occur during the second cycle, an additional incremental training is performed on the hybrid classification model. The manual feedback optimization includes: every second cycle (a set number of manually reviewed and corrected cases are fed back to the training set to optimize the parameters and feature weights of the hybrid classification model, requiring the monthly anomaly identification accuracy of the hybrid classification model to increase by more than or equal to a preset threshold). The terrain and equipment adaptation update includes: updating the wind farm DEM data every third cycle, recalibrating the terrain barrier coefficient and slope-aspect correction function; and automatically reading terrain and equipment parameters when adding a new wind measurement tower or replacing a sensor to complete the adaptation update of the hybrid classification model.

[0027] The method for identifying and intelligently cleaning up data anomalies in automatic weather stations for wind farms in complex terrain provided by this invention has at least the following beneficial effects: 1) This invention addresses the data quality control of automatic weather stations in wind farms with complex terrain. It constructs a closed-loop system covering the entire process from data preprocessing, anomaly identification, accurate judgment, quality control cleaning, to dynamic iteration for effect verification. This solves the technical problems of traditional methods, such as poor adaptability to complex terrain, insufficient accuracy in identifying extreme weather scenarios, and weak data processing capabilities due to sensor aging. It comprehensively improves the integrity, accuracy, and usability of meteorological data in wind farms, providing high-quality and standardized data support for the operation and maintenance of wind farm power prediction and wind energy resource assessment equipment.

[0028] 2) This invention pioneers a multi-terrain differentiated correction system. Targeting typical wind farm scenarios in three complex terrain types—plateau, mountain, and canyon—it fits differentiated slope and aspect correction functions based on wind tunnel experimental data. The correction coefficient is dynamically adjusted by combining the slope aspect with the prevailing wind direction angle, resolving the misjudgment of high wind speeds on steep slopes and data deviations on gentle slopes caused by traditional uniform correction standards. Simultaneously, it constructs spatial detection optimization rules related to terrain, sets a tiered reference station screening mechanism, and dynamically sets wind speed difference thresholds based on terrain barrier coefficients. This adapts to the large differences in wind fields between adjacent stations under complex terrain, significantly reducing the false negative rate of spatial anomaly identification and comprehensively improving the accuracy and scene adaptability of meteorological data preprocessing in complex terrain scenarios.

[0029] 3) This invention breaks through the traditional feature construction method that relies solely on the observation data itself. It constructs a multi-dimensional feature matrix covering basic statistical physical correlation, terrain adaptation, equipment health, and extreme weather. For the first time, it incorporates the sensor's service life, calibration cycle, and historical failure count into the feature system. By quantifying the equipment health status through weighted scoring, the model can accurately adapt to systematic deviation scenarios caused by sensor aging, significantly improving the recognition accuracy of aging data. Simultaneously, it designs exclusive features and instantaneous fault recognition rules for extreme weather scenarios such as typhoons and thunderstorms, filling the technical gap in traditional methods for processing extreme weather data. This significantly improves the recognition accuracy of abnormal data under extreme weather conditions, providing comprehensive and suitable feature inputs for subsequent anomaly classification and judgment, tailored to the wind farm business scenario.

[0030] 4) This invention constructs a two-level anomaly identification system combining multi-dimensional initial detection collaboration and precise judgment using a hybrid model. It employs a four-dimensional initial detection system—threshold and logical verification, temporal consistency detection, distribution offset detection, and extreme weather adaptation verification—to comprehensively cover various anomaly signals and avoid the missed detection problems of single detection methods. A hybrid classification model combining LightGBM bidirectional LSTM and an attention mechanism is used. The attention mechanism dynamically allocates weights for static and temporal features, adaptively matching the identification needs under different terrain complexities and equipment health states. This significantly improves the classification accuracy of anomaly data in complex scenarios while effectively reducing false positive and false negative rates, preventing the erroneous exclusion of real extreme wind energy data.

[0031] 5) This invention formulates differentiated intelligent quality control and cleaning strategies for different types of abnormal data, and matches suitable repair algorithms to erroneous data, missing data, and sensor aging data, significantly improving data repair accuracy and the proportion of effective data, ensuring the data integrity of core meteorological elements in wind farms meets business needs. Simultaneously, it constructs a full-process data traceability management system for quality control, using blockchain technology to build a multi-node data traceability chain, ensuring that the data throughout the quality control process is tamper-proof and tamper-proof, supporting multi-dimensional combined queries and visualization, providing complete and traceable data support for wind farm operation and maintenance optimization model iteration, supervision, auditing, and extreme weather review, enabling every step of the operation to be traceable and every piece of data to be verifiable.

[0032] 6) This invention establishes a closed-loop dynamic iterative mechanism for the entire process of quantitative verification, incremental training, feedback, and optimization. It comprehensively verifies the quality control effect through four wind farm-specific indicators: data integrity, accuracy, business practicality, and economic benefits. Through regular incremental training and manual feedback optimization, it continuously improves the model's adaptability and recognition accuracy, adaptively matching the iterative needs of wind farm terrain changes, equipment updates, and extreme weather scenarios. After implementation, the method can significantly reduce wind farm power prediction errors, improve the accuracy of power generation assessment and wind energy resource assessment, reduce unplanned downtime caused by sensor failures, lower equipment operation and maintenance costs, and achieve a dual improvement in the business value and economic benefits of wind farms. It provides a standardized, end-to-end solution for meteorological data quality control in wind farms with complex terrain. Attached Figure Description

[0033] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0034] Figure 1 This is a system framework diagram of an automatic weather station data anomaly identification and intelligent quality control cleaning method for wind farms in complex terrain, provided by an embodiment of the present invention. The data preprocessing module is used to execute step S10 (data preprocessing), the multi-dimensional preliminary inspection module is used to execute step S20 (multi-dimensional preliminary inspection), the hybrid model accurate judgment module is used to execute step S30 (hybrid model accurate judgment), the intelligent quality control cleaning module is used to execute step S40 (differentiated quality control cleaning), and the effect verification and dynamic iteration module is used to execute step S50 (effect verification and dynamic iteration). Figure 2 The flowchart illustrates a method for identifying and intelligently cleaning data anomalies at an automatic weather station in a wind farm with complex terrain, as provided in an embodiment of the present invention. The input layer is used to input a multidimensional feature matrix, the feature extraction layer is used to extract static and temporal features, the attention mechanism layer is used to assign feature weights, and the output layer is used to output the data anomaly determination result. Figure 3The data processing flowchart of the hybrid classification model provided in this embodiment of the invention is as follows: the first layer outputs a static feature classification probability vector through a static feature extraction model; the second layer outputs a temporal feature classification probability vector through a temporal feature extraction model; and the third layer calculates dynamic weights through an attention mechanism and fuses them to obtain the final classification probability vector. After the model makes its judgment, high-confidence results directly output the final classification label (0 / 1 / 2 / 4), while low-confidence results are marked as suspicious data in the intermediate state (label 3) and trigger manual review. After review, the data is reclassified to the final classification label.

[0035] The accompanying drawings have illustrated specific embodiments of the invention, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the invention in any way, but rather to illustrate the concept of the invention to those skilled in the art through reference to specific embodiments. Detailed Implementation

[0036] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0037] It should be noted that in the embodiments of the present invention, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of the present invention. However, they do not mean that the inventor has used or necessarily used the solution.

[0038] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.

[0039] This invention provides a method for anomaly identification and intelligent quality control cleaning of automatic weather station data in wind farms with complex terrain. Specifically, it is a method for multimodal anomaly identification, machine learning-driven precise quality control, and adaptive data cleaning of observation data from automatic weather stations in wind farms with complex terrains such as mountains, canyons, and plateaus. This method is applicable to the entire lifecycle of wind farms, including early-stage wind energy resource assessment, mid-stage power prediction optimization, and late-stage equipment operation and maintenance monitoring. In particular, it can solve problems such as low accuracy in anomaly identification, inconsistent quality control standards, and poor timeliness of data cleaning for meteorological elements such as wind speed, wind direction, and temperature in complex terrains, providing high-quality meteorological data support for the efficient operation of wind farms.

[0040] like Figure 1 The diagram shown is a system framework diagram of an automatic weather station data anomaly identification and intelligent quality control cleaning method for wind farms in complex terrain, provided by an embodiment of the present invention. The system includes a data preprocessing module, a multi-dimensional anomaly identification module (including an extreme weather adaptation sub-module), a machine learning precision judgment module, an intelligent quality control cleaning module, and an effect verification and iteration module. These five modules are respectively configured to perform the following... Figure 2 Steps S10-S50 in the method shown.

[0041] It should be noted that in some embodiments, step S50 is not necessary. That is, the method of the present invention can still achieve automatic weather station data anomaly identification and intelligent quality control cleaning without executing step S50. Executing step S50 is a preferred embodiment of the method of the present invention.

[0042] To facilitate a thorough understanding of the technical solution of the present invention by those skilled in the art, this embodiment describes the meteorological data used in the method of the present invention.

[0043] The meteorological data used in this embodiment is designed around the needs of the entire life cycle of wind farm wind energy assessment, power prediction and operation and maintenance monitoring. Specifically, it includes four categories: automatic weather station observation data, wind farm geographic information data, wind farm equipment operation and maintenance data, and wind farm extreme weather and operation data, which are different from the general meteorological station data system.

[0044] Automatic weather station observation data includes timestamps (accurate to the second, matching the time synchronization requirements of unit SCADA data), wind speed ( v Includes 10-minute average wind speed, instantaneous wind speed, unit m / s, core for power curve fitting), and wind direction ( d (unit: °C) Temperature ( T (unit: °C) and air pressure ( P (unit: hPa) and relative humidity ( U (unit: %); This type of data directly reflects the wind energy characteristics of the wind farm and the operating environment of the unit, and is the core foundation for anomaly identification and quality control cleaning; Geographic information data for wind farms includes the latitude and longitude (Lat, Lon) and altitude of the meteorological tower. h (Unit: m) Wind farm digital elevation model (DEM) data (resolution 30m×30m, used to extract slope, aspect and terrain barrier coefficient, to adapt to the micro-site selection requirements of wind farms), relative positional relationship between the wind measurement tower and the unit (e.g., 500m upwind and 300m downwind, used to determine the impact of wake interference). Wind farm equipment operation and maintenance data records include: wind measurement tower sensor model (e.g., ultrasonic anemometer, mechanical wind vane), installation height (matching the turbine hub height, commonly 70m / 80m / 100m), calibration records (every 6 months, related to sensor sensitivity attenuation assessment), fault repair log (including wind farm-specific fault types, such as wind direction sensor jamming caused by wake impact, sensor loosening caused by strong wind), and turbine operating status (e.g., shutdown / grid connection, used to correlate data anomalies and turbine interference). Extreme weather and operational data for wind farms include early warning information for extreme weather such as typhoons and thunderstorms (warning level, period of impact, expected wind speed range, and determination of wind speed cut-off for associated units), historical extreme wind energy data (such as maximum gust wind speed and duration of continuous strong winds, used to distinguish between real extreme data and fault data), and unit power curves (used to verify the power prediction adaptability of the cleaned data).

[0045] To ensure the consistency and accuracy of data processing, this embodiment performs time resolution standardization on all input data, unifying it to the 1-minute level. The specific processing method is as follows: For high-frequency data (such as 10-second data), a moving average downsampling method is used, with a 1-minute time window. The average value of the data within the window is calculated as the observation value for that minute. This preserves the main trend of data change and avoids the impact of high-frequency data redundancy on processing efficiency. For low-frequency data (such as 10-minute intervals), a linear interpolation upsampling method is used. Based on the values ​​of two adjacent 10-minute data points, the observation value of each 1-minute interval is calculated by linear interpolation to fill in the time interval, so that the data can meet the time resolution requirements of subsequent time series analysis and real-time quality control.

[0046] During the initial data collection process, data omissions are prone to occur due to factors such as complex terrain, communication link stability, equipment operating status, and extreme weather. These omissions mainly manifest in the following two categories: Complete missing data: No data was collected in the message, meaning that all meteorological data for a certain period of time is missing. Common causes include communication interruption (such as signal transmission being blocked by terrain), power failure of equipment, and complete sensor failure. This situation is more common in wind farms with poor communication conditions, such as mountainous areas and canyons. The duration of complete missing data in a single day may reach 2 to 4 hours. Partial data missing: The message only collected some meteorological element data, while other meteorological element data were missing. For example, only wind speed and temperature data were obtained, while wind direction, air pressure, and relative humidity data were missing. The main reasons include some sensor failures, loss of some fields during data transmission, and interference from extreme weather with some sensor data collection.

[0047] Missing or underreported data can lead to incomplete data sequences, affecting the accuracy of subsequent anomaly identification and reducing data availability. If not effectively addressed, this will further impact the reliability of services such as wind farm power forecasting and power generation assessment.

[0048] To address the aforementioned data gaps and omissions, this embodiment employs the following methods during the data preprocessing stage for preliminary processing, laying the foundation for subsequent anomaly identification and quality control cleaning: Missing data segments are marked according to their consecutive duration. For example, T5 represents 5 consecutive minutes of missing data and T30 represents 30 consecutive minutes of missing data. The time range and duration of missing data are clearly recorded, which provides a basis for developing differentiated repair strategies based on the duration of missing data. Preliminary screening: Remove data segments that are missing, making basic physical quantity calculations or logical verification impossible, such as data within a certain time period. v、T、P When all data is missing and there is no data from adjacent reference stations to assist in the judgment, the data segment is marked as pending further evaluation to avoid invalid data entering the subsequent processing flow and improve overall processing efficiency.

[0049] The method first executes step S10 through a data preprocessing module. Step S10 specifically includes: acquiring multi-source data and performing standardization processing such as unified time resolution, missing data labeling, and initial screening of invalid data; performing physical quantity transformation on the obtained standardized time-series data; and constructing a multi-dimensional feature matrix. The multi-source data includes wind farm automatic weather station observation data, geographic information data, equipment operation and maintenance data, and extreme weather and operation data. The multi-dimensional feature matrix includes basic statistical features, cross-parameter physical correlation features, terrain adaptation features, equipment health features, and time and extreme weather features.

[0050] In some embodiments, step S10 is specifically implemented through the following steps S101-S102.

[0051] S101; Data cleaning and physical quantity conversion.

[0052] In this embodiment, the data cleaning method is as follows: invalid characters (such as "NULL", "fault" and other non-numerical identifiers) and duplicate records (data whose timestamps are completely consistent with the wind tower ID) are removed from the original data. The 3σ criterion is used to initially filter obvious outliers (such as wind speed > 100m / s, temperature > 60℃), and repairable suspected abnormal data are retained.

[0053] Physical quantity conversion includes wind direction angle conversion, air pressure and altitude correction, and humidity unit conversion.

[0054] Wind direction angle conversion is the process of changing the wind direction. d (0°~360°) converted to Cartesian coordinate system components: u=sin( d × π / 180), v =cos( d ×π / 180) is used to eliminate periodic angular deviations, facilitating subsequent calculations of wind field stability characteristics, such as the ratio R of wind direction standard deviation to wind speed. dv ; where u and v For Cartesian coordinate system components.

[0055] Barometric altitude correction adjusts the barometric pressure data from each meteorological tower to sea-level pressure to facilitate cross-wind farm comparisons and wind energy resource assessments. The correction formula is as follows:

[0056] in, To observe air pressure, h The altitude of the wind measurement tower, T 0 represents the sea-level reference temperature (288.15 K). c Let γ be the vertical temperature lapse rate (γ = 0.0065 K / m), and M be the molar mass of dry air. R This is the universal gas constant. It is the acceleration due to gravity. T 0=288.15K, γ= 0.0065K / m, g=9.80665m / s 2 M = 0.0289644 kg / mol R =8.31432 J / (mol・K)).

[0057] Humidity unit conversion is the conversion of relative humidity to specific humidity. q : q =0.622× e / ( P -0.378 e ) in P This refers to the ambient atmospheric pressure measured on-site by the wind measurement tower. eFor water vapor pressure, e pass U and T Derivation of the Magnus formula enhances the physical correlation between humidity and temperature; the Magnus formula is expressed as: e = e 0×exp[ a ×( T - T 0) / ( T - b )] in, e 0 = 0.61078 hPa a =17.27、 b =35.86℃).

[0058] S102; Construct the feature matrix.

[0059] This embodiment constructs a multi-dimensional feature matrix containing five types of wind farm-specific features. The dimension is the number of samples × the number of features. All features are designed around the accuracy of wind farm anomaly identification and the adaptability of power prediction. Specifically, they include basic statistical features adapted to the fluctuation characteristics of wind farm data, cross-parameter physical correlation features, terrain adaptation features, equipment health features, and time and extreme weather features.

[0060] When calculating basic statistical characteristics, an adaptive window size rule is applied during data sampling, for example... v , d A 10-minute window was used to capture turbulent fluctuations. T , P , U A 1-hour window is used to accommodate slow changes.

[0061] Basic statistical characteristics include mean wind speed, standard deviation of wind speed, extreme values ​​and ranges, and rate of change of wind speed.

[0062] The formula for calculating the average wind speed is: m =(1 / W )×Σ( x i ), i =1~ W , W The window size is used; the average wind speed can be used for effective wind time statistics in wind energy resource assessment.

[0063] Wind speed standard deviation s Used for turbulence intensity I = s / m Calculations show that if the turbulence intensity exceeds 0.25, the unit power needs to be adjusted. s = .

[0064] Extreme values ​​and ranges include maximum wind speed. x max Minimum value x min and the extreme temperature difference R= x max - x min Maximum wind speed x max Used for verifying the cut-off wind speed of the unit. For example, wind speeds exceeding 25 m / s need to be determined to determine whether they are true extreme winds, and extreme temperature differences are used for assessing the risk of blade icing.

[0065] Wind speed change rate k =| x ( t )- x ( t -1)| / Δ t , where Δ t For sampling intervals, Δ is used for 1-minute intervals. t =60s, reflecting the degree of data change; when the wind speed change rate exceeds 2m / s·min, it is necessary to be alert to the unit load exceeding the limit and mark it as suspicious data.

[0066] Cross-parameter physical correlation characteristics include wind-pressure correlation characteristics, temperature-humidity-density correlation characteristics, and wind field stability characteristics.

[0067] The wind-pressure correlation characteristic can be calculated by: calculating the pressure gradient Δ between adjacent meteorological towers. P / Δ h (Unit: hPa / m) and wind speed v Pearson coefficient r The calculation formula is: r = Cov (Δ P / Δ h , v ) / [ s (Δ P / Δ h )× s ( v )] in Cov (Δ P / Δ h , v ) represents the covariance between the pressure gradient at altitude and the wind speed. s (Δ P / Δ h ) represents the standard deviation of the barometric altitude gradient. s (v) represents the standard deviation of wind speed. Under normal circumstances, within a wind farm... rA pressure gradient ≥0.3 indicates a higher wind speed, consistent with the funneling effect in wind farms. r A value less than 0.2 indicates abnormal data and requires investigation of power prediction errors.

[0068] Temperature-humidity-density correlation characteristics through air temperature T relative humidity U Calculate air density r The formula for calculating air density is: r =1.293×(273.15 / (273.15+ T ))×( P sea / 1013.25)×(1-0.378× U / 100) in P sea The sea level pressure is then calculated after altitude correction. r With wind speed v The product of, i.e. r × v 3 As a core parameter of wind energy density, if this product exceeds the historical normal range of the wind farm, for example, exceeding 500W / m², it will be considered a significant factor. 2 Marked as suspicious data; Wind field stability characteristics include wind direction standard deviation s d With wind speed v The ratio R dv = s d / v Under normal circumstances, R dv ≤0.5, if R dv A value ≥0.5 indicates that the wind direction sensor is stuck or there is abnormal turbulence.

[0069] Regarding terrain adaptation features, this embodiment defines the coefficient rules of the slope-aspect correction function f(s,a) for plateau, mountain, and canyon terrains commonly found in wind farms, as shown in Table 1 below.

[0070] Table 1. Applicable Terrain Range for Adjustment Rules of Corrected Function Coefficients ; Terrain-corrected wind speed: v '= v / f ( s , a ),in s The slope (°) of the location of the wind measurement tower corresponds to the included angle in Table 1. a Slope aspect (°), f( s, a ) is the slope-aspect correction function. v The corrected wind speed is shown in Table 1 above, and the specific calculation rules are as follows. The corrected wind speed is used for wind farm energy resource assessment (such as calculating annual theoretical power generation) to avoid wind speed deviations caused by terrain.

[0071] Terrain adaptation features also include spatial correlation features, adapting to the layout of wind farm meteorological towers. The calculation method can be: by calculating the wind speed difference Δ between the target meteorological tower and reference stations within a 3km radius. v =| v target - v ref |, among which v target The measured wind speed values ​​of the target anemometer tower to be verified during the same period. v ref To reference the concurrently measured wind speed values ​​from the wind measurement tower, and in conjunction with the terrain barrier coefficient... k block (No obstruction 1.0, with mountain obstruction 0.6, with unit wake obstruction 0.7), set threshold Δ v max =2× k block ×(1+ s / 30), the greater the slope and the stronger the wake, the greater the allowable wind speed difference; when the reference station is abnormal, it will automatically expand to 5km (typical spacing of wind farm meteorological towers) to ensure the validity of spatial correlation data.

[0072] Equipment health characteristics include sensor lifespan, calibration cycle, and historical fault characteristics. The sensor's service life characteristics are obtained by dividing the sensor's service life into 5 intervals (0~1 year, 1~3 years, 3~5 years, 5~8 years, >8 years), and assigning values ​​of 1.0, 0.9, 0.8, 0.6, and 0.4 respectively. The longer the service life, the lower the value assigned. The average life of wind farm sensors is 5~8 years. Sensitivity decay needs to be closely monitored after 5 years.

[0073] The calibration cycle characteristics can be obtained as follows: 1.0 for the interval between the current time and the last calibration ≤ 3 months, 0.8 for 3~6 months, 0.6 for 6~12 months, and 0.4 for > 12 months. The calibration cycle for wind farm sensors is usually 6 months. Exceeding this period can easily lead to data deviation.

[0074] Historical fault characteristics can be obtained as follows: 0 faults in the past year are assigned a value of 1.0, 1-2 faults are assigned a value of 0.8, 3-5 faults are assigned a value of 0.6, and >5 faults are assigned a value of 0.4. Sensors with a high number of faults have low data reliability and require enhanced quality control.

[0075] The time and extreme weather characteristics include hourly characteristics, seasonal characteristics, and extreme weather characteristics.

[0076] The hourly segment characteristics can be obtained by dividing a day into 24-hour segments (0~23) and using one-hot encoding, focusing on distinguishing between "high load periods" (such as the period corresponding to wind speeds of 8~18m / s) and low load periods. Abnormal data during high load periods have a greater impact on power. Seasonal characteristics can be obtained by dividing them into four categories according to the season (1=spring, 2=summer, 3=autumn, 4=winter), using unique thermal coding, combined with the monsoon patterns of wind farms (such as southeast monsoon in summer and northwest monsoon in winter), and the climatological thresholds of different seasons (such as the wind speed threshold being higher in winter than in summer). Extreme weather characteristics can be obtained by marking the typhoon / thunderstorm warning level (level 1-4) and the period of impact, calculating the difference between the real-time wind speed and the wind speed cut off by the turbine, for example, if the wind speed exceeds the cut-off speed by 25 m / s, it is necessary to determine whether it is a real extreme wind and link it to the wind farm shutdown plan.

[0077] Based on step S10, step S20 is further executed through the multi-dimensional abnormal signal identification module. Step S20 specifically includes: based on the standardized time series data and multi-dimensional feature matrix, performing at least one of the following preliminary checks: threshold and logic verification, time series consistency detection, distribution offset detection, and extreme weather adaptation verification; marking data that triggers any preliminary check abnormal condition as suspicious data; and recording the abnormal type through binary encoding.

[0078] In some embodiments, the multi-dimensional abnormal signal identification module adopts a four-dimensional preliminary detection of threshold verification, time series analysis, distribution detection and extreme weather adaptation to achieve comprehensive screening of abnormal signals. The specific steps are as follows: steps S201-S205.

[0079] S201: Execution threshold and logic verification, including the following three aspects.

[0080] Firstly, a physical limit threshold is set, which is based on the sensor's measurement range and cannot be exceeded. For example, the physical limit threshold is set as follows: the wind speed threshold is... v ≤ 70m / s (upper limit of mainstream wind speed sensor range); Temperature threshold is -40℃ ≤ T ≤ 50℃ (industrial-grade temperature and humidity sensor range); air pressure threshold is 500hPa ≤ P ≤ 1100hPa (covering high-altitude to low-altitude scenarios); humidity threshold is 0% ≤ H ≤ 100%.

[0081] Secondly, climatological thresholds are set seasonally based on normal data from wind farms over the past three years, and these thresholds are used to identify suspicious data. For example, the climatological thresholds set using the 95th percentile are as follows: Summer (June to August): Wind speed v A speed >30m / s is suspicious; Winter (December to February): Wind speed v >35m / s is suspicious; Spring / Autumn (March-May, September-November): Wind Speed v A speed >32m / s is suspicious.

[0082] Thirdly, logical conflict rules are designed to identify suspicious data, specifically as follows: Wind speed v When the wind speed is 0 m / s, the wind direction data is invalid and marked as suspicious; humidity U <0%RH or U When RH is >105%, the humidity sensor is considered faulty; sea level air pressure P sea The difference between the pressure and the neighboring station is >5 hPa, indicating an abnormality in the barometric pressure sensor; the daily temperature range (daily high temperature - daily low temperature) is >25℃, which exceeds the normal daily temperature range for mountainous areas, raising suspicion.

[0083] It should be noted that the three aspects mentioned above are not in any particular order during actual execution. They can be executed in parallel or in a set order, such as executing the first aspect first, and then the second and third aspects.

[0084] S202: Perform timing consistency check.

[0085] This embodiment constructs a VAE (Variational Autoencoder) architecture, which includes an encoder and a decoder. The encoder consists of two LSTM layers and a fully connected layer. The LSTM has 64 hidden units, and the fully connected layer outputs a mean value μ. z and variance σ z The latent space dimension is 16; the decoder consists of two LSTM layers and a fully connected layer; the LSTM of the decoder has 64 hidden units, and the fully connected layer outputs a reconstructed sequence with a dimension of 60. The input data for this VAE architecture is preprocessed temporal data, such as... v、T、U The 1-minute time series has a window size of 60, which is equivalent to 1 hour of data.

[0086] During training, normal data from wind farms over the past year (excluding manually labeled outliers) were used as training data. The training epochs were 100, the batch size was 32, the optimizer was Adam, and the learning rate was 0.001.

[0087] The anomaly detection method based on the trained VAE architecture is as follows: Calculate the mean squared error (MSE) between the real-time data sequence and the VAE reconstructed sequence: MSE=(1 / W )×Σ( x real - x rec ) 2 in, W To adjust the sliding window size, x real For real-time data sequences, x rec Reconstruct the sequence for VAE.

[0088] Set the abnormal threshold MSE th =3×μ MSE μ MSE Let MSE be the mean of the training set. If MSE > MSE th It is marked as a timing anomaly.

[0089] S203: Construct a kernel density estimation KDE model to perform distribution shift detection.

[0090] Based on normal data from the past year (categorized by element, such as...) v、T、U Data sources were selected by removing labeled outliers and data segments with missing durations exceeding 30 minutes. A kernel density estimation (KDE) model was constructed, with its Gaussian kernel function... K ( x ) = It possesses the characteristics of continuous differentiability and strong symmetry, making it suitable for the continuous distribution characteristics of meteorological data; Bandwidth selection: The Silverman rule is used to ensure that the model can capture the details of the data distribution while avoiding overfitting. The Silverman rule is expressed as follows:

[0091] Where σ is the sample standard deviation, IQR is the sample interquartile range (the difference between the upper and lower quartiles), n is the sample size, and min is the minimum value function.

[0092] Model training: KDE models are trained separately for each meteorological element. For example, wind speed data is used for training to obtain a wind speed KDE model. Each model independently stores the distribution parameters of the corresponding element, which can be called in real time later.

[0093] The anomaly detection method based on the trained KDE model is as follows: Real-time observation data x The probability density of its historical normal distribution is calculated using the KDE model. ,in x i This is a historical normal sample.

[0094] Take the probability density threshold f th =0.001 corresponds to an extremely low probability event, meaning that only 0.1% of the samples in historical normal data have a probability density lower than this value. f ( x )< f th Mark as distribution anomalies; during extreme weather periods, such as typhoons and low winter temperatures, simultaneously lower the threshold values ​​of corresponding elements: wind speed. f th =0.0005, Temperature f th =0.0005, to avoid misjudgment based on real extreme data; In particular, for low temperatures in winter (such as T <-20℃), strong summer gusts (such as v In extreme but reasonable situations such as wind speeds exceeding 40 m / s, thresholds are calibrated using historical extreme weather records to avoid misjudgments (e.g., temporarily lowering the KDE model threshold for winter temperatures). f th =0.0005).

[0095] S204: For extreme weather scenarios such as typhoons and thunderstorms, the extreme weather adaptation submodule supplements the instantaneous fault identification rules and data processing strategies, as detailed below: (1) Typhoon scene adaptation.

[0096] Enhanced data collection: During the typhoon warning period (from the issuance of the warning to 2 hours after the warning is lifted), the frequency of wind speed and wind direction data collection will be increased to the level of 1 second, while other elements will be kept at the level of 1 minute to capture instantaneous changes in wind speed; Anomaly identification rules: If the wind speed changes by more than 30 m / s within 30 seconds, and is accompanied by a sudden drop in air pressure of more than 8 hPa within 5 minutes, it is judged as an anomaly in the instantaneous impact of a typhoon; if the wind direction changes by more than 180° within 1 minute, and there is no obvious terrain obstruction, it is judged as an anomaly in the typhoon wind direction. Data processing strategy: For abnormal data on instantaneous impact of typhoons, the typhoon impact coefficient correction method is used for repair. The correction coefficient k = 1 - (real-time wind speed - historical average wind speed of typhoons in the same period) / historical maximum wind speed of typhoons in the same period, to ensure that the repaired value conforms to the actual impact pattern of typhoons; for abnormal data with disordered wind direction, it is marked as pending review and further judgment is made in combination with the typhoon's movement path.

[0097] (2) Thunderstorm scene adaptation.

[0098] Anomaly identification rules: If the wind speed changes by more than 20 m / s within 10 seconds, and is accompanied by a sudden increase in humidity of >20%RH within 1 minute, it is judged as an anomaly of instantaneous interference from lightning; if the air pressure rises by >5 hPa within 2 minutes and then drops rapidly, and there is no other weather system influence, it is judged as an anomaly of thunderstorm pressure pulse. Data processing strategy: For abnormal data of instantaneous interference from lightning strikes, wavelet threshold denoising combined with mean calibration during thunderstorm periods is used for repair. First, high-frequency interference is removed using the db4 wavelet basis, and then the mean of the normal data in the first 10 minutes of the thunderstorm period is used to calibrate and repair the data. For abnormal data of thunderstorm pressure pulses, linear interpolation of the adjacent 5 minutes of normal data is used for repair.

[0099] S205: Initial inspection results integration and dynamic threshold refresh.

[0100] The initial inspection results integration specifically includes: performing an OR operation on four types of results: threshold verification, temporal consistency detection, distribution offset detection, and extreme weather adaptation. If any one of the detection conditions is met, it is marked as suspicious data, and the suspicious type is recorded through binary encoding. For example, 1001 indicates threshold abnormality + thunderstorm and lightning strike abnormality, and 0110 indicates temporal abnormality + distribution abnormality, which facilitates subsequent accurate judgment module calls.

[0101] The dynamic threshold refresh mechanism includes periodic refresh, emergency adjustment, and terrain adaptation adjustment.

[0102] Regularly updated: Historical samples from the past year are updated quarterly, and climatological thresholds (such as summer) are recalculated. v Suspicion threshold and KDE probability density threshold are used to adapt to the distribution shift of meteorological data caused by seasonal changes. Emergency Adjustment: If new extreme event samples are added within the past 30 days (e.g.) v >40m / s T If the number of cases (<-30℃) exceeds 5% of the historical average for the same period, the threshold for the corresponding season and element will be automatically raised by 10% (e.g., winter). v The suspicious threshold was adjusted from 35 m / s to 38.5 m / s, and the reason for the adjustment was recorded. The threshold will be automatically rolled back in the next quarter. Terrain adaptation adjustment: When a new wind measurement tower is added to the wind farm, its terrain parameters (slope, aspect) are read, and the terrain correction function is recalculated. f (s, a), and update the threshold of the corresponding wind measurement tower (e.g., steep slope wind measurement tower). v The physical limit threshold is adjusted from 70m / s to 75m / s; the threshold priority of the extreme weather adaptation submodule is higher than the periodically refreshed threshold. That is, during the typhoon / thunderstorm warning period, the scenario-based threshold of step S204 is used first. For example, if the typhoon wind speed jumps to 30m / s, it will automatically return to the dynamic refresh threshold after the warning is lifted.

[0103] Based on step S20, step S30 is further executed through the machine learning precision judgment module. Step S30 specifically includes: taking the standardized time series data, multidimensional feature matrix and suspicious data as input, using a hybrid classification model including LightGBM, bidirectional LSTM and attention mechanism to classify and judge the suspicious data, directly outputting the corresponding classification label for high confidence results with the maximum probability of the classification label greater than or equal to a set threshold (the set threshold is preferably 0.9, and the value range is 0.85~0.95), and triggering manual review for low confidence results with all label probabilities less than the set threshold. The final classification label includes normal data, erroneous data, missing data and sensor aging data. Low confidence results are marked as suspicious data in the intermediate state (label 3) and trigger manual review.

[0104] In some embodiments, please combine Figure 3 The hybrid classification model shown is used to classify and determine the suspicious data marked in the initial inspection, so as to solve the problem of confusion between real extreme data and erroneous data. Specifically, it includes the following steps S301-S304.

[0105] S301: Construct the training dataset.

[0106] The training dataset used to train the hybrid classification model includes basic data, quantitative auxiliary data, and labeled data.

[0107] For example, the basic data comes from the wind farm's historical observation data over the past 3 years (including 1-minute data). v , d、T、P、U The data must include complete timestamps, terrain parameters, equipment operation and maintenance data (including service life and calibration cycle), and extreme weather warning data; quantitative auxiliary data should replace traditional weather phenomenon data, and pressure gradient force (ΔP / Δh) and turbulence intensity should be introduced. Difference in elements between adjacent meteorological towers (wind speed difference Δ) v Temperature difference Δ T The equipment health score (calculated by weighting the years of use, calibration cycle, and number of historical failures, with weights of 0.4, 0.3, and 0.3 respectively) improves data reliability. The labeled data is manually annotated by maintenance personnel in combination with equipment failure logs (such as sensor replacement records), communication interruption records (such as operator failure notifications), on-site inspection records, and extreme weather impact records to ensure the accuracy of the labels.

[0108] In this embodiment, the labeling system used for the labeled data includes four types of final classification labels and one type of intermediate state labels before manual review, as detailed below: Label 0 (Normal Data, Final Classification Label): Data that conforms to physical laws, temporal consistency, and terrain adaptation characteristics, without any anomaly markers, including real extreme data such as strong gusts in canyons, reasonably high wind speeds during typhoons, and extreme low temperatures in winter, which are consistent with the patterns of terrain and extreme weather. v Located at 5~15m / s and d As the terrain changes, or v Actual gust data for winds between 40 and 75 m / s (after terrain correction) during typhoon warning periods.

[0109] Tag 1 (Error Data, Final Category Tag): Data anomalies caused by sensor malfunction, transmission interference, or environmental interference, such as low readings due to icing of the wind speed sensor, random pulse values ​​caused by communication interruption, and wind speed jumps caused by electromagnetic interference.

[0110] Label 2 (Missing Data, Final Classification Label): Data is completely missing (e.g., no records for any element within a certain time period) or key elements are missing (e.g., ... v Data that is missing but other elements are normal, and cannot be temporarily supplemented by interpolation; Label 3 (Suspicious Data, Intermediate State Label): Intermediate state data determined by the model with low confidence. It is not the final classification label and is only used for samples whose probability of all final classification labels is less than the set threshold. After manual review, it needs to be reclassified to label 0 / 1 / 2 / 4 and does not participate in the subsequent quality control cleaning process.

[0111] Label 4 (deviation data caused by sensor aging, final classification label): The sensor has been used for more than 3 years or the calibration interval is more than 6 months. The data has systematic deviations, such as wind speed being 5% to 10% lower for a long time or temperature reading being 1 to 2°C higher for a long time, but it has not reached the fault level.

[0112] The data augmentation strategies employed in this embodiment when constructing the training dataset include temporal augmentation, terrain stratified sampling, device health stratified sampling, and class balancing.

[0113] Time series augmentation: Time reversal is applied to normal data sequences (e.g., reversing the time of day). v (Sequence reversal simulates afternoon symmetrical trend), additive Gaussian noise (signal-to-noise ratio 10:1, noise intensity ≤0.5m / s) v ), ≤0.3℃ ( T New samples are generated, with a focus on supplementing real extreme data samples (such as simulated strong gusts in canyons and wind speed sequences during typhoons) to solve the problem of scarce extreme data samples. Terrain-based stratified sampling: Data is divided into strata based on slope (<15° / 15°~30° / >30°), aspect (sunny / shady slope), and terrain type (plateau / mountain / canyon), ensuring that the sample proportion deviation of each terrain scene is ≤5% (e.g., the proportion of steep slope samples is not less than 20% of the total sample), and the proportion of real extreme data samples in each stratum is ≥10%. Stratified sampling of equipment health: Stratified sampling is conducted based on sensor service life (0~1 year / 1~3 years / 3~5 years / 5~8 years / >8 years) and calibration cycle (≤3 months / 3~6 months / 6~12 months / >12 months) to ensure that the proportion of samples in sensor aging scenarios is not less than 15% of the total sample. Category balancing: The SMOTE algorithm is used to expand minority class samples such as label 2 (missing data) and label 4 (sensor aging data) to make the ratio of the number of samples of the four label classes reach 1:1:1:1, thus avoiding logical contradictions.

[0114] S302: Construct a hybrid classification model architecture.

[0115] For the initially flagged suspicious data, a hybrid model combining LightGBM, bidirectional LSTM, and attention mechanism is used for accurate classification, distinguishing between real extreme data and erroneous data. This model also adapts to sensor aging scenarios. The hybrid classification model architecture has three layers, and the specific data processing steps for each layer are as follows: First layer: LightGBM static feature extraction.

[0116] Input: The basic statistical features constructed in step S10 (mean, standard deviation, rate of change, etc., a total of 8 dimensions), cross-parameter physical correlation features (wind-pressure correlation coefficient, wind field stability features, etc., a total of 6 dimensions), terrain adaptation features (terrain-corrected wind speed, spatial correlation difference, etc., a total of 7 dimensions), equipment health features (service life, calibration cycle, historical fault features, etc., a total of 3 dimensions), and time and extreme weather features (hourly segment, season, extreme weather warning level, etc., a total of 5 dimensions), totaling 29 static features; Model parameters: learning rate 0.05, tree depth 8, number of leaf nodes 32, regularization coefficient. L 1 = 0.1 L 2=0.2, adopting the gradient boosting decision tree algorithm, supporting feature importance evaluation (such as terrain-corrected wind speed feature importance weight ≥0.2, equipment health feature importance weight ≥0.15); Output: Static feature classification probability vector k=0,1,2,3,4 corresponds to 5 types of labels, with probability values ​​ranging from [0,1], and satisfies .

[0117] Second layer: Bidirectional LSTM for temporal feature capture.

[0118] Input: Preprocessed 1-minute time series data ( v、T、U The data consists of 30 consecutive time-step sequences (30×3 dimensions), covering the trend of changes within 15 minutes before and after the data. For extreme weather periods, an additional 10 time-step sequences (10×2 dimensions) of wind speed and direction at the 1-second level are input to enhance the temporal features. The network structure includes a forward LSTM layer, a backward LSTM layer, a feature fusion layer, and an output layer.

[0119] The forward LSTM layer has 64 hidden units, an activation function of tanh, and a dropout rate of 0.2 to prevent overfitting and capture past-to-current temporal dependencies, such as the upward trend of wind speed. The backward LSTM layer has 64 hidden units, uses the tanh activation function, and has a dropout rate of 0.2 to capture current-to-future temporal dependencies, such as... v An impending downward trend; The feature fusion layer concatenates the forward and backward LSTM outputs into a 128-dimensional feature vector. If it contains 1-second-level extreme weather data, it is compressed to 32 dimensions through a fully connected layer and then fused with the 128-dimensional vector to obtain a 160-dimensional feature vector, covering bidirectional time information and extreme weather time series features. The output layer uses a softmax activation function to map 160-dimensional features into a 5-dimensional temporal classification probability vector P. Bi-LSTM ( k ), which satisfies the probability normalization condition.

[0120] The third layer: Attention mechanism terrain weight optimization, which is implemented through the following steps: 1) Terrain feature input: Take the slope-aspect correction function calculated in the above steps. f ( s , a ), and introduce the terrain complexity coefficient τ=min( s / 30°, 1) ( s For the slope, τ∈[0, 1], the larger the slope, the closer τ is to 1, which means the terrain has a stronger influence on the data, and the equipment health score H (calculated from the years of use, calibration cycle, and historical fault characteristics, H∈[0,1]). 2) Weight parameter learning: Optimize feature weights through model training. (static feature weight parameters) and (Temporal feature weight parameters), all initial values ​​are set to 1.0, and are dynamically adjusted according to the terrain scene samples during training; 3) Attention score calculation: score static = The more complex the terrain ( The larger the value of H, the lower the equipment health (the smaller the value of H), and the higher the proportion of static features (including terrain adaptation and equipment health features) scores. The simpler the terrain ( The smaller the value (the higher the H value), the higher the device health (the larger the H value), and the higher the weight of the time feature score. 4) Weight normalization: The weights are normalized using the Softmax function to ensure that the sum of the weights is 1. , ; 5) Implement terrain and equipment reinforcement rules: When τ=1 (slope > 30°), An additional 20% increase (i.e.) =1.2× When H < 0.5 (low equipment health), An additional 15% increase (i.e.) =1.15× This study aimed to strengthen the influence of terrain features and equipment health features on the classification results.

[0121] 6) Final probability fusion: ; in, For the final classification probability vector after fusion, take... As a preliminary classification label.

[0122] S303: Model Training and Optimization.

[0123] In this embodiment, the data is segmented according to time series (to avoid future data leakage): training set (data from the first two years, accounting for 70%), validation set (data from the first half of the third year, accounting for 20%), and test set (data from the second half of the third year, accounting for 10%). A weighted cross-entropy loss function is used, assigning a weight of 2.0 to label 1 (erroneous data) and label 4 (sensor aging data), and an additional sampling weight of 1.5 (achieved through oversampling) to the true extreme data subcategory in label 0. The weights for normal data in label 0, label 2 (missing data), and label 3 (suspicious data) are all 1.0. The formula is as follows:

[0124] in, This is the one-hot encoded value for the tag (e.g., tag 1 corresponds to y=[0,1,0,0,0]). The category weights are used; in label 0, the real extreme data is determined by wind speed ≤75m / s and KDE probability density ≥0.0005 after terrain correction. During training, an oversampling strategy (sampling weight 1.5) is used to improve the model's ability to identify this type of data.

[0125] The AdamW optimizer (weight decay of 0.01) is used, and the learning rate is adjusted according to the cosine annealing strategy (initially 0.001, decaying by 10% every 10 rounds). The training run consists of 50 rounds. When the F1-score on the validation set shows no improvement for 5 consecutive rounds, an early stopping mechanism is triggered to save the optimal model parameters. At the same time, an oversampling strategy is used for real extreme data samples during training, for example, the sampling weight is set to 1.5 to improve the model's ability to recognize this type of data.

[0126] The model evaluation metrics include core metrics and auxiliary metrics. The core metrics include: F1-score (combined accuracy and recall), which requires an F1-score ≥ 0.93 for labels 0 to 4, with an F1-score ≥ 0.92 for identifying true extreme data in label 0, and an average F1-score ≥ 0.95; The auxiliary indicators include: error data identification accuracy ≥ 96% (accuracy of tag 1), missing data identification accuracy ≥ 94% (accuracy of tag 2), sensor aging data identification accuracy ≥ 92% (accuracy of tag 4), suspicious data misclassification rate ≤ 5% (the proportion of tag 3 that is misclassified as other tags), normal data missed detection rate ≤ 3% (the proportion of tag 0 that is misclassified as other tags), and true extreme data misdetection rate ≤ 4% (the proportion of true extreme data in normal data that is misclassified as abnormal).

[0127] S304: Execute decisions based on the trained hybrid classification model, including high-confidence decisions and low-confidence reviews.

[0128] For high-confidence determinations, if the probability of a certain final classification label is... P final ( k If the value is greater than or equal to 0.9, the final category label is output directly. Subsequent processes are only executed for labels 0 / 1 / 2 / 4. Label 0 (Normal data, final classification label): Directly output to the result set and marked as valid-cleaned; If it is real extreme data, it is necessary to additionally meet the following conditions: wind speed after terrain correction ≤ 75m / s (threshold for steep slope wind towers, ≤ 70m / s for gentle terrain) and KDE probability density ≥ 0.0005 (threshold after extreme weather calibration), and mark it as valid-cleaned-real extreme to ensure that real extreme data is traceable; Tag 1 (Error Data, Final Category Tag): Enters the intelligent quality control cleaning process in the subsequent step S40, marked as pending cleaning - error, indicating that the data is erroneous and needs to be cleaned to correct the errors in the data; Tag 2 (Missing data, final classification tag): Enters the missing data repair process in the next step, marked as pending cleaning - missing data, indicating that this data is missing data and needs to be repaired to fill in the missing data; Tag 4 (Sensor aging data, final classification tag): Enters the sensor aging deviation correction process in the subsequent steps, marked as to be cleaned - aging deviation, and the data deviation is adjusted by the aging correction model.

[0129] For low-confidence judgments, if the probability of all final classification labels is <0.9, label 3 (suspicious data) is marked as an intermediate state. A review list is automatically generated and pushed to the operation and maintenance platform, triggering a manual review mechanism. After the review is completed, label 3 is reclassified to label 0 / 1 / 2 / 4 and the corresponding process is executed.

[0130] The qualification requirements for reviewers are as follows: Education and major: Associate degree or above, majoring in meteorology, new energy science and engineering, data science or related fields; Work experience: At least 1 year of experience in wind farm meteorological data processing or equipment operation and maintenance, and familiar with the working principle of automatic weather stations; Training and assessment: Candidates must complete specialized training in complex terrain meteorological data anomaly identification, extreme weather data processing, and sensor aging characteristic analysis, and pass the assessment (out of 100 points, passing score 80 points) before they can participate in the review.

[0131] The review time requirements are as follows: Regular low-confidence data (non-extreme weather periods): review to be completed within 24 hours from the date of the list generation (time base is the local time of the wind farm location); Low-confidence data during extreme weather periods (typhoon / thunderstorm warning periods): review to be completed within 2 hours from the date of the list generation, to ensure that the data is used in a timely manner for power forecasting and operation and maintenance decisions.

[0132] The review checklist should include the following structured information, entered using a fixed Excel template, including basic information, model output information, supporting evidence information, and a review comments section.

[0133] Basic information includes data timestamps (accurate to the second), wind tower ID, wind tower terrain parameters (slope, aspect, terrain type), equipment information (sensor model, years of service, last calibration time), and whether the period of anomaly occurrence was during an extreme weather warning period (yes / no, if yes, the warning type and level must be specified); model output information includes the probability distribution of five types of labels (e.g., P 0 = 0.4, P 1 =0.3, P 2=0.2, P 3 = 0.1, P4=0.0), feature importance ranking (top 5 features and their contribution, such as terrain-corrected wind speed contributing 25%, equipment health score contributing 18%); supporting evidence information includes data from adjacent reference stations (within 3km, or up to 5km if none), equipment operation and maintenance logs (fault records and calibration records within 12 hours before and after the abnormal period), historical data comparison (average and fluctuation range of data for the same period ±3 days on the same date in the past 3 years), records of extreme weather impact (if any, must include correlation analysis between actual weather conditions and data anomalies); the review opinion column includes the review conclusion (normal data / erroneous data / missing data / sensor aging data), judgment basis (must be combined with basic information, model output, and detailed explanation of supporting evidence), processing suggestions (direct retention / repair / removal / further observation), reviewer's signature, and review time field.

[0134] Regarding the application and iteration cycle of the review results, after the review is completed, the maintenance personnel must enter the structured review results into the meteorological data quality control platform within 1 hour. The system will automatically associate the corresponding data entries and update the label information. Every Sunday at 24:00, the system will automatically extract cases that were manually reviewed and corrected within the week, such as data that was misjudged by the model as incorrect data or confirmed as normal extreme data by manual review, and add them to the training set of the machine learning precision judgment module to ensure that the data is used for model optimization in a timely manner. At the end of each month, the model correction effect of the review data for the month will be evaluated, and the improvement in model accuracy after review will be calculated. The improvement is required to be ≥2%. If the standard is not met, the reasons must be analyzed and the model parameters or review standards adjusted.

[0135] Based on step S30, step S40 is executed through the intelligent quality control cleaning module. Step S40 specifically includes: performing differentiated repair on different final classification labels, isolating and removing invalid data that meet the removal rules, and generating full-process traceability information for each data point in the entire process, including original data information, quality control cleaning process information, operation subject and time information, to obtain a standardized meteorological dataset after quality control.

[0136] In some embodiments, for the error data (label 1), missing data (label 2), and sensor aging data (label 4) output by the accurate judgment module, the intelligent quality control cleaning module formulates differentiated cleaning strategies according to the anomaly type to achieve classified repair, invalid removal, and traceability. Specifically, it includes the following steps S401-S403.

[0137] S401: Perform corresponding category repair for different category labels.

[0138] The intelligent quality control cleaning process only performs differentiated repair on the final classification labels (labels 1 / 2 / 4) output by the model with high confidence. Label 3 in the intermediate state is manually reviewed and classified, and then the corresponding repair process is performed according to the final classification label. For erroneous data repair (label 1), it is divided into three categories according to the cause of the error: sensor failure repair, transmission interference repair, and environmental noise repair, as detailed below.

[0139] Sensor fault repair (continuous error duration 30min~6h): Virtual sensor model based on cross-parameter physical correlation (random forest regression); input features include P、T、U Topographic parameters ( s、a The following data were collected: synchronous data from nearby reference stations (selection criteria: priority within 3km, if abnormal, then expand to 5km, altitude difference ≤250m), historical calibration records of equipment (such as sensor sensitivity parameters), and equipment health scores; the model was trained using synchronous data from the normal sensor period of the past 6 months, with a training set sample size ≥100,000, ensuring that the RMSE between the corrected value and the true value was ≤0.8m / s (wind speed), ≤15° (wind direction), and ≤1℃ (temperature).

[0140] An example of repair is as follows: the wind speed sensor is icing up, causing the reading to be lower than expected (actual wind speed 8 m / s, observed value 3 m / s). The model input is the same-day pressure gradient ΔP / Δh = 0.02 hPa / m, air temperature -5℃, reference station wind speed 9 m / s, equipment health score 0.7 (service life 2 years, calibration interval 2 months), output repair value 7.8 m / s, error ≤ 0.2 m / s.

[0141] Transmission interference repair (single-point / short-term error, duration <30min): Multi-source weighted fusion interpolation is used, with weights dynamically allocated according to data reliability. The weight allocation rules are as follows: Data from different layers within the same tower (e.g., repairing 80m data from a 70m height wind tower): Weight 50% (data from the same tower has the highest correlation, with a Pearson correlation coefficient ≥0.8), so the weight is set at 50%; Nearby reference station data (meeting the screening conditions mentioned above, prioritizing within 3km, and extending to 5km for anomalies): Weight 30% (strong spatial correlation, correlation coefficient ≥0.6); ERA5 reanalysis data (spatial resolution 0.25°×0.25°, temporal resolution 1h): Weight 20% (as a fallback, suitable for scenarios without nearby reference stations or missing data from the same tower, correlation coefficient ≥0.4). If data from a certain data source (e.g., a nearby reference station) is abnormal (e.g., exceeding the normal fluctuation range), its weight is automatically allocated to other valid data sources to ensure that the total weight sum is 100%. Environmental noise remediation (such as random fluctuation errors caused by electromagnetic interference) employs a combination of wavelet thresholding and moving average methods. The db4 wavelet basis is selected, and the original data is decomposed into three levels. A soft threshold is set (the threshold calculation formula is: ...). l = s × ,in s (where n is the standard deviation of noise and n is the data length) High-frequency noise components are removed; the denoised data is processed by a 5-point moving average to smooth local fluctuations and retain the overall trend of the data (such as the short-term trend of wind speed); the rate of change of the repaired data is compared with that of the normal data at adjacent time points. If the rate of change of wind speed is ≤2m / s·min and the rate of change of temperature is ≤0.5℃ / min, the repair is deemed effective; otherwise, the wavelet threshold and the sliding window size are re-optimized.

[0142] For missing data repair (tag 2), it includes short-term missing data repair (missing data duration ≤ 1h), medium-term missing data repair (missing data duration 1h~12h) and long-term missing data repair (missing data duration > 12h).

[0143] Short-term data loss repair is based on a linear interpolation and neighboring data fusion method using temporal correlation; it utilizes normal data (such as wind speed sequences) from 10 minutes before and after the data loss period. v 1, v 2,..., v The missing values ​​were calculated using a linear interpolation formula: v 缺测 = v 前 +( t 缺测 - t 前 ) / ( t 后 - t 前 )×( v 后 - v 前 ),in t 前 , t 后 These are the timestamps of normal data before and after the missing data period; corresponding data from nearby reference stations (distance ≤ 3km) during the same period are also included. v 时序插值 The final repaired value is obtained by weighting the values ​​according to time synchronization (the smaller the time deviation, the higher the weight, with a weight range of 0.3 to 0.7) and fusing them with the time series interpolation results. v 最终 The formula is: v 最终 = v 时序插值 ×(1- w )+ v 邻近参考站 × w ( w(Weighted by neighboring data); the error between the repaired value and subsequent measured data (if any) must be ≤1.2m / s (wind speed), ≤1.8℃ (temperature), and ≤6% (humidity).

[0144] Mid-term missing data remediation is based on a multi-feature prediction model of LightGBM; input features include historical data from the same period 24 hours before and after the missing data period (e.g., wind speed and temperature trends during the same period last week), complete data from adjacent meteorological towers (prioritizing within 3km, expanding to 5km for anomalies), ERA5 reanalysis data (extracting corresponding grid data according to the missing data period timestamp), and terrain correction parameters (slope). s Slope aspect a (Corresponding wind speed correction coefficient) and equipment health score; trained with complete data from the past year without missing data, and optimized parameters using 5-fold cross-validation to ensure the model's prediction of missing data. R 2 ≥0.92 (wind speed) R 2 ≥0.90 (temperature); Input the input features into the trained LightGBM model, output the predicted value of missing data minute by minute, and then smooth it through a sliding window (window size 10 minutes) to eliminate prediction fluctuations.

[0145] Long-term missing data restoration primarily relies on spatial substitution, supplemented by historical analogy. Two to three reference stations with similar terrain characteristics to the missing meteorological towers (slope difference ≤ 5°, aspect difference ≤ 30°, elevation difference ≤ 150m) and complete data are selected (prioritizing stations within 3km, extending to 5km for anomalies). The correlation coefficient between each reference station and the historical data of the missing towers (Pearson correlation coefficient ≥ 0.75) is calculated. The restoration baseline value is obtained by weighting the values ​​based on correlation, using the following formula: v 基础 =Σ( v 参考站i × r i ) / Σ r i (in the formula r i For the first i (Correlation coefficients of each reference station); extract the variation patterns of missing data from the towers over the past three years (such as daily average wind speed fluctuations and peak occurrence times), and use this data for spatial substitution. v The baseline is adjusted for trend correction, with the correction coefficient k = historical fluctuation range of the missing tower during the same period / fluctuation range of the reference station during the same period, resulting in the final correction value. v 最终 = v 基础 ×k; If a reference station that meets the conditions cannot be found (e.g., the reference station data also has missing measurements), it is marked as unrepairable and enters the invalid data removal process.

[0146] For sensor aging data repair (Label 4), an aging deviation correction model based on equipment health is adopted; the aging deviation coefficient is calculated according to the sensor's service life, calibration cycle, and historical failure count. β The formula is: β =1-0.3×min(service life / 8,1)-0.4×min(calibration interval / 12,1)-0.3×min(historical failure count / 10,1)), where service life is in years, calibration interval is in months, and historical failure count is the number of failures in the past year; β The value range is 0.5~1.0 (corresponding to 50%~100% sensitivity, consistent with sensor aging patterns); the repair formula is as follows: Wind speed repair value v 修复 = v 观测 × β +0.5×(Wind speed at the nearest reference station during the same period) v 观测 By using neighboring data for calibration, the error of a single coefficient correction is reduced; Temperature restoration value T 修复 = T 观测 +0.3×(historical normal temperature for the same period - T 观测 )×(1- β Adjustments were made based on historical data from the same period.

[0147] The rate of change of the repaired data is calculated by dividing the absolute value of the difference between the repaired data and the normal data for the preceding and following 5 minutes by the time interval, with each time interval being 1 minute. The repaired data must meet the following requirements: a rate of change of ≤1.5 m / s·min (wind speed) and ≤0.3℃ / min (temperature) compared to the adjacent normal time period; otherwise, the data must be readjusted. β The weighting of the coefficients has been adjusted (the weighting for years of use has been changed from 0.3 to 0.4, and the weighting for calibration interval has been changed from 0.4 to 0.3), and the calculation has been recalculated. β Then perform the repair; if it still fails to meet the standard after two consecutive adjustments, it is marked as invalid aging repair and enters the invalid data removal process.

[0148] S402: Invalid data removal.

[0149] Set the rejection criteria: Data whose errors exceed the standard after error correction, such as wind speed correction value RMSE > 1.2m / s, temperature correction value error > 1.5℃, and wind direction correction value error > 20°; missing data that cannot be supplemented by any correction algorithm, such as continuous missing data duration > 24h, and no nearby reference station, different layers of the same tower, or ERA5 reanalysis data available; data that still cannot meet the accuracy requirements after sensor aging data correction, such as wind speed correction value change rate > 2m / s·min.

[0150] Data that could not be clearly categorized even after multiple manual reviews, and which has no reference value for subsequent wind farm power prediction and operation and maintenance analysis: such as data without timestamps, without meteorological tower identification, and with chaotic data format.

[0151] Based on the above rejection criteria, the following rejection process will be executed: 1) Automatic system marking: Based on the above rejection criteria, the system automatically marks records that exceed the standard after error data repair or that cannot be repaired due to missing data, generating an "Invalid Data Rejection List". The list includes data timestamp, wind tower ID, data type (wind speed / temperature / humidity, etc.), anomaly type (error / missing data), and rejection reason (e.g., wind speed RMSE after repair = 1.5m / s > 1.2m / s, no valid reference data for 26 consecutive hours of missing data, wind speed change rate after aging repair = 2.3m / s·min > 2m / s·min). 2) Manual review and confirmation: The list is automatically pushed to the wind farm operation and maintenance platform. The operation and maintenance personnel need to review the list within 24 hours in conjunction with the equipment operation and maintenance log (such as whether there is a record of complete sensor failure during this period), on-site inspection (such as whether the wind measurement tower is damaged due to extreme weather), and extreme weather impact assessment report to confirm whether it meets the exclusion criteria. 3) Data Isolation and Storage: Data that needs to be removed after review is not directly deleted from the database, but is transferred to an invalid data isolation database for storage, retaining complete traceability information (such as original data, repair process records, and review opinions) for subsequent auditing and problem tracing; at the same time, it is marked as removed-invalid in the corresponding position of the normal dataset to avoid misuse when calling data later; if subsequent manual review finds that the data in the isolation database is valid data (such as real extreme data), it needs to be added to the training set according to the data return rules and removed from the isolation database, and the normal dataset label is updated.

[0152] S403: Data traceability management.

[0153] This embodiment records traceability information in the following way: For each data point in the entire intelligent quality control cleaning process, the following complete traceability information is recorded to ensure that every step of the operation is traceable and every piece of data is verifiable. The original data information includes the data acquisition time (accurate to the second), the acquisition device number (e.g., wind speed sensor number WS-001, data acquisition device number DC-05), the original observation value, the device status at the time of data acquisition (e.g., normal operation, low battery warning, aging warning), and whether the acquisition period is during an extreme weather warning period (yes / no, warning type and level); the quality control cleaning process information includes the initial inspection results (abnormality type and binary code, e.g., threshold abnormality + thunderstorm lightning abnormality - 1001), the accurate judgment label and probability (e.g., label 1 - erroneous data, probability 0.96, label 4 - sensor aging data, probability 0.93), and the cleaning and repair algorithm type (e.g., sensor fault repair - random forest regression, short-term missing data repair - linear interpolation). The data includes: value + neighbor fusion, aging data repair - deviation coefficient correction), repair data source and weight (e.g., 70m data in the same tower -50%, neighbor station A -30%, ERA5 -20%, 3km reference station abnormal, extended to 5km reference station B -30%), comparison of data before and after repair (original value, repaired value, error value), removal decision information (e.g., removal reason: continuous missing measurement for 26 hours, reviewer: Zhang San); operation subject and time include records of automatic system operation (e.g., 2025-10-10 08:30 system automatic repair), records of manual operation (e.g., 2025-10-10 10:15 maintenance personnel Zhang San reviewed and confirmed removal, 2025-10-10 11:30 maintenance personnel Li Si entered the review result).

[0154] This embodiment uses blockchain technology to construct a data traceability chain. The nodes on the chain include the wind farm data center, the operation and maintenance management platform, and third-party supervision agencies (such as wind power industry regulatory departments). Each node stores a complete copy of the traceability log to ensure that the data is tamper-proof and tamper-proof. At the same time, a structured copy of the traceability information is retained in the local database (using a MySQL database, stored in separate tables by meteorological tower ID-year-month) to improve query efficiency. The query function supports multi-dimensional combined queries, specifically including basic queries, process queries, and result queries.

[0155] Basic queries can be performed by time range (e.g., 2025-10-01 00:00 to 2025-10-01 23:59), wind tower ID (e.g., WT-03), and data type (e.g., wind speed); process queries can be performed by cleaning status (e.g., repaired, removed, pending review), repair algorithm type (e.g., random forest regression repair, wavelet denoising repair, aging bias correction), and anomaly type (e.g., sensor failure, thunderstorm lightning strike anomaly, sensor aging); result queries can be performed by repair error range (e.g., wind speed repair error 0~0.5m / s), and original data removal criteria. The system supports queries based on factors such as inability to repair or exceeding limits after repair, as well as queries related to extreme weather events (such as typhoon data or thunderstorm data). Query results can be displayed in the form of time series graphs (showing the time distribution of original data values, repaired values, and removal markers), statistical reports (showing the proportion of different types of abnormal data, repair success rate, and aging data proportion for each meteorological tower), and flowcharts (showing the entire process of quality control and cleaning for a single data point, including manual review nodes). It also supports exporting to Excel (for data statistical analysis) and PDF (for report archiving) formats.

[0156] This embodiment provides an exemplary application scenario for traceability information. In this scenario, the traceability information is used to analyze abnormal data types of different wind measurement towers and sensors (e.g., the WT-05 tower wind speed sensor experiences an average of 3 icing failures per month, and the WT-08 tower sensor has an aging deviation data ratio of 15% after 4 years of use). Targeted maintenance plans are then developed (e.g., adding a heating and de-icing device to the WT-05 tower wind speed sensor, and shortening the calibration cycle of the WT-08 tower sensor from 6 months to 3 months). Cases of manual review and correction in the traceability information are extracted (e.g., data misjudged by the system as erroneous, but confirmed as normal extreme data after manual review) and added to the training set of the machine learning precision judgment module to optimize the model's classification accuracy. When industry regulatory authorities conduct inspections, complete traceability logs are provided to prove the compliance of the data quality control and cleaning process (e.g., all removed data was manually reviewed, and original records were retained; aging data repair was based on equipment health parameters, and the correction process is traceable), thus meeting regulatory requirements. For abnormal data during extreme weather periods such as typhoons and thunderstorms, the data processing flow is reviewed by tracing the source information, and the identification rules and repair algorithms of the extreme weather adaptation sub-module are optimized (such as adjusting the judgment threshold for abnormal typhoon wind speed changes).

[0157] Based on step S40, step S50 is executed through the effect verification and dynamic iteration module. Step S50 specifically includes: based on the standardized meteorological dataset after quality control, the quality control effect is verified through four categories of wind farm indicators: data integrity, data accuracy, business practicality, and economic benefits. Based on the verification results, the feedback quality control process data, and the manual review results, incremental training, manual feedback optimization, and dynamic iteration of terrain and equipment adaptation updates are performed.

[0158] In some embodiments, step S50 is used to ensure the long-term effectiveness of the method, which specifically includes the following steps S501-S502.

[0159] S501: Build multi-dimensional quality verification metrics to perform data integrity verification, data accuracy verification, business usability verification, and economic benefit verification.

[0160] Quality verification metrics used for data integrity verification include: Valid data percentage = (total data volume - invalid data volume) / total data volume, requiring ≥90%, of which the valid percentage of core elements such as wind speed and wind direction is ≥95% (to meet the power prediction data input requirements); Missing data repair rate = repaired data volume / missing data volume, requiring ≥85%, and 100% repair of data missing for ≤6 hours during the grid connection period (data during the grid connection period directly affects the calculation of power generation); Wind farm-specific integrity indicators: The integrity of the data associated with the wind measurement tower and the turbine is ≥98% (such as the time synchronization rate of wind speed data of the wind measurement tower and power data of the turbine), to avoid power prediction deviations due to data asynchrony.

[0161] For data accuracy verification (quality verification indicators for matching the accuracy of wind farm calculations include:) Accuracy of repaired values: 1000 manually annotated wind farm verification data points (including wind speed, wind direction, temperature, and air density) were selected. The RMSE between the repaired values ​​and the actual values ​​was calculated as follows: wind speed ≤ 0.5 m / s (meeting the accuracy requirements for power curve fitting), wind direction ≤ 15° (adapting to turbine yaw control), temperature ≤ 1℃ (ensuring air density calculation error ≤ 2%), and air density ≤ 0.05 kg / m³. 3 ; Anomaly identification accuracy: Error data identification accuracy ≥95% (focusing on wind farm-specific faults such as wind speed sensor icing and wake interference), false detection rate ≤5% (avoiding the exclusion of real extreme wind energy data), and false negative rate ≤3% (preventing faulty data from entering the power prediction model). Specific accuracy indicators for wind farms: the deviation between the wind speed after terrain correction and the measured wind speed is ≤3% (to ensure the accuracy of wind energy resource assessment), and the deviation in air density calculation is ≤2% (to ensure that the deviation in annual power generation assessment is ≤1%).

[0162] The quality verification metrics used for business usability verification include: Power prediction error: Input the cleaned data into the wind farm power prediction model (such as the XGBoost power prediction model) and calculate the NMAE (Normalized Mean Absolute Error). The requirement is that it should be reduced by ≥20% compared to before cleaning. According to industry practice, for every 1% reduction in power prediction NMAE, the wind farm's curtailment rate can be reduced by 0.5%-0.8%, and the annual power generation can be increased by 0.3%-0.5%. In the test of this method at the Shanxi Canyon Wind Farm, the NMAE decreased from 28% to 21%, a reduction of 25%. Power generation assessment deviation: The annual power generation is calculated based on the cleaned data and compared with the actual grid-connected power generation. The deviation is ≤3%. Among them, the proportion of power generation contributed by the real extreme data must match the frequency of extreme weather in the same period in history (e.g., the proportion of power generation during typhoons is ≤5%, which is consistent with the proportion of the average number of days affected by typhoons in a year). Operation and maintenance decision support: After optimizing the operation and maintenance plan through traceability information, the unplanned downtime caused by sensor failures is reduced by ≥20% compared with before optimization, and the proportion of invalid data caused by sensor aging is reduced by ≥10%; at the same time, the unit switch-out strategy based on real extreme data can reduce the number of unit overload failures by ≥30%.

[0163] Quality verification indicators used for economic benefit verification include: Annual revenue increase = (Wind curtailment rate before cleaning - Wind curtailment rate after cleaning) × Annual power generation × Electricity price - Data processing cost - Additional cost of operation and maintenance optimization The wind curtailment rate is calculated using data from the same period of the past year for wind farms. The pre-cleaning rate is the rate before using this method, and the post-cleaning rate is the rate after using this method. The data source is the wind farm's grid connection dispatch records. This method reduces the NMAE from 28% to 21% (a 25% reduction), corresponding to a wind curtailment rate reduction from 8% to 5% (a 3 percentage point reduction), echoing the business verification indicators mentioned earlier. Annual power generation is calculated as wind farm installed capacity (e.g., 100MW) × annual utilization hours (calculated based on the wind farm's average annual utilization hours over the past 3 years, with data sourced from wind farm production reports), i.e., annual power generation = installed capacity × annual utilization hours. A 3 percentage point reduction in the wind curtailment rate corresponds to a 3% increase in annual power generation. For a 100MW wind farm with 2500 hours of annual utilization, the annual increase in power generation = 100MW × 2500h × 3% = 7.5 × 10 6 kWh.

[0164] It should be noted that the annual increase in power generation includes two parts: 7.5 million kWh from a 3 percentage point reduction in the wind curtailment rate, and 500,000 kWh from the recovery of unplanned power generation due to reduced sensor malfunctions, totaling 8 million kWh. Data processing costs include GPU computing power costs (calculated based on GPU model, runtime, and electricity price) and storage costs (hardware and software costs for storing traceability information), with data sourced from the wind farm IT operation and maintenance cost accounting report; additional operation and maintenance optimization costs, such as the extra costs incurred by installing de-icing devices for sensors and shortening calibration cycles, are also sourced from the wind farm operation and maintenance cost accounting report. S502: Implements a dynamic model iteration mechanism, including incremental training strategies, human feedback optimization, and terrain adaptation updates.

[0165] The incremental training strategy includes: automatically adding new data weekly (100,000 records / week), using a sliding time window (retaining data from the past year), and removing expired data (>1 year) to avoid data distribution drift; new data must include complete equipment health parameters (years of use, calibration cycle updates) and extreme weather-related information (if any); monthly incremental training (wind farm operation and maintenance cycles are typically monthly), loading historically optimal model parameters, updating the model with new data, and training for 10 rounds (only fine-tuning parameters to reduce computational costs); if extreme weather such as typhoons or thunderstorms occur in the month, an additional incremental training is added to optimize the extreme weather adaptation submodule; parallel computing using a GPU cluster (2 NVIDIA A100 GPUs) is employed, with a single incremental training session taking ≤2 hours, without affecting real-time quality control operations; actual testing shows that 2 NVIDIA GPUs... The A100 cluster processes 1 million data points at the 1-minute level, with a total processing time of ≤3 minutes including feature engineering and model inference. When adding extreme wind energy data exceeding 40m / s after terrain correction (such as gusts caused by canyon narrowing effects), an additional dedicated training session is added to optimize the extreme data identification logic. The terrain correction function needs to be updated synchronously during training. f The coefficient library of (s,a) ensures that new terrain scene data can be adapted.

[0166] The manual feedback optimization includes: establishing a wind farm quality control result feedback platform, where maintenance personnel can mark misjudged data by the model (such as misjudging wind direction anomalies caused by wake interference as sensor failure), and supplementing unit operation logs and on-site inspection records (such as unit location during wake interference periods and terrain obstruction conditions); the platform supports uploading on-site photos (such as sensor icing and wake impact marks) as supporting evidence to improve the credibility of feedback; and adding 500-1000 manually marked wind farm-specific cases (such as extreme gusts, wake interference, and sensor icing) to the training set each month to retrain the model. Classification parameters, adjust terrain adaptation coefficient (e.g., fine-tune the correction coefficient for steep windward slopes from 1.6 to 1.62), and anomaly detection threshold (e.g., adjust the winter wind speed suspicious threshold from 35m / s to 36m / s); after feedback optimization, calculate the improvement in recognition accuracy of wind farm-specific anomaly types (e.g., wake interference, icing failure, sensor aging) monthly, requiring an improvement of ≥5% / month, and stabilizing at ≥95% within 6 months; if the target is not met, backtrack on feedback cases and adjust the model attention mechanism weight allocation rules (e.g., in wake interference scenarios, add an extra 10% weight to spatial correlation features).

[0167] Terrain adaptation updates include: when a new wind measurement tower is added to a wind farm, its terrain parameters (slope, aspect, terrain type) are automatically read, and the terrain correction function is recalculated with reference to the applicable terrain range table. f ( s , a The reference station range of the spatial correlation model is updated synchronously (3km is preferred, and if there is no effective reference station within 3km, it is expanded to 5km, matching the typical layout spacing of wind farm meteorological towers); the new meteorological tower data needs to undergo a one-month trial operation quality control (an additional 30% of the data is manually reviewed), and after confirming that the adaptation is correct, it is included in the regular quality control process; the wind farm DEM data is updated annually (e.g., changes in terrain barrier caused by road construction and vegetation clearing), and the terrain barrier coefficient k is recalibrated. block (e.g., after the road is opened, the original barrier area k) block (Adjusted from 0.6 to 1.0); densely vegetated areas k block (Adjusted from 0.9 to 0.85); After the update, backtesting with nearly 3 months of historical data is required to ensure that the deviation between the wind speed after terrain correction and the measured wind speed is ≤3%; After the sensor is replaced, the equipment health profile (initial service life of the new sensor, calibration cycle) is automatically updated, and the virtual sensor model is retrained (based on normal data from the new sensor). The adaptation cycle is shortened from the traditional 15 days to 24 hours. During this period, the new sensor data needs to be cross-validated with the old sensor data to ensure that the repair value RMSE ≤0.8m / s (wind speed) and ≤1℃ (temperature).

[0168] To verify the technical effectiveness of the present invention, data from the entire year of 2023 for three typical complex terrain wind farms were selected for testing.

[0169] The test objects are as follows: Shanxi Canyon Wind Farm (35° slope, 8 wind measurement towers, approximately 1.2 million data points per day); Qinghai Plateau Wind Farm (slope 12°, 6 wind measurement towers, approximately 900,000 data entries per day). Yunnan mountain wind farm (slope 25°, number of wind measurement towers 10, daily data volume approximately 1.5 million).

[0170] The obtained test indicators and results are shown in Table 2.

[0171] Table 2 Test results of the method of the present invention ; As shown in Table 2, the method of the present invention meets the preset objectives in three types of complex terrain wind farms, and has strong adaptability, which is significantly better than the prior art.

[0172] In summary, this invention addresses the industry pain points of data quality control for automatic weather stations in wind farms with complex terrain. It constructs a closed-loop system covering the entire process from data preprocessing, anomaly identification, accurate judgment, quality control and cleaning to dynamic iteration for effect verification. This system overcomes the technical bottlenecks of traditional methods, such as poor adaptability to complex terrain, insufficient accuracy in identifying extreme weather scenarios, and weak data processing capabilities due to sensor aging. It comprehensively improves the integrity, accuracy, and usability of meteorological data from wind farms, providing high-quality and standardized data support for the operation and maintenance of wind farm power prediction and wind energy resource assessment equipment.

[0173] This invention integrates a multi-dimensional initial detection mechanism that combines threshold-verified temporal variational autoencoder analysis with spatial correlation verification and distribution offset detection. This mechanism, combined with a hybrid classification model using a LightGBM bidirectional long short-term memory network and an attention mechanism, achieves synergistic effects. The accuracy of erroneous data identification increases from 72% to 96% compared to traditional methods, while the false positive rate decreases from 28% to 4%. This performance has been verified using 2 million measured data points from a canyon wind farm in Shanxi. Simultaneously, the accuracy of identifying sensor aging data reaches 92%, addressing the industry pain point that traditional methods cannot identify systematic data biases caused by sensor aging. Furthermore, through terrain-adaptive feature attention mechanism weight optimization and the collaborative design of an extreme weather adaptation submodule, the accuracy of identifying real extreme data such as strong gusts in canyons reaches 98%, and the accuracy of identifying abnormal data under extreme weather conditions such as typhoons and thunderstorms reaches 90%, effectively avoiding the false positive problem of traditional fixed threshold methods and reducing the overall false positive rate by 30%. This invention dynamically adjusts the probability density threshold of the kernel density estimation model during extreme weather periods, adjusting the probability density threshold corresponding to wind speed and temperature to 0.0005, which can reduce the misjudgment rate of real extreme wind speeds such as typhoon gusts of 55 m / s from 25% in traditional methods to below 2%.

[0174] Compared to traditional general meteorological data quality control methods that do not incorporate terrain parameters, this invention, through targeted design of the spatial correlation characteristics of terrain barrier coefficients using slope and aspect correction functions, improves the anomaly identification F1-score of wind farms in different terrains such as mountainous areas, plateaus, and canyons from 0.78 to 0.95. The quality control accuracy deviation for different terrain scenarios, including gentle slopes, steep slopes, sunny slopes, and shady slopes, does not exceed 3%, solving the industry pain point of inconsistent quality control standards within the same wind farm. The terrain-corrected wind speed obtained based on the slope and aspect correction function avoids the overestimation of wind speed caused by traditional terrain coefficient overlay, reducing the wind energy resource assessment deviation of plateau wind farms from 8% to 3%. This invention adds an extreme weather adaptation submodule and sensor aging adaptation function, enabling accurate processing of sensor aging scenarios in extreme weather conditions such as typhoons and thunderstorms, filling the technical gap in traditional methods for such scenarios. The accuracy of data processing during extreme weather periods is improved by 25% compared to traditional methods, and the effective utilization rate of sensor aging data is increased from 30% to 90%. The sensor aging deviation coefficient calculation logic designed in this invention ensures a lower limit of 0.5 for the deviation coefficient, perfectly aligning with the actual operating principle that the sensitivity decay of sensors does not exceed 30% within an 8-year service life. This improves the accuracy of aging data repair by 40% compared to the traditional fixed coefficient method. When a new wind measurement tower is added to a wind farm or when terrain conditions change, this invention can automatically calibrate the terrain correction function and spatial correlation model, shortening the adaptation period from 15 days to 24 hours. After sensor replacement, the system can automatically update the equipment health parameters and model correlation coefficients, shortening the adaptation period from 7 days to 12 hours, significantly reducing the wind farm operation and maintenance adaptation costs. This invention sets up a trial operation quality control mechanism for new wind measurement towers, manually reviewing 30% of the processed data in the first month of operation. This ensures that the data adaptation accuracy of the new tower is not less than 95%, effectively avoiding the power prediction deviation problem caused by abnormal data during the traditional adaptation period.

[0175] This invention employs the Dask distributed computing framework combined with NVIDIA A100 GPU acceleration, processing 1 million data entries per day in less than 8 minutes, far shorter than the 60+ minutes required by traditional serial methods. During extreme weather periods, when the data acquisition frequency is increased to the 1-second level, daily data processing time is less than 12 minutes, with a processing latency of less than 5 minutes, fully meeting the timeliness requirements of wind farm power prediction outputting quality control results within 15 minutes for real-time grid dispatch. The incremental training mechanism used in this invention requires only 2 hours per month for incremental training, far shorter than the 24 hours required for traditional full training. Additional incremental training after extreme weather events requires only 1 hour, and model parameter fine-tuning requires no manual intervention from professional algorithm engineers, reducing annual computing costs by 60%. Furthermore, this invention standardizes the manual review process, improving manual review efficiency by 30% and effectively preventing review delays from affecting the normal use of data.

[0176] When the cleaned data processed by this invention is used for short-term power prediction from 0 to 4 hours, the normalized average absolute error (NAR) is reduced by at least 23% compared to traditional methods, far exceeding the 10% to 15% reduction achieved by traditional methods. During extreme weather periods, the NAR is reduced by at least 18%, meeting the grid dispatch requirement that the NAR not exceed 15%, effectively reducing wind curtailment caused by prediction deviations, and lowering the curtailment rate by 2% to 3%. The annual power generation assessment based on the cleaned data deviates from the actual grid-connected power generation by no more than 3%, far exceeding the deviation level of over 5% achieved by traditional methods, effectively avoiding investment misjudgments during the wind energy resource assessment stage. Taking a 100MW wind farm with a total investment of 1 billion yuan as an example, reducing the assessment deviation from 5% to 3% can reduce investment risk by 30 million yuan over a 20-year operating cycle. This calculation is based on a grid-connected electricity price of 0.6 yuan per kilowatt-hour. When the terrain-corrected wind speed data is used for the early-stage wind energy resource assessment of wind farms, the effective wind time calculation error does not exceed 2%, and the turbulence intensity assessment error does not exceed 5%. It can provide an accurate basis for the selection of different models of units such as 3MW and 4MW, and avoid the mismatch problem of large units matching small wind energy resources or small units matching large wind energy resources.

[0177] This invention, through the synergy of equipment health characteristics and anomaly identification models, can reduce the sensor fault false detection rate to no more than 3%, and can provide early warnings of faults such as wind speed sensor icing and wind direction sensor jamming 7 to 15 days in advance, effectively reducing unplanned downtime of the units. The number of unplanned downtimes per wind farm per year can be reduced from 5 to 1, which, based on a loss of 10,000 yuan per downtime, can reduce downtime losses by 40,000 yuan per year. This invention can analyze the anomaly types of different sensors on different meteorological towers based on full-process traceability information, such as periodic icing faults of wind speed sensors on specific meteorological towers, or wind direction data anomalies caused by wake interference on specific meteorological towers, and then formulate targeted operation and maintenance plans. By installing heating and de-icing devices on meteorological towers with frequent icing faults, annual operation and maintenance costs can be reduced by 120,000 yuan. By adjusting the position of meteorological towers affected by wake interference to avoid wake areas, annual fault repair costs can be reduced by 80,000 yuan. This invention has an accuracy rate of no less than 98% in identifying extreme weather anomaly data, and can effectively distinguish between real extreme wind conditions and sensor fault data. For example, a genuine gust of 55 m / s during a typhoon can be accurately identified, triggering the unit to shut down to avoid equipment overload and damage. False high wind speed data of 99 m / s caused by sensor malfunction can be precisely filtered out, preventing erroneous unit shutdowns and achieving a balance between unit operational safety and power generation optimization.

[0178] Based on calculations for a 100MW wind farm, adopting this invention can reduce the wind curtailment rate by 2% to 3%, increasing annual power generation by approximately 7.5 million kWh. At the local benchmark onshore wind power price of 0.38 yuan per kWh, this translates to an annual increase in revenue of approximately 2.85 million yuan. Simultaneously, annual operation and maintenance costs for the wind farm can be reduced by 500,000 yuan, equipment procurement costs by 300,000 yuan, and sensor lifespan extended from 5 years to 7 years. The annual data processing cost of this invention is approximately 200,000 yuan, and the additional cost for operation and maintenance optimization is approximately 280,000 yuan. The comprehensive calculation shows an annual net increase in revenue of 2.37 million yuan, far exceeding the approximately 800,000 yuan increase in annual revenue achieved by traditional methods, demonstrating significant economic benefits. This invention provides a standardized meteorological data quality control solution for wind farms in complex terrain, addressing the industry pain point of inconsistent quality control standards within the same wind farm. It is adaptable to wind farm scenarios with different terrains and turbine layouts, such as plateaus, mountains, and canyons, effectively lowering the industry's technical barriers. This invention achieves a value upgrade of meteorological data from basic observation and recording to operational decision support through a closed-loop process of data quality control, power prediction, and operation and maintenance optimization, providing technical support for the digital transformation of the wind power industry. Simultaneously, this invention can effectively improve the power generation efficiency of wind farms; a single 100MW wind farm can increase annual power generation by 500,000 to 1,000,000 kWh. Based on a coal consumption of 300 grams per kWh for thermal power, this is equivalent to a reduction of 400 to 800 tons of carbon dioxide emissions annually, demonstrating significant environmental benefits.

[0179] The above embodiments are only used to illustrate the present invention and are not intended to limit the present invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the present invention, and the patent protection scope of the present invention should be defined by the claims.

Claims

1. A method for identifying and intelligently cleaning data anomalies from automatic weather stations in wind farms with complex terrain, characterized in that, The method includes: Multi-source data is acquired and standardized by unifying time resolution, marking missing data, and initially screening invalid data. Physical quantity transformation is performed on the obtained standardized time-series data to construct a multi-dimensional feature matrix. The multi-source data includes wind farm automatic weather station observation data, geographic information data, equipment operation and maintenance data, and extreme weather and operation data. The multi-dimensional feature matrix includes basic statistical features, cross-parameter physical correlation features, terrain adaptation features, equipment health features, and time and extreme weather features. Based on the standardized time-series data and multi-dimensional feature matrix, at least one of the following preliminary checks is performed: threshold and logic verification, time-series consistency detection, distribution offset detection, and extreme weather adaptation verification. Data that triggers any preliminary check anomaly condition is marked as suspicious data, and the anomaly type is recorded by binary encoding. Using the standardized time-series data, multidimensional feature matrix, and suspicious data as input, a hybrid classification model including LightGBM, bidirectional LSTM, and attention mechanism is used to classify and determine the suspicious data. For high-confidence results where the maximum probability of the classification label is greater than or equal to a set threshold, the corresponding classification label is directly output. Differentiated repair is performed on different classification labels, invalid data that meets the removal rules is isolated and removed, and full-process traceability information containing original data information, quality control cleaning process information, operation subject and time information is generated for each data point in the whole process, so as to obtain a standardized meteorological dataset after quality control.

2. The method according to claim 1, characterized in that, The physical quantity conversions include wind direction angle conversion, barometric altitude correction, and humidity unit conversion; The wind direction angle conversion includes converting wind directions within the range of 0° to 360°. d The conversion formula to Cartesian coordinates is as follows: u=sin( d × π / 180), v =cos( d ×π / 180); Where u and v For Cartesian coordinate system components; The barometric altitude correction includes correcting the barometric pressure observed by each meteorological tower to sea-level pressure. The correction formula is: ; in, To observe air pressure, h T0 represents the altitude of the wind measurement tower and T0 represents the sea level reference temperature. γ Let M be the vertical temperature lapse rate, and M be the molar mass of dry air. R This is the universal gas constant. It is the acceleration due to gravity; The humidity unit conversion includes converting relative humidity to specific humidity. q The conversion formula is: q =0.622× e / ( P -0.378 e ); in, P This refers to the ambient atmospheric pressure measured on-site by the wind measurement tower. e The vapor pressure is calculated using the following formula: e = e 0×exp( a ×( T - T 0) / ( T - b )); in, e 0 represents the baseline saturated water vapor pressure, and exp represents the natural exponential function. a These are the empirical fitting coefficients. b It is an empirical constant. T This is the actual measured temperature.

3. The method according to claim 1, characterized in that, The terrain adaptation features include a slope-aspect correction function and a target-reference station wind speed difference rule, wherein: The slope-aspect correction function f ( s , a (), based on wind tunnel experimental data fitting from wind farms; targeting slope s For plateau terrain with a slope of ≤15°, a basic slope correction factor should be set. f s =1.0+0.008× s For 15° < s For mountainous terrain with a slope of ≤30°, a base slope correction factor should be set. f s =1.15+0.006×(s-15); For s For canyon terrain with a slope greater than 30°, a basic slope correction factor should be set. f s =1.45; the basic slope correction coefficient is dynamically adjusted in conjunction with the angle between the slope aspect and the prevailing wind direction. s When the altitude is ≤30°, the plateau terrain f ( s , a )= f s ×1.08, Mountainous terrain f ( s , a )= f s ×1.1, Canyon Topography f ( s , a )= f s ×1.12; 30°< s When <150°, f ( s , a )= f s When the included angle is ≥150°, the plateau terrain... f ( s , a )= f s ×0.88, Mountainous terrain f ( s , a )= f s ×0.9, Canyon terrain f ( s , a )= f s ×0.85; The terrain-corrected wind speed formula is: v '= v / f ( s , a ),in v The measured wind speed is given by s, where s is the slope of the location of the wind measurement tower. a Slope direction, v 'This is the corrected wind speed; The target-reference station wind speed difference rule calculates the wind speed difference Δ between the target meteorological tower and reference meteorological towers within a designated area. v =| v target - v ref |, among which v target The measured wind speed values ​​of the target anemometer tower to be verified during the same period. v ref To reference the concurrently measured wind speed values ​​from the wind measurement tower, and in conjunction with the terrain barrier coefficient... k block Set the difference threshold Δ v max =2× k block ×(1+ s / 30), of which s The slope of the target anemometer tower location; when Δ v >Δ v max If the location is not specified, it is marked as an abnormal spatial correlation; if there is no valid reference meteorological tower in the surrounding set area, the search range is automatically expanded.

4. The method according to claim 1, characterized in that, The cross-parameter physical correlation features include wind-pressure correlation features, air density features, and wind field stability features; wherein: Calculate the Pearson correlation coefficient between the pressure elevation gradient and wind speed of adjacent anemometer towers. r As a characteristic of wind-pressure correlation, the calculation formula is: r = Cov (D P / D h , v ) / [ σ (D P / D h )× σ ( v )]; in, Cov (Δ P / Δ h , v ) represents the covariance between the pressure gradient at altitude and the wind speed. σ (Δ P / Δ h ) represents the standard deviation of the barometric altitude gradient. σ ( v ) represents the standard deviation of wind speed; R is the ratio of the standard deviation of wind direction to wind speed. dv As a characteristic of wind field stability; The air density characteristic includes air density. ρ And wind energy density, air density ρ The calculation formula is: ρ =1.293×(273.15 / (273.15+ T ))×( P sea / 1013.25)×(1-0.378× U / 100); in, T To measure the actual temperature, P sea To correct air pressure at sea level, U Relative humidity; Wind energy density is ρ × v 3 .

5. The method according to claim 1, characterized in that, The temporal consistency detection includes: based on the standardized time-series data, setting the sliding window size, constructing a VAE model containing a 2-layer LSTM encoder and a 2-layer LSTM decoder, calculating the mean square error (MSE) between the real-time data sequence and the VAE reconstructed sequence; and setting an anomaly threshold (MSE). th =3×μ MSE , where μ MSE The mean MSE of the training set is given when MSE > MSE. th When this occurs, it is marked as a timing anomaly.

6. The method according to claim 1, characterized in that, The distribution offset detection includes: Using historical normal data from wind farms as samples, a Gaussian kernel function was employed. K ( x Construct a KDE model for each meteorological element, with the kernel function formula as follows: K ( x ) = ; in, The data consists of real-time meteorological observations of the wind farm to be verified. The bandwidth of the KDE model is determined using the following formula. : ; Where σ is the sample standard deviation, IQR is the sample interquartile range, n is the sample size, and min is the minimum value function; Real-time observation data x The probability density of its historical normal distribution is calculated using the KDE model. f ( x ): ; in, For the first i One historical normal sample; Set probability density threshold f th ,when f ( x )< f th When the weather is abnormal, it is marked as an anomaly; during periods of extreme weather, the probability density threshold is lowered. f th .

7. The method according to claim 1, characterized in that, A hybrid classification model, comprising a static feature extraction model, a temporal feature extraction model, and a dynamic weight fusion mechanism, is used to classify suspicious data. The static feature extraction model is selected from LightGBM, XGBoost, or Random Forest; the temporal feature extraction model is selected from bidirectional LSTM, GRU, or Transformer; and the dynamic weight fusion mechanism is an attention mechanism. The first layer extracts static features using a static feature extraction model and outputs a static feature classification probability vector. ; The second layer extracts temporal features through a temporal feature extraction model and outputs a temporal feature classification probability vector. P Bi-LSTM ( k ); The third layer calculates dynamic weights and completes probability fusion through a dynamic weight fusion mechanism, specifically including: Calculate static feature attention score score static : score static = ; in, The first learnable weight parameter, This represents the terrain complexity coefficient. =min(s / 30°,1), where s is the slope of the wind tower and H is the equipment health score; Calculate attention score for temporal features : ; in, This is the second learnable weight parameter; The static feature attention scores and temporal feature attention scores are weighted and normalized using the Softmax function, and the final classification probability fusion formula is determined as follows: ; in, For the final classification probability vector after fusion, take... As a preliminary classification label, The attention score is the normalized static feature. The attention score for the normalized temporal features; When the slope of the meteorological tower is greater than 30°, the tower should be raised. When the equipment health score H < 0.5, improve... .

8. The method according to claim 1, characterized in that, Perform differentiated repairs on different category labels, including: For sensor aging data, first calculate the aging deviation coefficient. β The formula is: β =1-[0.3×min(service life / 8, 1)+0.4×min(calibration interval / 12, 1)+0.3×min(historical number of failures / 10, 1)]; The service life is in years, the calibration interval is in months, and the historical failure count is the number of failures in the past year. β The value range is 0.5 to 1.0; The data is then corrected using a repair formula, which is: v 修复 = v 观测 × β +0.5×( v 参考 - v 观测 ); in, v 修复 The wind speed after repair. v 观测 The actual observed wind speed, v 参考 The wind speed is the same as that at a nearby reference station. The temperature correction formula is: T 修复 = T 观测 +0.3×( T 参考 - T 观测 )×(1- β ); in, T 修复 For the restored temperature, T 观测 The actual observed temperature. T 参考 Temperatures were within the normal range for this time of year. For short-term missing data with a duration of ≤1 hour, linear interpolation and fusion with neighboring data are used for data restoration. The restoration formula is as follows: v 最终 = v 时序插值 ×(1- w )+ v 邻近参考站 × w; in, v 最终 This is the final repair value for short-term missing data. v 时序插值 This is the result of linear interpolation of normal data before and after the period of missing data. w Dynamic weights for data from nearby reference wind towers. v 邻近参考站 This refers to the corresponding data from nearby reference stations during the same period; For transmission interference errors, multi-source weighted fusion interpolation is used for repair. The weight allocation rule is: weight of data from different layers of the same tower > weight of data from nearby reference wind measurement towers > weight of ERA5 reanalysis data. The weight ratios of the three types of data are 50%, 30%, and 20%, respectively, and are dynamically adjusted according to the validity of the data source.

9. The method according to claim 1 or the method according to claim 4, characterized in that, The entire process traceability information is stored using a data traceability chain in the form of a consortium blockchain constructed with blockchain technology. The traceability chain includes three types of consensus nodes: wind farm data center, operation and maintenance management platform, and third-party supervision agency. Each node synchronously stores a complete traceability copy containing the original data hash value, quality control and cleaning operation log, operation subject ID, and timestamp. The node consensus mechanism ensures that the data is tamper-proof and tamper-proof. The full-process traceability information supports multi-dimensional combined queries based on time range, meteorological tower ID, anomaly type, cleaning status, and repair algorithm type. The invalid data removal process includes: generating an invalid data removal list for data whose error exceeds the standard after repair, long-term missing data that cannot be repaired, invalid data after aging repair, and worthless data that cannot be clearly classified. After manual review and confirmation, the invalid data is transferred to an invalid data isolation library for storage, while retaining the original data. At the same time, a removal mark is made at the corresponding position in the normal dataset to retain a complete traceability link.

10. The method according to claim 1 or 5, characterized in that, After obtaining the standardized meteorological dataset after quality control, the method further includes: verifying the quality control effect based on the standardized meteorological dataset after quality control through four categories of wind farm indicators: data integrity, data accuracy, operational usability, and economic benefits; and performing dynamic iterations of incremental training, manual feedback optimization, and terrain and equipment adaptation updates based on the verification results, feedback data of the quality control process, and manual review results; wherein: The incremental training includes: updating the training data using a sliding time window every first cycle, retaining historical valid data, and performing incremental training on the hybrid classification model every second cycle; if extreme weather such as typhoons or thunderstorms occur during the second cycle, an additional incremental training is performed on the hybrid classification model. The manual feedback optimization includes: every second cycle, 500-1000 manually reviewed and corrected cases are fed back into the training set to optimize the parameters and feature weights of the hybrid classification model, requiring the monthly anomaly identification accuracy of the hybrid classification model to improve by more than or equal to 2%; The terrain and equipment adaptation update includes: updating the wind farm DEM data every third cycle, recalibrating the terrain barrier coefficient and slope-aspect correction function; and automatically reading terrain and equipment parameters when adding a new wind measurement tower or replacing a sensor to complete the adaptation update of the hybrid classification model.