A method for imputing missing wind measurement data that considers the trend and randomness of wind speed.

CN122173786BActive Publication Date: 2026-08-14CHINA POWER CONSRTUCTION GRP GUIYANG SURVEY & DESIGN INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-09
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

虽然这些既有技术在处理数据缺失问题上发挥了一定作用,但在应对复杂多变的风速特性时仍存在局限性

Benefits of technology

1、本发明通过构建基于滑动窗口的累积概率分布函数集合,并将原始测风塔风速数据转换为标准化概率序列,有效地将具有显著季节性和日变化趋势的风速物理量转化为平稳的概率空间数据,从而在插补过程中深度融合了风速的宏观演变趋势特征。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122173786B_ABST
    Figure CN122173786B_ABST
Patent Text Reader

Abstract

This invention relates to the field of wind measurement data processing, and particularly to a method for imputing missing wind measurement data that considers the trend and randomness of wind speed. The method includes: acquiring raw wind measurement data; constructing a set of cumulative probability distribution functions based on a sliding window to map wind speed to standardized values; constructing a Markov state space and calculating the transition probability matrix; performing probability sampling on the missing data segment to generate an imputation sequence; restoring the wind speed range through inverse function mapping and generating specific values; and correcting the results for bias using a particle swarm optimization algorithm. This invention can deeply integrate the macroscopic evolution trend and microscopic random fluctuation characteristics of wind speed, effectively suppressing trend drift in long sequence imputation, and improving the imputation accuracy and robustness under complex missing data conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wind measurement missing data interpolation technology, and more specifically, to a wind measurement missing data interpolation method that takes into account the trend and randomness of wind speed. Background Technology

[0002] The entire lifecycle of a wind farm, including early-stage planning and design, mid-stage construction, and later-stage operation and management, relies heavily on long-term, continuous, and high-quality wind speed observation data. This data is the core basis for accurate wind resource assessment, reliable power generation prediction, and scientific optimization of wind turbine layout. Therefore, ensuring the integrity and accuracy of wind measurement data is crucial for guaranteeing the economic benefits and operational safety of wind power projects. However, in actual observation processes, due to limitations such as the spatial and geographical distribution of wind measurement towers, complex measurement environments, and sensor equipment maintenance, raw wind measurement data often suffers from varying degrees of loss, resulting in incomplete data sequences. How to effectively interpolate wind measurement data to improve the temporal continuity of the data is a key issue in current wind power project research.

[0003] Currently, there are many methods for wind data interpolation, mainly including traditional mathematical interpolation methods such as linear interpolation, polynomial interpolation, and regression analysis; as well as methods based on artificial intelligence and machine learning, such as deep learning networks, random forest algorithms, and various time series models. While these existing technologies have played a role in addressing the problem of missing data, they still have limitations when dealing with the complex and variable characteristics of wind speed. For example, existing methods are not comprehensive enough in simultaneously capturing the temporal correlation, randomness, and macroscopic trend characteristics of wind speed. In practical applications, the interpolated data is prone to distortion when facing low-probability wind speed events, and when dealing with the interpolation requirements of ultra-long-term series, it is often accompanied by trend drift, causing the interpolated sequence to deviate from the physical evolution law of actual wind speed, affecting the accuracy of subsequent wind resource assessment.

[0004] In summary, existing technologies still have room for improvement in capturing multi-dimensional wind speed features and maintaining the stability of long-term evolution. Therefore, developing a wind measurement data interpolation method that can deeply integrate the trend and random characteristics of wind speed and effectively correct data distortion and trend drift has significant engineering application value for improving the scientific nature of wind farm construction and operation. Summary of the Invention

[0005] To address the aforementioned issues, this invention presents a wind speed data interpolation method that considers the trend and randomness of wind speed measurements. By constructing a cumulative probability distribution function that incorporates trends through seasonal standardization, and obtaining nonlinear wind speed data that considers temporal correlation and the randomness of low-probability events through Markov models and probability sampling, the method effectively characterizes the randomness, trend, and nonlinearity of wind speed data, thereby improving the reliability and accuracy of the interpolation results.

[0006] To achieve the invention objective of "effectively correcting data distortion and trend drift in wind measurement data interpolation," the wind measurement data interpolation method considering the trend and randomness of wind speed includes the following steps: acquiring original wind measurement data, time series standardization processing, Markov state space construction, state transition probability matrix calculation, probability sampling of missing sequences, reverse mapping of wind speed intervals, generation of specific wind speed values, evaluation of the interpolated sequence, and bias correction based on particle swarm optimization. Specifically, the present invention provides the following technical solution, including the following steps: S1. Collect the original wind speed data of the anemometer tower to be interpolated; S2. Calculate the cumulative probability distribution function of each time point in the original wind speed data of the meteorological tower by means of statistical frequency and sliding window, and denote it as the cumulative probability distribution function set. Then, convert the original wind speed data of the meteorological tower into a standardized time series according to the cumulative probability distribution function. S3. Discretize the standardized time series values ​​by partitioning them into intervals to obtain the Markov state set, and obtain the Markov chain based on the standardized time series values ​​and the state set. S4. Based on the Markov chain and its discrete state set (Markov state set), calculate the state transition probability matrix of the Markov chain; S5. Obtain the current state based on the previous state of the current value to be interpolated in the missing data segment. Then, find the probability distribution corresponding to the current state in the state transition probability matrix. Obtain the next state value through probability sampling. Iterate through all the values ​​to be interpolated in turn to obtain the missing data segment interpolation sequence. S6. Based on the timestamp of each interpolation state, find the corresponding cumulative probability distribution function in the set of cumulative probability distribution functions, calculate the inverse function value corresponding to each interpolation state, and thus obtain the wind speed interval sequence corresponding to each moment of the missing data segment. S7. Sample with the same probability within each wind speed interval of the wind speed interval sequence, and map each wind speed interval to a specific wind speed value to obtain an interpolated wind speed sequence; repeat S5-S7 to obtain multiple interpolated wind speed sequences. S8. Evaluate multiple interpolated wind speed sequences and determine the final missing data interpolation sequence based on the evaluation rules. Fill the missing data segment with the final missing data interpolation sequence.

[0007] As a preferred option, the calculation process of the cumulative probability distribution function incorporates short-term correlation analysis within the date extension time window. The span of the date extension time window is dynamically adjusted based on the distribution density of non-empty data points within the window. When the number of valid data points in the preset initial window is lower than a preset threshold, the date extension time window extends to both sides of the date until the number of valid data points covered reaches the preset calculation requirements.

[0008] The raw wind measurement data acquisition step involves retrieving the historical observation records of the wind measurement tower to be interpolated and arranging them in the order of timestamps to form the raw wind speed time series. The cumulative probability distribution function of the raw wind speed data of the original wind measurement tower at each corresponding moment in the raw wind speed time series is calculated by statistical frequency and sliding window method. A set of cumulative probability distribution functions is constructed, and the raw wind speed values ​​are mapped to standardized values ​​within a preset probability interval based on the cumulative probability distribution function.

[0009] Preferably, the formula for calculating the cumulative probability distribution function is: ; In the formula, This indicates the number of years covered by the dataset. This indicates the wind speed value. Indicates the first The year date is The time is The wind speed value, Indicates the length of the date slider window. Indicates an indicator function, This represents a timestamp. The timestamp is defined with precision down to the month, day, hour, and minute.

[0010] Preferably, in S2, the original wind speed data from the meteorological tower is converted into a standardized time series. The steps include: For the original wind speed data from the meteorological tower Each wind speed value Using the cumulative probability distribution function of its corresponding timestamp, calculate according to the following formula. Thus obtain ,in, This represents the total number of samples in the original anemometer tower wind speed data: ; in, Indicates the first time stamp of the entire period The cumulative probability distribution function of a timestamp, mathematically, The range of values ​​is within However, in the actual calculation of wind speed interpolation, it is stipulated that the cumulative probability is not zero. The core reason for this stipulation is that the distribution used for calculation is a finite-duration wind speed observation sample. If a certain wind speed value (e.g., 5 m / s) does not appear in this statistical sample, the empirical distribution calculated based on the sample will incorrectly determine that wind speed value as the 0th quantile, corresponding to a cumulative probability of 0, assuming "this wind speed cannot occur." However, in real meteorological scenarios, this wind speed value objectively exists and is entirely possible. To avoid the statistical bias caused by this finite sample from misleading the interpolation results, this invention mandates that the cumulative probability not be zero in subsequent wind speed interpolation calculation steps.

[0011] Preferably, continuous wind speed values ​​are mapped to a finite number of discrete "states," each corresponding to a wind speed interval, providing a discretization basis for subsequent Markov chain state transition analysis. Specifically, a set of states is generated based on the standardized time series value space, where the standardized time series value range is... Then at fixed intervals right By dividing the data into equal intervals, it can generate... Non-overlapping small intervals Round up to cover the entire range, where Among them, the first The range of values ​​for each interval is Each interval is denoted as By iterating through all the partitioned intervals, we obtain the discrete state set (Markov state set, Markov state space). Special note: The physical interpretations of discrete state sets, Markov state sets, and Markov state spaces are undisputed. The use of multiple names is a clear description of the data flow and transformation process in mathematical processing or their original mathematical expressions, and will not interfere with the understanding of those skilled in the art. The term "sequence" is used instead of "set" because of the trend and temporal nature of wind speed data.

[0012] Alternatively, within the full range of standardized values, a fixed interval is used to divide the data into multiple non-overlapping sub-intervals, and each sub-interval is defined as an independent Markov discrete state, forming a set of discrete states. Subsequently, based on the sub-interval to which the values ​​at each moment in the standardized time series belong, they are labeled as the corresponding discrete states, thus constructing a Markov state chain synchronized with the time series.

[0013] Specifically, construct the corresponding Markov chain: for any time in the standardized wind speed time series Observations By judgment The interval to which it belongs ,Sure state of time By iterating through all the observations at each time step and determining their state assignments, a temporal state sequence (Markov chain) is obtained. .

[0014] As a preferred option, calculate the state set in each state set. The transition probabilities of each state are independent of time. A Markov model assumes that the current state depends only on the previous state, and the transitions between states are described by these transition probabilities. Assume the Markov chain is... Markov state space The state transition probability formula is as follows: ; In actual sample data, the state transition probability can be obtained through frequency statistics, and the specific formula is as follows: ; in, Represents all of the Markov chains Momentary satisfaction Quantity, Represents all of the Markov chains Momentary satisfaction Quantity; Indicates the state value at the next moment. This represents the current state value. For example, if we're discussing wind speed, This indicates that the current value falls within the wind speed range. At that time, the wind speed value falls within the wind speed range at the next moment. The probability of.

[0015] The steps for calculating the transition probability matrix are based on the Markov state chain and its discrete state set. The frequency of state transitions is statistically analyzed. The transition probability between each state is determined by calculating the number of times the state chain transitions from any first state to any second state and dividing it by the total number of times the first state appears. This constructs a multi-dimensional state transition probability matrix, where the algebraic sum of the probabilities in each row of the matrix is ​​always equal to a unit value.

[0016] As a preferred approach, based on the Markov state space The current state value is calculated sequentially by counting frequencies. Transition to the next state value transition probability The state transition probability matrix can be obtained (e.g.) Figure 3Each row represents the probability of transitioning from the current state to any other state, and the sum of the probabilities in each row is 1.

[0017] Preferably, step S5 includes the following steps: S51, Based on missing data segments First value The timestamp query retrieves the state at the previous moment, and records it as a missing data segment. Current state The state value space and probability space at the next time step are obtained through the state transition probability matrix of the Markov chain, based on the current state. Query the state at the next time step from the state transition probability matrix of the Markov chain. value space (like Figure 3 The label values ​​of each column in the probability transition matrix, where Figure 3 Probability transition matrix interval (4%, with a total of 25 possible states) and probability space. (like Figure 3 probability transition matrix The column value corresponding to the column label in the same row, where the state and probability of these two corresponding positions are in a one-to-one correspondence. S52. Based on the correspondence between the state value space and the probability space described above, sample the states in the state value space with different probabilities to obtain the next state. Among them, probability sampling refers to selecting a sample from the population in a random manner. Each individual in the population has a non-zero, known probability of selection. This method can ensure the randomness and representativeness of the sample, so that the sample statistics can reasonably infer the population parameters. For example, when selecting balls from a box covered with a black cloth, it is known that there are 3 white balls and 1 red ball, but each time the ball is drawn, it may be a white ball or a red ball, which is uncertain. S53, Order Perform calculation steps S51-S52 sequentially until the missing data segment is obtained. The state sequence of all time steps of a segment is denoted as .

[0018] The missing sequence probability sampling step uses the known state of the previous moment before the starting position of the missing data segment as the starting point of the current state. It then retrieves the corresponding transition probability distribution from the state transition probability matrix and determines the state value for the next moment using a probability sampling algorithm. By sequentially traversing all time nodes within the cyclic missing time period, a state sequence of the missing data segment (also known as...) is generated. ).

[0019] As a preferred embodiment, S6 specifically includes: Because each state represents a certain interval to which the standardized time series value belongs, the state sequence of the known missing data segment... For ease of subsequent description, the standardized time series interval corresponding to each state of the missing data segment is denoted as . ,in, , This indicates the number of samples with consecutive missing data in the missing segment (i.e., the length of the missing segment). For each missing imputation state, the cumulative probability distribution function for the corresponding timestamp is found in the set of cumulative probability distribution functions. Then, the cumulative probability distribution function is calculated for each standardized time series region. The inverse function value is used to calculate the wind speed range corresponding to each state. , .

[0020] The wind speed interval reverse mapping step retrieves the corresponding function relationship in the cumulative probability distribution function set based on the timestamp corresponding to each state value in the discrete state interpolation sequence; by solving the inverse function value of the cumulative probability distribution function with respect to each standardized sub-interval, the wind speed physical quantity interval corresponding to the missing data segment at each time moment is calculated, forming a wind speed interval sequence.

[0021] As a preferred method, the wind speed space is obtained by using the inverse function of the cumulative probability distribution function, specifically calculated using the following formula: ; In the formula, Represents the set of cumulative probability distribution functions. The inverse function of the cumulative probability distribution of timestamps; Here, the specific logic for determining the wind speed range using the inverse function of the cumulative probability distribution function is as follows: substitute the left and right boundaries of the standardized sub-intervals into the inverse function of the cumulative probability distribution corresponding to the timestamp, and the output values ​​are used as the lower and upper limits of the wind speed range.

[0022] As a preferred option, in order to Each interval is mapped to a specific wind speed value. Specifically, it is assumed that the wind speed in each interval follows a uniform distribution, that is, the values ​​in the wind speed interval at each time moment are sampled with the same probability to obtain the wind speed interpolation value at each time moment. The specific wind speed value generation steps are as follows: the wind speed distribution within each wind speed interval follows a uniform distribution law; by performing equal probability sampling within the wind speed interval corresponding to each time moment, the interval-form data is converted into specific wind speed values, generating a complete interpolated wind speed sequence; by repeatedly executing the state sampling, interval mapping and value generation process, multiple candidate interpolated wind speed sequences are obtained.

[0023] As a preferred embodiment, in S8, multiple interpolated wind speed sequences are evaluated, and the final missing data interpolation sequence is determined based on the evaluation rules. The final missing data interpolation sequence is then filled into the missing data segment. Based on the absolute error between the average value of each interpolated wind speed sequence and the average value of the wind speed sequence corresponding to the preset quantile. The evaluation will be conducted using the following specific calculation formula: ; ; In the formula, Indicates the first In S7, to improve the reliability of the interpolated values ​​and reduce the uncertainty of probability sampling, the absolute error of the interpolation sequence is determined by iteratively performing state sampling, interval mapping, and numerical generation steps multiple times during the interpolation of missing data segments. Each iteration generates an interpolation sequence that conforms to the Markov property until the missing data segment to be interpolated is obtained. The iteration stops at the interpolated sequences, denoted as _ ... Therefore Indicates the first The average of all points in the interpolated sequence yes The cumulative distribution function The quantiles are selected empirically based on test results.

[0024] The interpolation sequence evaluation step is based on the statistical characteristics of the candidate interpolation wind speed sequences for optimization: the average value of each candidate interpolation wind speed sequence is calculated, and the average value of the historical wind speed sequence corresponding to the preset quantile is calculated. By comparing the absolute errors between the two, all absolute errors are iterated, and the interpolation wind speed sequence corresponding to the smallest absolute error is selected. If the absolute error of this interpolated wind speed sequence is less than or equal to the preset maximum permissible error, then this interpolated wind speed sequence is the final missing data interpolation sequence. If the absolute error of this interpolated wind speed sequence is greater than the preset maximum permissible error, it will be adjusted and corrected using quantile mean and particle swarm optimization to obtain the final missing data interpolation sequence.

[0025] As a preferred approach, in reality, the interpolated values ​​may exhibit trend drift due to the long interpolation time series. To bring the wind speed series back to normal levels, a particle swarm optimization algorithm is used to find the optimal wind speed adjustment coefficient that makes the mean of the interpolated wind speed series equal to a given mean. Assume the wind speed series to be interpolated is... The adjusted wind speed sequence is The optimal solution and the adjusted wind speed calculation formula are as follows: in, This represents the optimal wind speed adjustment coefficient. This indicates the wind speed adjustment coefficient. This represents the interpolated wind speed sequence (based on the previous assumptions, this calculation is only to illustrate the principle and does not conflict with the aforementioned physical interpretation). This means taking the average of the results after subtracting the points from each of the two sequences. This represents the wind speed sequence corresponding to the 50th quantile at each time point in the missing data segment. And the aforementioned They are different. The aforementioned approach, to avoid inaccuracies caused by the randomness of probability sampling, involves repeatedly interpolating for a given consecutive missing data segment to find the interpolation sequence with the smallest absolute error, which is then used as the optimal interpolation sequence for the current missing data segment. However, this optimal interpolation sequence, although iterated... While the result is optimal, trend drift may occur. Therefore, after imputing the missing data segments, the particle swarm optimization is needed to adjust the coefficients of the imputed missing data segments, i.e., to make linear modifications, in order to obtain the imputed sequence that is closest to the 50th percentile sequence.

[0026] The deviation correction step based on the particle swarm optimization algorithm performs a secondary adjustment on the interpolation results where the absolute error exceeds the preset maximum allowable error. When the absolute error exceeds the preset maximum allowable deviation range, the particle swarm optimization program is started to find the optimal wind speed adjustment coefficient. The particle swarm optimization program uses the wind speed adjustment coefficient as the search space, takes the degree of convergence between the mean of the initial interpolation sequence and the mean of the preset quantile as the objective function, and determines the final correction coefficient by iteratively updating the position and velocity of the particles.

[0027] The final missing data imputation sequence is obtained by multiplying the optimal wind speed adjustment coefficient by each wind speed value in the imputation result before correction to obtain the corrected sequence, and then filling the missing positions of the original data with this sequence.

[0028] As a preferred approach, the particle swarm optimization (PSO) algorithm generates a number of particles in each iteration, each representing a potential solution (here, the wind speed adjustment coefficient). The particles move within the solution space, updating their state by tracking two core optimal solutions: the individual optimal solution found during its own iterations, and the global optimal solution shared by all particles. Each particle updates its velocity and position according to a fixed formula, continuously moving closer to the optimal solution until a preset number of iterations is reached or the accuracy requirement is met, at which point the optimization stops and the global optimal solution is output. The key parameter settings are as follows: Number of iterations: Set to 50-200 iterations based on the sample size of the wind measurement data. For larger sample sizes (e.g., tens of thousands of data points), the number of iterations can be appropriately reduced to balance optimization accuracy and computational efficiency.

[0029] Number of particles: 20-50. Too few particles can easily lead to local optima, while too many will increase computation time. Usually, 30 particles are used as a baseline, which can be fine-tuned according to the fluctuation of wind speed data.

[0030] Inertia weight: A linear decreasing strategy is adopted, with an initial value of 0.9, which is gradually reduced to 0.4 during iteration. This setting can balance the particle's global exploration ability in the early stage and its local development ability in the later stage, adapting to the trend and random characteristics of wind speed.

[0031] Learning factors: Both the individual learning factor (c1) and the global learning factor (c2) are set to 2.0. This value promotes rapid convergence of particles towards their own optimal and collective optimal directions, improving optimization stability.

[0032] During the execution of the particle swarm optimization algorithm, its search parameters are configured as follows: the iterative process adopts a linearly decreasing inertial weight strategy, which enables the algorithm to have high global search capability in the early stage and high local convergence accuracy in the later stage; at the same time, the individual experience learning factor and the global social learning factor of the particles are both set to fixed constants to balance individual exploration and group cooperation in the search process.

[0033] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention constructs a set of cumulative probability distribution functions based on a sliding window and converts the original wind speed data from the meteorological tower into a standardized probability sequence. This effectively transforms the physical quantity of wind speed, which has significant seasonal and diurnal variation trends, into stable probability space data, thereby deeply integrating the macroscopic evolution trend characteristics of wind speed during the interpolation process.

[0034] 2. This invention utilizes Markov chains and transition probability matrices to model the randomness of standardized time series. By generating interpolation states through probability sampling, it can accurately capture the random fluctuations and temporal correlations of wind speed over short time scales, thus solving the data distortion problem that traditional interpolation methods are prone to when dealing with random changes in wind speed.

[0035] 3. This invention uses an inverse function mapping mechanism to restore the interpolation results in the probability space to the physical wind speed space, and combines it with the particle swarm optimization algorithm to dynamically correct the mean deviation of long sequences, effectively suppressing the trend drift phenomenon commonly seen in ultra-long time sequence interpolation, and ensuring that the interpolated sequence maintains a high degree of consistency with historical observation data in terms of statistical characteristics.

[0036] 4. This invention solves the problem of inaccurate estimation of probability distribution function when the original data is severely missing by adaptively adjusting the length of the extended date time window. It improves the robustness and interpolation accuracy of the algorithm under extremely complex missing conditions, and provides more reliable data support for wind resource assessment. Attached Figure Description

[0037] Figure 1 This is a flowchart of wind measurement data interpolation; Figure 2 This is a schematic diagram of the cumulative probability distribution function; Figure 3 This is a diagram illustrating how the numerical values ​​of the state transition probability matrix represent the transition probabilities from the current state to the next state. Figure 4 This is a comparison chart of quantiles and actual wind speeds at different times; Figure 5 This is a diagram illustrating the variable values ​​generated during the interpolation process; Figure 6 This is a schematic diagram of Markov imputation results for short missing sequences; Figure 7 This is a schematic diagram of the Markov imputation result for a long missing sequence. Detailed Implementation

[0038] Example 1: As Figures 1-7 As shown, this invention proposes a method for imputing missing wind speed data that considers the trend and randomness of wind speed. It uses wind speed data from all years of a specific wind tower to impute missing data. Its core idea is to use the state transition probability matrix of a Markov chain and random sampling to impute missing data.

[0039] Specifically, the first step is to use seasonal standardization to obtain the cumulative probability distribution function of the known data (e.g., Figure 2 This figure illustrates the cumulative probability distribution of wind speed at a certain moment. Using this CDF (Concurrent Distribution Function), the actual wind speed can be converted into a cumulative probability value. Furthermore, the original wind speed value is transformed into a standardized time series using the cumulative probability distribution function. Then, based on the value range of the standardized time series, the Markov state space and Markov chain are obtained, and the state transition probability matrix is ​​calculated (e.g., ...). Figure 3 The state transition probability matrix represents the probability of transitioning from the current state to the next state; the redder the color (or the larger the value), the higher the transition probability. This matrix quantifies the transition pattern of wind speed states (e.g., the transition probability from "4%-8%" state to "8%-12%" state is 20.7%), reflecting the Markov property of wind speed changes (the current state is determined only by the previous state), and serves as a "rule base" for imputing missing data. Finally, probability sampling is performed based on the state transition probability matrix, and uniform sampling maps the wind speed intervals corresponding to the Markov states to individual wind speed values, thus obtaining the imputed data. Through multiple rounds of sampling and imputation, multiple imputed sequences are obtained, which are then evaluated and corrected to gradually approximate the true data (e.g., ...). Figure 4-7 ).

[0040] Figure 4This is a comparison of quantiles and actual wind speeds at different times (the blue series in the figure represents the wind speed curves at different quantiles (20%, 40%, 60%, 80%), and the black series represents the actual wind speed curves. Most of the actual wind speed values ​​fluctuate within the range of each quantile curve, indicating that quantiles can be used as a reference benchmark for the "trend" of wind speed; at the same time, during some periods (such as around 0:00-4:00), the actual wind speed is significantly lower than the 20% quantile curve, dropping to the 1-2 m / s range. These low wind speed situations occur infrequently and are considered low-probability events).

[0041] Figure 5 These are the variable values ​​generated during the interpolation process (these are the variable values ​​generated during the interpolation process of the wind tower data; the entire interpolation process involves randomly generating missing values, generating a standardized time series, obtaining discrete states, sampling based on the probability matrix to obtain Markov interpolated states, obtaining wind speed ranges based on the state intervals, and then sampling to obtain the corresponding wind speeds). Probability sampling refers to selecting samples randomly from the population. Each individual in the population has a non-zero, known probability of selection. This method ensures the randomness and representativeness of the sample, allowing the sample statistics to reasonably infer population parameters.

[0042] Figure 6 This is the Markov imputation result for short missing sequences (here, the missing sequence refers to a short consecutive missing sequence; the horizontal axis represents the number of missing samples, and the vertical axis represents the wind speed value). As can be seen from the figure, at most sample points, the difference between the Markov imputation value and the true value is small, indicating a high degree of fit between the imputation curve and the true curve. This demonstrates that the method effectively preserves the trend and randomness of wind speed for short-term missing data. It should be noted that the plotting logic is as follows: First, the original wind measurement time series (original wind speed data from the wind tower) is traversed, and all short missing segments with consecutive missing values ​​not exceeding 70 samples are selected. Then, a representative missing segment is selected, and its original timestamps are mapped sequentially to sample numbers 1, 2, 3... at fixed observation intervals (e.g., 10 minutes) as the horizontal axis, with the imputed wind speed value and the true wind speed value at the corresponding time point plotted on the vertical axis.

[0043] Figure 7This is the Markov imputation result for a long missing sequence (here, the missing sequence refers to a continuous missing sequence with a relatively long length (missing for half a month). The horizontal axis represents the number of missing samples, and the vertical axis represents the wind speed value. As can be seen from the figure, even when faced with a long continuous missing sequence (half a month), the imputation curve can stably capture the long-term trend of wind speed and also exhibits random fluctuations). It should be noted that the plotting logic of this figure is as follows: a continuous period was extracted from the time series of a certain wind measurement tower. During this period, a sensor failure / data loss of more than 10 days occurred. Therefore, the first 300 points in the figure contain both true values ​​and short-term missing values, followed by imputation values ​​for more than 10 days.

[0044] It is achievable, such as Figure 1 As shown, the present invention may include the following steps: Step 101, Data Collection and Parameter Definition. Instance data is constructed by collecting wind speed data from the anemometer tower, clarifying the range of missing data segments and core parameters (used for subsequent probability distribution calculations, state division, and accuracy control). Specifically, wind speed data from all years of the anemometer tower to be interpolated are collected, arranged in chronological order, and denoted as [missing data]. A segment of consecutive missing data is denoted as , This represents the total number of samples of the original anemometer wind speed data. This represents the number of missing data segments. The number of samples in the missing data segment, with a separator length of . ,and For integers, the date extended time window is... The maximum number of interpolation operations is The standard quantile is The maximum permissible absolute error is ; Step 102: Calculate the cumulative probability distribution for each timestamp. By using a sliding time window, the wind speed probability distribution for each moment is statistically analyzed, preserving both long-term trends (data from the same time over many years) and incorporating short-term correlations (data from adjacent dates). Specifically, a sliding window is first used to obtain the samples needed to calculate the cumulative probability distribution function for each timestamp. If the number of non-empty data points in the window is too small, the window is expanded along the date dimension of the year until it contains at least 50 valid data points. Then, the cumulative probability distribution function for each timestamp in the time series data is calculated by statistical frequency counting, resulting in a set of cumulative probability distribution functions. The cumulative probability distribution function for each timestamp (e.g., ...) is used to calculate the cumulative probability distribution function for each timestamp. Figure 2 The specific calculation formula is as follows (as shown): ; in, For the number of years covered by the dataset, This is the wind speed value. Indicates the first The year date is The time is The wind speed value, Indicates the length of the date slider window. It is an indicator function, if If true, the value is 1; otherwise, the value is 0. Step 103: Generate standardized time series. (Through...) The original wind speed is converted into a cumulative probability value to eliminate seasonal and diurnal wind speed magnitude differences (such as differences between summer and winter wind speed magnitudes, and differences between daytime and nighttime wind speeds), retaining only the randomness feature so that the Markov model can capture the randomness of wind speed in the future.

[0045] Specifically, this is achieved by calculating the wind speed value for each timestamp. The cumulative probability value corresponding to the cumulative probability distribution function To obtain the standardized time series The specific calculation formula is as follows: ; In the formula, Indicates the first The cumulative probability distribution function of each timestamp, in wind speed interpolation, is considered to be... The range of values ​​is within between; Step 104: Generate the Markov state set (discrete state set). The range of values ​​for the standardized time series. Discretizing the states into equidistant intervals transforms continuous probability values ​​into discrete states, facilitating the analysis of state transition patterns using Markov chains.

[0046] Specifically, in Inner Intervals can be generated Then the first small interval, The range of values ​​for each interval is , recorded as By iterating through the discrete states sequentially, we can obtain the set of discrete states. Here, each state represents a division interval of the standardized time series; Step 105, obtain the Markov chain. According to standardized time series Each value in the set belongs to a range in the state set. The value of the range is transformed into the corresponding state, forming a state transition sequence (Markov chain) that changes over time, reflecting the temporal correlation of Markov states.

[0047] Specifically, for standardized time series Each observation value By determining which interval a value belongs to, we can obtain its state, denoted as . By iterating through the data and making judgments sequentially, we can obtain the Markov chain as follows: .

[0048] Step 106, calculate the state transition probability matrix. Statistically calculate the "from state" transition probability matrix in the Markov chain. Transition to state The frequency of "" is used to obtain the transition probability, quantifying the transition pattern between states. To distinguish between time and state values, let , This represents the state values ​​at two different times. , The values ​​all belong to the discrete state set.

[0049] Specifically, calculation arrive To calculate the transition probabilities, first calculate all pairs of values ​​in the Markov chain that are equal to... Number of samples Then calculate for all consecutive time intervals in the Markov chain where the current time interval equals The next moment equals Number of samples Finally, based on the probability calculation formula, the state is obtained. Transition to state transition probability The specific calculation formula is as follows: ; Traverse sequentially The probability of each state transitioning to another state in the Markov state set is calculated, and the state transition probability matrix is ​​finally obtained. Here, the probability of the wind speed at the next moment being in a different wind speed range is calculated given the current wind speed value.

[0050] Step 107: Determine the initial state of the missing data segment. According to the "no aftereffect" property of Markov chains, the first state of the missing data segment is determined by its state at the previous moment, i.e., according to... First value The timestamp query retrieves the state at the previous moment, and records it as a missing data segment. initial state ; Step 108: Obtain the value and probability space of the next state. Based on the current state, query the transition probability matrix to obtain the possible next states and their corresponding probabilities, capturing the dynamic behavior of low-probability wind speed events. Obtain the state value space and probability space for the next time step through the state transition probability matrix, based on the current state... Query the state at the next time step from the state transition probability matrix of the Markov chain. value space (like Figure 3 The label values ​​of each column in the probability transition matrix, where, Figure 3 Probability transition matrix interval (4%, with a total of 25 possible states) and probability space. (like Figure 3 In the probability transition matrix The column value corresponding to the column label in the same row, for example, if ,but ); Step 109: Sampling to obtain the next state. Random sampling is performed according to the probability space distribution to ensure that low-probability states (corresponding to extreme wind speeds) have a chance to occur, preserving the randomness of wind speed. Specifically, based on the correspondence between the state value space and the probability space in step 108, states in the state value space are sampled with different probabilities to obtain the next state. ; Step 110: Generate the state sequence for imputing missing data segments. Iteratively execute "current state - query probability - sample next state" to obtain the state sequence for the complete missing data segment, reflecting the continuous transition pattern of wind speed state. Specifically, let... Perform calculations 108-109 sequentially until the missing data segment is obtained. The state sequence at all times is denoted as ; Step 111: Convert the state to a wind speed range. Using the inverse function of the CDF, map the state to the actual wind speed range.

[0051] Specifically, firstly, based on the first Timestamp of each state In the set of cumulative probability distribution functions Find the cumulative probability distribution function corresponding to its timestamp. Since each state represents a division interval of the standardized time series, for ease of subsequent description, the imputed sequence of missing states will be used. ,use In fact, ,in, Indicates that in the known number of... Time equals First, based on the quantile intervals obtained from sampling using the probability transition matrix, then, by calculating the cumulative probability distribution function for each... The inverse function value yields the wind speed range for each state, denoted as . , The specific calculation formula is as follows: ; Let represent the inverse function of the cumulative probability distribution of the i-th timestamp; Step 112: Sampling to obtain the imputed wind speed sequence. Uniform sampling is performed within each wind speed interval, converting the interval into specific wind speed values ​​to generate a complete imputed sequence for missing data segments. Specifically, this is achieved by sampling within each interval... Sampling with equal probability, Each interval in the interval is mapped to a specific wind speed value, denoted as . ,and, By iterating through the data, the wind speed value at each moment can be obtained, which is the interpolated wind speed sequence, denoted as... ; Step 113, Generate A series of interpolated wind speed sequences are obtained. To improve the reliability of the interpolated wind speed sequences, steps 107-112 are performed iteratively multiple times. In each iteration, a probability sampling is performed to generate a sample sequence that conforms to the Markov property, until the missing data segment is obtained. of The interpolated wind speed sequences are denoted as follows: ; Step 114, Calculate absolute error of each interpolated sequence The specific formula for calculating the error is as follows: ; ; In the formula, Indicates the first The absolute error of each interpolated sequence, Indicates the first The average of all points in the interpolated sequence Indicates the first The inverse function of the cumulative probability distribution of timestamps. yes The cumulative distribution function quantiles (e.g.) Figure 4 As shown), You can select based on experience according to the test results; the default is... ; Step 115: Select the optimal sequence and correct it.

[0052] Iterate through all absolute errors and select the interpolated wind speed sequence corresponding to the smallest absolute error. If the absolute error of this interpolated wind speed sequence is less than or equal to the preset maximum permissible error, then this interpolated wind speed sequence is the final missing data interpolation sequence. If the absolute error of this interpolated wind speed sequence is greater than the preset maximum permissible error Then, adjustments and corrections are made based on the quantile mean and the particle swarm optimization algorithm. Particle swarm optimization parameters, such as the number of iterations and the number of particles, are set, and the algorithm is used to iteratively optimize and find the match between the quantile mean and the particle swarm optimization algorithm. A wind speed sequence whose overall mean is increasingly close, or one that uses a certain number of iterations as a stopping condition, is considered the optimal interpolated wind speed sequence. The final missing data imputation sequence is denoted as ; The particle swarm optimization algorithm is used to find the optimal wind speed adjustment coefficient that makes the mean of the wind speed sequence equal to a given mean. Assume the interpolated wind speed sequence is... The adjusted wind speed sequence is The optimal solution and the adjusted wind speed calculation formula are as follows: in, This represents the optimal wind speed adjustment coefficient. This indicates the wind speed adjustment coefficient. This represents the interpolated wind speed sequence (based on the previous assumptions, this calculation is only to illustrate the principle and does not conflict with the aforementioned physical interpretation). This means taking the average of the results after subtracting the points from each of the two sequences. This represents the wind speed sequence corresponding to the 50th percentile at each moment in the missing data segment.

[0053] Step 116: Complete imputation of all missing data segments. Iterate through each consecutive missing data segment, integrating all imputation results into the original sequence to obtain complete wind speed data. Specifically, let... That is, when starting to process the next missing data segment, repeat steps 107-115 until... That is, the traversal stops after all consecutive missing data segments have been imputed. 1 is the iteration variable of the outer loop. The algorithm will iteratively process each missing data segment. Additionally, steps 107-115 can be a complete sub-process for processing a single missing data segment (depending on the number of missing data segments included in the manually set time interval). In this sub-process, using... The fixed variable refers to the amount generated by the currently being processed missing data segment, and the outer loop variable. 1 Correspondingly, when 1 change, with The quantity of the index also needs to be updated.

[0054] Example 2: In one embodiment of the present invention, the raw wind measurement data acquisition step S1 first preprocesses the raw database collected by the wind measurement tower. The preprocessing process includes removing obvious outliers, such as wind speed values ​​exceeding physical limits or long-term unchanging dead values. The processed data is aligned according to a time step of minutes or hours, and missing parts are marked as null values, forming the basic sequence to be processed.

[0055] In performing the time series standardization step S2, the core of this invention lies in constructing a probabilistic model that reflects the change of wind speed over time. Since wind speed exhibits distinct annual and daily cycle characteristics, this invention calculates the cumulative probability distribution function corresponding to each timestamp by setting a sliding window on the date axis. In this case, the cumulative probability distribution function for each timestamp represents the periodic statistical regularity of change at each moment. This statistical method based on a "date-time" dual-dimensional window essentially condenses the seasonal and diurnal trends of wind speed into the distribution function.

[0056] In the Markov state space construction step S3, the present invention constructs a continuous probability space. Divided into The number of discrete intervals is used. The fineness of this division directly affects the interpolation result's ability to reproduce the original wind speed fluctuations. If the division is too coarse, it will fail to reflect subtle changes in wind speed; if the division is too fine, the transition probability matrix will contain a large number of zero elements, leading to sampling failure. This invention uses an equidistant division method, dividing each interval into discrete intervals. This is considered as a state. Subsequently, each value in the standardized time series is mapped to a time series state sequence based on the interval index it falls into. This process transforms the complex nonlinear time series prediction problem into a discrete state transition probability problem.

[0057] In step S4 of the transition probability matrix calculation, the present invention iterates through the temporal state sequence. The frequency of state transitions between all adjacent time points is counted. The state transition probability can be obtained through frequency statistics. This matrix comprehensively depicts the statistical pattern of wind speed fluctuations at this wind measurement point, that is, the probability that the wind speed will shift to another probability level at the next time point when the current wind speed is at a certain probability level.

[0058] In the missing sequence probability sampling step S5, this invention employs the concept of Monte Carlo simulation. For each missing point, the algorithm does not provide a definite predicted value, but instead randomly samples from the corresponding row of the probability matrix based on the state at the previous time step. This sampling mechanism preserves the random nature of wind speed, making the interpolated sequence highly similar to the true wind speed in second-order statistical properties such as power spectral density and autocorrelation function. To improve the stability of interpolation, this invention generates multiple candidate sequences through repeated sampling, and selects the optimal solution through subsequent evaluation steps.

[0059] In the wind speed range reverse mapping step S6 and the specific wind speed value generation step S7, this invention performs an inverse transformation of "probability-physical quantity". This is achieved through the inverse function... The probability interval boundaries obtained from sampling are transformed into wind speed interval boundaries. This step is crucial to ensuring that the interpolated values ​​conform to the local and current climate characteristics. For example, at the same probability level, the mapped wind speed in winter is usually higher than in summer, which reflects the trend in the interpolation results. Finally, uniform sampling is performed within the defined wind speed intervals to eliminate the step effect caused by discretization, ensuring that the final generated wind speed sequence has continuous physical characteristics.

[0060] In step S8 of the interpolated sequence evaluation, this invention introduces an evaluation criterion based on quantiles. Since wind speed distribution typically exhibits Weibull distribution characteristics, simple mean comparisons are insufficient to reflect the quality of long sequences. This invention calculates the historical average level corresponding to a preset quantile (such as the median) as a benchmark to measure whether the interpolated sequence has experienced trend drift. Selecting the sequence with the smallest absolute error ensures the accuracy of the interpolation results in terms of macroscopic statistical distribution.

[0061] In the deviation correction step S9 based on the particle swarm optimization algorithm, this invention considers that although Markov sampling can maintain local randomness, when dealing with extremely long missing data segments, the accumulated random error may cause the overall mean of the sequence to deviate from the historical expectation. Therefore, this invention designs a linear correction operator sequence. Each particle in the particle swarm optimization algorithm represents a set of possible values. The mechanism by which the particle swarm optimization algorithm finds the optimal wind speed adjustment coefficient is as follows: A group of candidate particles is randomly generated in the solution space, with each particle representing a set of wind speed adjustment coefficient sequences. During optimization, the proximity of the mean of the adjusted sequence to the mean of the 50th percentile of the historical data of the missing segment is used as the main evaluation criterion. At the same time, in order to prevent point-by-point correction from destroying the original fluctuation texture of the sequence, a smoothing constraint is applied to the adjustment coefficients at adjacent times, that is, the absolute value of the difference between adjacent coefficients does not exceed a preset threshold (preferably 0.05-0.10), and each coefficient is restricted to a preset reasonable range (preferably [0.8, 1.2]).

[0062] During the iteration process, each particle tracks its own historical best coefficient sequence and the group's global best coefficient sequence, dynamically adjusting its search direction and step size. A strategy of initially large and then decreasing inertial weights is adopted, allowing the algorithm to focus on global exploration in the early stages and on refined local search in the later stages, until the preset number of iterations or accuracy requirements are reached, at which point the algorithm stops.

[0063] In the final optimal adjustment coefficient sequence, the coefficients at each time point fluctuate slightly and gradually around the baseline level. This is achieved by both pulling the mean of the interpolated sequence back to the historical expected level through the overall statistical characteristics of the coefficient sequence and ensuring that the scaling ratios of adjacent time points are basically consistent through smoothness constraints. Thus, while correcting trend drift, it preserves the random texture and relative fluctuation patterns captured by Markov sampling. Through this heuristic optimization, this invention can fine-tune the interpolation results, correct potential systematic biases, and ultimately obtain interpolated data that conforms to both short-term fluctuation patterns and long-term statistical distributions.

[0064] In summary, this invention decomposes wind speed into trend components (through... By combining the capture of random components (captured via Markov chains) and utilizing particle swarm optimization for global bias control, a complete and self-correcting wind measurement missing data interpolation system was constructed. This method does not rely on external reference station data and can complete high-quality interpolation based solely on single-tower historical records.

[0065] In a further embodiment of the present invention, the dynamic adjustment mechanism of the date extension time window is specifically manifested as follows: the algorithm first attempts to collect samples with a small date radius; if the sample size is insufficient to construct a smooth time window... If the curve (e.g., encountering data gaps for several consecutive years due to extreme weather) is affected, the radius is automatically expanded. This design ensures that the cumulative probability distribution function is sufficiently representative in any missing environment, avoiding probability estimation distortion caused by insufficient samples.

[0066] Furthermore, the transition probability matrix can be constructed using a higher-order Markov model, which considers not only the impact of the current moment on the next moment but also the combined effects of multiple previous moments, further enhancing the ability to characterize complex wind speed evolution patterns. In the particle swarm optimization stage, the correction coefficients can also be extended to piecewise correction functions, providing differentiated compensation for deviations in different wind speed ranges, thereby achieving higher-precision interpolation correction across the entire range.

[0067] Through the synergistic effect of the aforementioned technical means, this invention can generate interpolation sequences that highly match the original data in terms of time-domain fluctuations, frequency-domain characteristics, and probability distribution dimensions, providing a solid data foundation for the accurate site selection and efficient operation of wind farms. The method of this invention has a clear flow, high computational efficiency, and can handle missing data of various lengths and distributions, showing broad prospects for application.

[0068] Example 3: To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present invention and are not intended to limit the present invention.

[0069] In wind energy resource assessment and wind farm operation management, the completeness of wind measurement data is fundamental to ensuring assessment accuracy and operational efficiency. However, due to factors such as communication failures, sensor icing, equipment power outages, or regular maintenance, data gaps are inevitable in the original wind measurement data sequence. Existing linear interpolation, mean interpolation, or spatial interpolation methods based on nearby sites often struggle to simultaneously account for both the macroscopic trends (such as seasonality and diurnal variation) and microscopic randomness (such as gusts and instantaneous fluctuations) of wind speed, resulting in significant deviations in statistical characteristics between the interpolated data and the actual observed values.

[0070] This invention provides a method for imputing missing wind speed data that considers the trend and randomness of wind speed. By constructing a standardized probability space, the complex non-stationary wind speed sequence is transformed into a stationary Markov state chain. Furthermore, a heuristic optimization algorithm is used to correct the bias, thereby achieving high-quality data recovery.

[0071] The first step is to acquire raw wind measurement data. In this step, the system connects to the wind tower data management center and retrieves historical observation records from the wind towers to be processed. These records typically include sampled data from multiple dimensions such as wind speed, wind direction, temperature, and air pressure. The system first preprocesses this data, including removing outliers (such as negative values ​​or values ​​exceeding the local limit wind speed) and dead values ​​that have remained unchanged for a long time. Subsequently, the system arranges the data at equal intervals according to the observation timestamps, marking any times for which no data was collected as null values, forming a raw wind speed time series containing missing data segments. This series serves as the basic input for all subsequent calculations, and its completeness directly affects the extraction of statistical features.

[0072] Next, the time series standardization process (step 2) is initiated. This is the core step in processing trend characteristics in this invention. Since wind speed exhibits significant seasonal and diurnal cycles, directly modeling the original physical quantity would lead to overly complex model parameters. This invention maps wind speed values ​​to a probability space by calculating a cumulative probability distribution function. For each timestamp in the original sequence, the system searches for corresponding sample points in the historical database based on its month, date, and specific time (accurate to the minute). To ensure statistical validity, this invention introduces a date-extended time window. This window slides left and right on the date axis, for example, extending forward and backward by several days from the current date, and collecting similar samples across multiple years. This sliding window-based statistical method can capture the distribution pattern of wind speed within a specific time period. When the distribution density of effective sample points within the initial window is low, the system automatically initiates a dynamic adjustment mechanism, extending the search range to both sides of the date axis until the preset computational requirements are met.

[0073] In the specific calculation logic of the cumulative probability distribution function, the system uses an indicator function to iterate through the samples. For each historical wind speed sample at a specific timestamp, the number of samples with values ​​less than or equal to the current observation is counted, and this number is divided by the total number of samples within that window to obtain the cumulative probability of the wind speed value at that moment. In this way, the original wind speed time series is transformed into a standardized probability series with values ​​ranging from zero to one. This standardized time series eliminates macroscopic trends caused by seasonal and diurnal variations, exhibiting stronger stationarity and laying the foundation for subsequent Markov modeling.

[0074] The Markov state space construction step (step 3) is then executed. Markov models require discrete input data, therefore the standardized probability sequence needs to be discretized. The system divides the probability range from zero to one at predetermined fixed intervals. For example, the entire range is divided into several non-overlapping sub-intervals, each representing a relative wind speed intensity. These sub-intervals are defined as independent Markov discrete states. By matching each value in the standardized time series with its corresponding sub-interval, the system constructs a Markov state chain synchronized with the original time axis. This chain records the evolution trajectory of wind speed in the probability space, and its state changes reflect the inherent laws governing wind speed fluctuations.

[0075] After obtaining the Markov state chain, the system calculates the transition probability matrix (step 4). This step aims to quantify the probability of wind speed transitions between different states. The system traverses the entire state chain, counting the frequency of transitions from any state to another. The state transition probability is calculated by dividing the transition frequency by the total number of times the initial state occurs. This ultimately constructs a multi-dimensional state transition probability matrix. Each row of this matrix corresponds to an initial state, and the sum of all elements in the row is always equal to the unit value of one, representing the probability distribution of transitioning to all possible states at the next moment, given the current state. This matrix is ​​the core basis for subsequent probability sampling, deeply characterizing the randomness of wind speed fluctuations at this wind measurement point.

[0076] Step 5 involves probabilistic sampling of the missing sequence. When encountering a missing data segment in the original data, the system uses the known state of the moment preceding the missing point as the initial state and retrieves the corresponding transition probability distribution from the transition probability matrix. The system then uses a probabilistic sampling algorithm to randomly determine the state value at the first missing moment based on this distribution. Subsequently, using this generated predicted state as a benchmark, the system continues to deduce the state at the next moment until the entire missing time period is filled. Through this iterative sampling, the system generates a discrete state imputation sequence. This method generates a sequence that is not only numerically random but also conforms to the physical laws of wind speed evolution in terms of temporal logic. To enhance the robustness of the imputation results, the system typically performs multiple sampling operations in parallel, generating multiple candidate discrete state sequences.

[0077] Then, a reverse mapping of the wind speed interval is performed (step 6). Since the interpolation result generated in the previous steps is in a discrete state space, it needs to be restored back to the physical wind speed space. The system retrieves the corresponding cumulative probability distribution function based on the timestamp of the interpolation moment. For each discrete state, there is a left boundary and a right boundary of the standardized probability interval. The system maps these two probability boundary values ​​back to the corresponding physical wind speed boundary by solving the inverse function of the cumulative probability distribution function. In this way, the discrete state at each moment is transformed into a specific wind speed value interval. This reverse mapping mechanism ensures that the interpolated wind speed value can automatically adapt to the local and current trend characteristics. For example, in seasons with high wind speeds, the same probability level will automatically map to a higher physical wind speed value.

[0078] Next, the specific wind speed value generation step (step 7) is executed. Within the defined wind speed range, to eliminate the stepped effect caused by discretization, the system assumes that the wind speed distribution within the range follows a uniform distribution law. By sampling with equal probability within the range, the system converts the range data into continuous specific wind speed values. This process completes the final transformation from a discrete state to a continuous physical quantity. At this point, the system has obtained multiple complete candidate interpolated wind speed sequences. These sequences exhibit fluctuation textures similar to the real wind speed on a short time scale and conform to historical statistical patterns on a macro scale.

[0079] To select the optimal solution from multiple candidate sequences, the system also performs imputation sequence evaluation (step 8). The core criterion for evaluation is the statistical consistency of the sequences. The system calculates the average value of each candidate imputation sequence and compares it with the historical average wind speed for the same period at a preset quantile. Here, the quantile mean represents the expected level of wind speed for that specific period. The system calculates the absolute error between the two, iterates through all absolute errors, and selects the candidate sequence with the smallest absolute error as the initial imputation result. If the absolute error of this imputation wind speed sequence is less than or equal to the preset maximum permissible error, then this imputation wind speed sequence is the final missing data imputation sequence; if the absolute error of this imputation wind speed sequence is greater than the preset maximum permissible error, then adjustments and corrections are made based on the quantile mean and particle swarm optimization algorithm to obtain the final missing data imputation sequence. This screening process ensures that the imputation sequence maintains a high degree of consistency with historical observations in terms of both total quantity and average level.

[0080] Finally, to further eliminate potential trend drift in long-sequence interpolation, the system executes a deviation correction step based on the particle swarm optimization algorithm (step 9). This step is a closed-loop optimization process. When the absolute error of the initial interpolation result exceeds a preset tolerance range, the system initiates the particle swarm optimization program. In the particle swarm optimization algorithm, each particle represents a potential wind speed adjustment coefficient. These particles move freely within the search space; their position determines the magnitude of the adjustment coefficient, and their velocity determines the search direction and step size.

[0081] During optimization, the system multiplies the initial imputation sequence by the adjustment coefficients represented by the particles to obtain the corrected sequence. The system uses the closeness between the mean of the corrected sequence and the historical expected mean as the fitness function. Particles continuously update their state based on their own historical best position and the global best position of the population. To balance search efficiency and accuracy, the algorithm employs a linearly decreasing inertia weight strategy. In the early stages of iteration, a larger inertia weight gives particles a stronger global exploration ability, helping them escape local optima; in the later stages of iteration, a smaller inertia weight enhances local convergence accuracy, enabling precise locking of the optimal adjustment coefficients. Through multiple iterations, the system determines the final correction coefficients and applies them to the initial imputation sequence to obtain the final missing data imputation result. This result is then filled back into the missing positions of the original data, completing the entire imputation process.

[0082] In practical applications, for example, a wind farm might experience extreme cold weather in winter, causing the sensors on its wind measurement tower to freeze and resulting in a continuous data gap of up to three days. If traditional linear interpolation is used, the interpolated data will be a smooth diagonal line, completely losing the random fluctuations in wind speed, leading to a significant deviation from the actual wind energy resource assessment results. However, using the method of this invention, the system first obtains historical data before and after the missing data segment in step 1. In step 2, the system retrieves wind speed samples from the wind measurement tower over the past five years for the same dates and time periods, constructing a refined cumulative probability distribution function.

[0083] In steps 3 and 4, the system analyzes the wind speed evolution pattern at this location during winter. For example, if the wind speed is currently high (high probability state), the transition probability matrix shows the probability distribution of whether it will remain high or slowly decrease in the next moment. During the sampling process in step 5, the discrete state chain generated by the system can simulate the suddenness of gusts and the intermittency of wind speed. Through the reverse mapping in steps 6 and 7, these probability states are restored to high wind speed physical values ​​that conform to the characteristics of winter climate.

[0084] Because the missing period is as long as three days, the cumulative effect of random sampling may cause the mean of the interpolated sequence to be slightly higher or lower than the average winter level for that location. In this case, step 8 removes sequences with significant deviations from hundreds of candidate sequences. Subsequently, step 9 uses a particle swarm optimization algorithm to fine-tune the size of each wind speed point, ensuring that the interpolated sequence for these three days maintains its original fluctuation pattern while accurately matching the statistical characteristics of the same historical period in terms of total energy and average wind speed. The final interpolated sequence is visually indistinguishable from the actual observation data and exhibits extremely high consistency in power spectrum analysis and frequency distribution statistics.

[0085] The steps in this invention are closely coordinated and logically necessary. Standardization step 2 addresses non-stationarity through probability mapping, enabling Markov modeling steps 3 and 4 to operate efficiently within a stationary space. The random states generated in sampling step 5 are regressed to the physical space through back mapping step 6 and generation step 7, achieving a balance between randomness and trend. Evaluation step 8 and correction step 9 serve as quality control mechanisms, ensuring the reliability of the final output data.

[0086] Furthermore, the dynamic date extension window used in calculating the cumulative probability distribution function in this invention greatly enhances the algorithm's adaptability to small sample sizes. In scenarios with newly built meteorological towers or short data records, if the samples within a fixed window are insufficient, the system can still obtain a representative statistical distribution by automatically extending the date range, avoiding systematic biases in the interpolation results. This adaptive feature makes this method highly versatile under different geographical environments and climatic conditions.

[0087] The application of particle swarm optimization (PSO) in bias correction not only solves the mean shift problem but also optimizes the overall distribution of the sequence through a global search of the adjustment coefficients. Compared with traditional fixed-proportion correction, PSO can find the optimal balance point in a nonlinear solution space based on the feedback of the objective function. This heuristic search method is particularly effective in handling extremely long missing data segments (such as data missing for several consecutive weeks), effectively suppressing the "wandering" phenomenon caused by random sampling and greatly improving accuracy.

[0088] Through the aforementioned complete technical chain, this invention can not only fill in the gaps in individual data points, but also handle long-distance, large-area data repair tasks. The interpolated data sequence can be directly used for refined calculation of wind power generation, wind turbine load simulation, and wind farm power prediction. Because this method is entirely based on the statistical characteristics of the meteorological tower itself and does not rely on satellite data or data from surrounding meteorological stations, it has extremely high independent application value in remote areas or wind farms lacking reference stations.

[0089] During system operation, the data flow at each step exhibits strict temporal order. After the original sequence is input, it undergoes a standardization transformation into the probability domain, where stochastic evolution simulation is performed. Subsequently, it reverts to the physical domain via inverse transformation, and finally, global optimization completes the output. This "from physics to probability and back to physics" processing architecture is the fundamental reason why this invention can simultaneously accommodate multiple wind speed characteristics. The calculation results of each step provide necessary constraints and guidance for the next step, collectively forming a high-precision wind measurement data recovery system.

[0090] In practical software implementation, this method can be integrated into the preprocessing module of wind resource assessment software. Users only need to upload the raw data file, and the system can automatically execute the entire process from step 1 to step 9. Regarding computational resource allocation, since the calculation of the transition probability matrix and particle swarm optimization involve numerous iterations, the system can employ multi-threaded parallel processing technology, significantly improving the processing speed of long data sequences. This efficient computational architecture ensures the timeliness of the method when processing massive amounts of wind measurement data, providing strong technical support for the rapid advancement of wind power projects.

[0091] This invention, through the comprehensive application of cumulative probability distribution, Markov chains, and particle swarm optimization, successfully solves the challenge of balancing trend and randomness in wind measurement data interpolation. The synergistic cooperation between each step ensures that the interpolation results achieve a high level of simulation in both statistical and physical terms, providing a high-quality, highly reliable data source for subsequent wind energy resource development and refined wind farm management.

[0092] Example 4: In order to enable those skilled in the art to fully understand and implement the present invention, the specific implementation principle of the present invention is further supplemented below with a specific application scenario.

[0093] In this embodiment, a 100-meter-high wind measurement tower at a large mountain wind farm in northwestern China is used as the application object. During the transition from spring to summer, the tower experienced a severe sandstorm, which physically blocked the ultrasonic anemometer sensor probe, resulting in a data gap of 120 consecutive hours (5 days). Since this period coincides with a critical time for wind farm output fluctuations, failure to accurately recover the missing data would directly lead to serious deviations in the monthly wind resource assessment report for that month.

[0094] Step 1: Acquiring raw wind measurement data. First, the historical time series of the wind measurement tower before and after the fault is retrieved via the remote interface of the Supervisory Control and Data Acquisition (SCADA) system. This step not only extracts wind speed values ​​but also simultaneously acquires auxiliary parameters related to wind speed, such as air temperature, air pressure, and turbulence intensity. During processing, the system identifies all records from 00:00 on May 12th to 23:50 on May 16th as invalid values. The controller performs time-stamp alignment on the raw series, standardizing the sampling interval to one data point every 10 minutes, thus providing a standard, equidistant time benchmark for subsequent statistical analysis. During this process, the system automatically removes pulse anomalies generated at the moment of the fault, ensuring the physical authenticity of the input data.

[0095] Step 2: During time series standardization, the system retrieves observation samples for the same month from the historical database of the past eight years for the specific seasonal window of May. The core principle of this step is to project the non-stationary wind speed physical quantity into a uniformly distributed probability space using a cumulative probability distribution function. Specifically, for each missing time point (e.g., 10:00 AM on May 12th), the system automatically opens a dynamic date extension window centered on that time point and spanning 15 days before and after it. By searching for all valid observations within this window in the historical samples, the system constructs a local wind speed probability distribution model for that specific time point. By calculating the percentile ranking of each historical wind speed point in the sample set, the system achieves a transformation from wind speed value to... Nonlinear mapping of interval probability values. This processing method can significantly eliminate the changes in macroscopic wind speed trends caused by temperature rises and air pressure fluctuations during the spring and summer seasons, resulting in a processed sequence exhibiting stable time series characteristics.

[0096] Step 3: During the construction of the Markov state space, the system divides the standardized probability space into 20 equally spaced discrete state intervals. This step involves... The intervals are divided with a step size of 0.05, and each interval corresponds to a specific relative wind speed intensity level. For example, state 1 represents a wind speed in an extremely low probability interval, while state 20 represents a wind speed in an extremely high probability interval. The system assigns each value in the historical standardized time series to the corresponding state interval, thereby discretizing the continuous probabilistic evolution process into a Markov state chain. This discretization process can effectively filter out high-frequency noise in the wind speed signal while retaining the core evolutionary logic of wind speed fluctuations, providing structured data for subsequent stochastic modeling.

[0097] Step four, when calculating the state transition probability matrix, involves traversing the entire Markov state chain and counting the frequency of state transitions between adjacent time points. This step constructs a 20×20 matrix to record the conditional probability of wind speed transitioning from the current state to the next. For example, in mountainous wind fields, due to the funneling effect caused by the terrain, wind speed often exhibits strong persistence, which is represented in the matrix by higher values ​​for diagonal elements (i.e., the probability of the state remaining unchanged). For moments with frequent gusts, the elements far from the diagonal in the matrix record the probability distribution of sudden wind speed changes. In this way, the system solidifies the unique wind speed fluctuation pattern at that location into a set of mathematical parameters, achieving precise quantification of the randomness of wind speed.

[0098] Step 5: When performing probability sampling of missing sequences, the system uses the last known state at 23:50 on May 11th as the starting point for iteration. This step utilizes the Monte Carlo sampling principle to extract the corresponding probability distribution vector from the state transition probability matrix and generate random numbers for state selection. To simulate the diversity of wind speed fluctuations, the system generates 500 independent interpolation state sequences in parallel. Each sequence exhibits different fluctuation paths at the microscale, but all follow the physical constraints defined by the transition probability matrix in terms of macroscopic statistical characteristics. This multi-path sampling technique can effectively cover low-probability wind speed events, avoiding the overly smoothing problem of traditional interpolation methods, thus achieving a high degree of restoration of the instantaneous randomness of wind speed.

[0099] Step Six: During the reverse mapping of wind speed intervals, the system re-maps the sampled discrete states back to the physical wind speed intervals. This step utilizes the inverse function of the cumulative probability distribution function constructed in Step Two to find the corresponding physical wind speed boundary based on the specific time of the interpolation point. For example, if the sampled state at a certain time is state 15, and the corresponding probability interval is (0.70, 0.75], the system will calculate, based on the historical CDF curve at noon on May 13th, that the physical wind speed corresponding to this probability interval is likely between 8.5 m / s and 9.2 m / s. This dynamic reverse mapping based on timestamps ensures that the interpolated values ​​can automatically inherit the local diurnal variation patterns and seasonal trends, effectively injecting trend characteristics.

[0100] Step 7: When generating specific wind speed values, the system introduces uniformly distributed random perturbations within the mapped wind speed range. This step determines the final continuous wind speed values ​​by performing secondary sampling within the range of 8.5 m / s to 9.2 m / s. This step eliminates the numerical staircase effect caused by discrete states, resulting in a natural smoothness and fluctuation in the interpolated wind speed sequence over time. By reconstructing 500 candidate sequences, the system obtained 500 complete physical wind speed interpolation schemes, which are fully compatible with the original observation data in terms of physical dimensions.

[0101] Step 8: When performing the interpolation sequence evaluation, the system introduces multidimensional statistical verification indicators. This step first calculates the average wind speed for each candidate sequence within the 5-day missing data period and compares it with the expected average wind speed for the wind farm in mid-May over the past 8 years. Simultaneously, the system also performs consistency checks on the sequence's variance, skewness, and kurtosis. By calculating the comprehensive statistical distance between the candidate sequences and historical baseline sequences, the system automatically eliminates sequences that, while conforming to randomness, significantly deviate from the long-term climate average. Finally, the sequence with the smallest statistical distance is selected for the final fine-tuning stage, thus ensuring the statistical rigor of the interpolation results.

[0102] Step Nine: During deviation correction based on the particle swarm optimization algorithm, the system performs global optimization on the selected initial interpolation sequence. In this step, the wind speed adjustment coefficient at each moment is defined as the particle's position vector. During optimization, the controller sets the objective function to minimize the sum of squared residuals between the daily average wind speed of the interpolation sequence and the historical daily average wind speed for the same period. Each particle in the particle swarm searches within a 50-dimensional space (corresponding to key time nodes in the interpolation segment), continuously updating its optimal individual and group positions to find correction parameters that best approximate the true overall energy distribution of the sequence. Through this heuristic optimization, the system can automatically correct trend drift caused by the long-term operation of the Markov chain. For example, if the overall energy is too high due to sampling, the particle swarm optimization algorithm will automatically fine-tune the amplitude at each point, bringing the simulated values ​​of missing data from the past 5 days back to a reasonable range.

[0103] In this application scenario, through the coordinated operation of the above nine steps, the final interpolated sequence filled the data gaps from May 12th to 16th. Compared to traditional linear interpolation, this invention can reconstruct the drastic wind speed fluctuations before and after the sandstorm; compared to a single deep learning model, this invention ensures the interpretability of the data in terms of physical logic through Markov matrices and CDF mapping, avoiding the physically meaningless extreme values ​​that may be generated by black-box models.

[0104] Furthermore, during execution, the system addresses the sample sparsity issue caused by the short duration of mountain wind field observation records by dynamically adjusting the length of the date extension window. When the number of historical samples falls below 100 at a given moment, the system automatically extends the window from 30 days to 45 days, ensuring the smoothness of the CDF curve. This adaptive mechanism guarantees the robustness of standardization step 2 under any complex climatic conditions. Simultaneously, the linearly decreasing inertial weights introduced by the particle swarm optimization algorithm enable the correction process to quickly cover potential deviations in the initial stages and accurately locate the optimal correction value in the later stages, significantly improving computational efficiency.

[0105] Building upon this foundation, the present invention achieves excellent interpolation results by combining discrete state transitions in the probability space with continuous distribution restoration in the physical space. The principle behind this effect lies in the fact that the Markov chain captures the short-term memory (i.e., randomness) of wind speed, while the cumulative probability distribution function locks in the long-term evolution (i.e., trend) of wind speed. These two are coupled through a reverse mapping step 6, ensuring that the final output data incorporates both.

[0106] Furthermore, the closed-loop feedback mechanism employed in step 9 of the deviation correction process in this invention can automatically adjust the optimization strategy for missing data segments of different lengths. For short-term (e.g., several hours) missing data, the system primarily relies on the inertia of the Markov chain for recovery; for long-term (e.g., several days or weeks) missing data, the system focuses on using the particle swarm optimization algorithm for total quantity control. This scaled processing logic effectively solves the "mean regression failure" problem that easily occurs in existing technologies when processing long sequence imputation, significantly improving the accuracy of energy density calculation in wind resource assessment.

[0107] All content not described in detail in this specification belongs to existing technology known to those skilled in the art, and the specific parameters of each algorithm, such as the number of Markov states and the size of the particle swarm, can be routinely adjusted according to the sampling frequency and environmental complexity of the actual wind measurement tower, and can be run using a conventional high-performance computing server. Furthermore, the various embodiments in this case share the same principles, although the specific descriptions differ slightly (especially in mathematical notation). Each embodiment is not mutually exclusive or contradictory; the different descriptions are extensions of the principles, and the differences are intended to assist those skilled in the art in understanding and implementation. They should not be construed as unclear or insufficient disclosure in this specification. It is hereby stated that the probability and statistical calculations and heuristic optimization logic involved in this technical solution, being mature mathematical methods combined with the unique physical constraints of this invention, are existing technologies at the software code level and therefore are not described in detail.

Claims

1. A method for imputing missing wind measurement data considering the trend and randomness of wind speed, characterized in that: S1. Collect the original wind speed data of the anemometer tower to be interpolated; S2. Calculate the cumulative probability distribution function of each time point in the original wind speed data of the meteorological tower by statistical frequency and sliding window method, construct a set of cumulative probability distribution functions, and then convert the original wind speed data of the meteorological tower into a standardized time series according to the cumulative probability distribution function. In S2, the calculation of the cumulative probability distribution function takes into account the short-term correlation within the date extended time window, and the length of the date extended time window is adaptively adjusted according to the number of non-empty data points in the window. When the number of non-empty data points in the preset window is less than the preset number, the length of the date extended time window is increased until it contains at least the preset number of valid data points. S3. Discretize the standardized time series values ​​by partitioning them into intervals to obtain the Markov state set, and obtain the Markov chain based on the standardized time series values ​​and the state set. S4. Based on the Markov chain and its Markov state set, calculate the state transition probability matrix of the Markov chain; S5. Obtain the current state based on the previous state of the current value to be interpolated in the missing data segment, find the probability distribution corresponding to the current state in the state transition probability matrix, obtain the next state value through probability sampling, and iterate through all the values ​​to be interpolated in turn to obtain the missing data segment interpolation sequence. S6. Based on the timestamp of each interpolation state, find the corresponding cumulative probability distribution function in the set of cumulative probability distribution functions, calculate the inverse function value corresponding to each interpolation state, and thus obtain the wind speed interval sequence corresponding to each moment of the missing data segment. S7. Sample with the same probability in each wind speed interval of the wind speed interval sequence, and map each wind speed interval to a specific wind speed value to obtain an interpolated wind speed sequence; repeat S5-S7 to obtain multiple interpolated wind speed sequences. S8. Evaluate multiple interpolated wind speed sequences and determine the final missing data interpolation sequence based on the evaluation rules. Fill the missing data segment with the final missing data interpolation sequence. The evaluation rules specifically include: Iterate through all absolute errors and select the interpolated wind speed sequence corresponding to the smallest absolute error. If the absolute error of this interpolated wind speed sequence is less than or equal to the preset maximum permissible error, then this interpolated wind speed sequence is the final missing data interpolation sequence. If the absolute error of this interpolated wind speed sequence is greater than the preset maximum allowable error, then adjustments and corrections are made based on the quantile mean and particle swarm optimization algorithm to obtain the final missing data interpolation sequence. In step S8, the evaluation of multiple interpolated wind speed sequences includes evaluating the absolute error between the average value of the interpolated wind speed sequences and the average value of the wind speed sequences corresponding to preset quantiles. The specific calculation formula is as follows: ; ; In the formula, Indicates the first The absolute error of each interpolated sequence, Indicates the first The average of all points in the interpolated sequence yes The cumulative distribution function Quantiles Indicates the number of samples in the missing data segment; Adjustments and corrections are made based on quantile mean and particle swarm optimization algorithm. The specific calculation formula is as follows: in, This represents the optimal wind speed adjustment coefficient. This indicates the wind speed adjustment coefficient. This represents the interpolated wind speed sequence. This means taking the average of the results after subtracting the points from each of the two sequences. This represents the wind speed sequence corresponding to the 50th quantile at each time point in the missing data segment. This represents the adjusted wind speed sequence.

2. The wind speed missing data interpolation method considering the trend and randomness of wind speed according to claim 1, characterized in that, The formula for calculating the cumulative probability distribution function is: ; In the formula, This indicates the number of years covered by the dataset. This indicates the wind speed value. Indicates the first The year date is The time is The wind speed value, Indicates the length of the date slider window. Indicates an indicator function, Represents a timestamp.

3. The wind speed missing data interpolation method considering the trend and randomness of wind speed according to claim 1, characterized in that, In step S2, the step of converting the original wind speed data from the meteorological tower into a standardized time series includes: For each wind speed value in the original anemometer tower wind speed data, the cumulative probability value is calculated using the cumulative probability distribution function of its corresponding timestamp according to the following formula, resulting in the set of cumulative probability distribution functions: ; in, Indicates the first The cumulative probability distribution function of wind speed values ​​corresponding to each timestamp The range of values ​​is between, This represents the total number of samples of the original anemometer wind speed data.

4. The wind speed missing data interpolation method considering the trend and randomness of wind speed according to claim 1, characterized in that, In step S3, the discretization of the standardized time series values ​​through interval partitioning is specifically as follows: Based on the standardized time series value range, at fixed intervals Perform equidistant division to generate A number of non-overlapping small intervals, among which Round up to cover the entire range; The range of values ​​for each interval is ,in Each interval is denoted as By iterating through all the partitioned intervals, we obtain the Markov state set.

5. The wind speed missing data interpolation method considering the trend and randomness of wind speed according to claim 4, characterized in that, The specific process of constructing a Markov chain is as follows: For any moment in the standardized wind speed time series Observations By judgment The interval to which it belongs Determine the state at time t The observations at all times are iterated sequentially and their state assignments are determined to obtain the time-series state sequence.

6. The method for interpolating missing wind measurement data considering the trend and randomness of wind speed according to claim 1, characterized in that, S4 specifically includes: The current state is calculated by counting frequencies based on the Markov state set. Transition to the next state value The transition probabilities are used to construct a multi-dimensional state transition probability matrix.

7. The wind speed missing data interpolation method considering the trend and randomness of wind speed according to claim 6, characterized in that, The formula for calculating the state transition probability is as follows: ; In the sample data, the state transition probability is obtained through frequency statistics, and the specific formula is as follows: ; in, Represents all of the Markov chains Momentary satisfaction Quantity, Represents all of the Markov chains Momentary satisfaction Quantity; Indicates the state value at the next moment. This indicates the state value at the current moment.

8. The method for interpolating missing wind measurement data considering the trend and randomness of wind speed according to claim 1, characterized in that, S5 specifically includes: S51. Query the status at the previous moment based on the timestamp of the first value of the missing data segment. The state value space and probability space at the next time step are obtained through the state transition probability matrix, and the state value space at the next time step is queried from the state transition probability matrix. and probability space; S52. Based on the correspondence between the state value space and the probability space, sample the state value space with different probabilities to obtain the next state. ; S53, Order Then, perform calculation steps S51-S52 sequentially until the state sequence of all times in the missing data segment is obtained.

9. The method for interpolating missing wind measurement data considering the trend and randomness of wind speed according to claim 1, characterized in that, Specifically, S6 includes: Based on the standardized time series intervals corresponding to the state sequence of the missing data segment, the cumulative probability distribution function corresponding to the timestamp of each imputed state is found in the cumulative probability distribution function set. By calculating the inverse function value of the cumulative probability distribution function with respect to each standardized time series interval, the wind speed interval sequence corresponding to each moment of the missing data segment is obtained.

Citation Information

Patent Citations

  • Method for supplementing missing measurement data of anemometer tower

    CN120974065A

  • Sewage treatment time sequence data missing value interpolation method based on random configuration network

    CN121434715A