Wind power plant short-term wind power prediction method based on data mode pre-judgment
By using differentiated encoding and adaptive learning mechanisms for historical multidimensional feature data, a frequency-balanced spatiotemporal encoding matrix is constructed. Combined with a three-level index and a composite prediction model, the real-time performance and accuracy issues of short-term wind power prediction for wind farms are resolved, achieving high-precision short-term wind power prediction for wind farms.
Patent Information
- Application Number
- CN202510904115.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-31
AI Technical Summary
Existing methods for short-term wind power forecasting in wind farms are insufficient in handling complex spatiotemporal data characteristics and dynamically adapting to weather changes and unit operating status, making it difficult to meet real-time and accuracy requirements.
By using differentiated encoding and adaptive learning mechanisms for historical multidimensional feature data, a frequency-balanced spatiotemporal encoding matrix is constructed. Combined with a three-level index and a composite prediction model, high-precision short-term wind power prediction for wind farms is achieved.
It achieves high-precision prediction in complex terrain and sudden weather scenarios, reducing prediction error by 30%-50%, meeting the real-time requirements of power grid dispatch, and has adaptive capabilities.
Smart Images

Figure CN120875128A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of short-term wind power prediction technology for wind farms, and specifically to a method for short-term wind power prediction for wind farms based on data pattern prediction. Background Technology
[0002] In the field of wind power generation, accurate short-term wind power forecasting is crucial for grid dispatching, optimized unit control, and improving wind energy utilization efficiency. Existing technologies are mainly divided into two categories: physical model methods based on numerical weather prediction (NWP) and statistical model methods based on historical data, but both have significant drawbacks.
[0003] Firstly, physical modeling methods based on numerical weather prediction rely on mesoscale weather forecast data provided by meteorological departments and simulate the local effects of wind farms through fluid dynamics models. For example, the Spanish SIPREOLICO system and the commonly used dynamic downscaling methods in China require complex three-dimensional spatial wind flow field modeling and directional calculations to generate wind acceleration factors and wind direction angles. However, these methods suffer from the following problems: they require hundreds of directional calculations for different wind directions and thermal stability, resulting in a lengthy process for generating the basic database, making it difficult to meet the real-time requirements of short-term forecasts; they cannot simultaneously consider the comprehensive impact of atmospheric thermal stability, wind direction, and other parameters on wind power, especially in complex terrain where near-surface atmospheric vertical motion is intense, leading to a significant decrease in forecast accuracy; and they are sensitive to meteorological data resolution and initial conditions, lacking a dynamic correction mechanism, making it difficult to cope with sudden changes in meteorological conditions.
[0004] Secondly, there are data-driven statistical modeling methods, which use machine learning algorithms to uncover spatiotemporal correlations in historical data. For example, the widely used WPFSVER1.0 system and the model from North China Electric Power University rely heavily on SCADA data and numerical weather prediction inputs. However, these methods suffer from the following problems: model performance is affected by the quality of meteorological data sources and the accuracy of SCADA data acquisition, and they lack robustness to anomalous data; the decision-making process of core algorithms (such as LSTM and GRU) is difficult to interpret, parameter optimization relies on empirical tuning, resulting in long training cycles and high computational costs; and they are slow to respond to emerging meteorological models or changes in generator operating conditions, leading to accumulated prediction errors over time.
[0005] Therefore, there is an urgent need for a short-term wind power prediction method that can efficiently process multi-dimensional feature data, accurately match historical patterns, and has adaptive capabilities. Summary of the Invention
[0006] The purpose of this invention is to provide a method for predicting short-term wind power of wind farms based on data pattern prediction. By differentially encoding, extracting patterns and performing correlation analysis on historical multidimensional feature data, combined with an adaptive learning mechanism, a high-precision prediction of short-term wind power of wind farms can be achieved.
[0007] The specific technical solution of the present invention is as follows:
[0008] One of the technical solutions of this invention is to provide a method for short-term wind power prediction of wind farms based on data pattern prediction, comprising:
[0009] Historical multidimensional feature data is collected and divided into core group and edge group according to the frequency of occurrence of historical multidimensional feature data. The core group generates fine-grained coding by statistically analyzing extreme values through a sliding window, while the edge group generates sparse coding based on event duration and intensity. An adaptive weighting algorithm is combined to balance the density of the two types of coding and construct a frequency-balanced spatiotemporal coding matrix.
[0010] The spatiotemporal coding matrix is divided into segments according to time windows. The intensive pattern and envelope pattern of each segment are extracted. Based on the envelope pattern, a three-level index of weather type, unit topology and time dynamics is constructed. The two types of patterns are stored in the historical pattern library. A fast retrieval structure is established based on the three-level index.
[0011] The system extracts intensive and envelope patterns from real-time data, matches them against a historical pattern library using a three-level index, performs inclusive judgment, and classifies them into four types of association patterns: strong association, moderate association, weak association, and invalid association.
[0012] Analyze strong, moderate, and weak associations, and make predictions based on hierarchical weighted fusion of association patterns;
[0013] Prediction is made through a composite prediction model, and the prediction results are compared with the actual results in real time. If a certain type of pattern fails continuously, the core group encoding decomposition is triggered, the weights are redistributed according to the feature contribution and the three-level index is updated to achieve adaptive learning.
[0014] As a further improvement to this method, a dynamic joint frequency statistics and density clustering fusion strategy is adopted when dividing the core group and the edge group:
[0015] Statistically analyze the joint frequency of multidimensional feature combinations and identify high-frequency feature combinations as core group candidates;
[0016] Candidate feature combinations are grouped a second time using a density-based clustering algorithm to remove noise points and generate the final core group and edge group division results.
[0017] As a further improvement to this method, the core group generates fine-grained codes using a variable step-size sliding window technique. The window length is adaptively adjusted according to the wind speed fluctuation characteristics, with a sliding step of 10 minutes and a window length of 1 hour. Within each window, the maximum value, minimum value, extreme value timestamp, and trend slope are extracted as encoding features.
[0018] The edge group generates sparse coding using event intensity-time coding technology. Events are detected by a dual-threshold hysteresis comparator. For the detected events, event parameters are extracted and combined into a three-dimensional sparse vector using a three-dimensional design.
[0019] As a further improvement to this method, the adaptive weighting algorithm's strategy for balancing the two types of coding densities includes:
[0020] Calculate the density of fine-grained coding and sparse coding;
[0021] Among them, the density of fine-grained coding is the number of codes per unit time, and the density of sparse coding is the frequency of occurrence of low-frequency events.
[0022] Weights are dynamically allocated based on the density ratio of the encoding.
[0023] The dynamic weight allocation includes reducing the weight of the core group and increasing the weight of the edge group proportionally if the density of the core group is higher than the frequency of the edge group.
[0024] As a further improvement to this method, the intensive pattern extraction adopts a hierarchical redundancy removal and dynamic clustering fusion method, the steps of which include:
[0025] The first layer filters out high-volatility features by using a variance threshold, and removes feature columns with variances below a preset threshold.
[0026] The second layer uses Pearson correlation coefficient analysis to merge highly linearly correlated feature groups while retaining representative features.
[0027] K-means dynamic clustering is performed on the simplified feature vectors to generate cluster centers as the intensive pattern, and the density parameters within the clusters are statistically analyzed.
[0028] The construction of the extreme boundary of the envelope pattern includes spatiotemporal correlation coding, and the steps include:
[0029] Extract the core feature extrema at the beginning and end of the time window to generate a four-dimensional boundary vector [min start ,max start ,min end ,max end ];
[0030] For edge group events, add a binary marker vector to record the event type and the number of occurrences, and concatenate it with the core boundary vector to form a hybrid envelope code.
[0031] As a further improvement to this method, when constructing the three-level index of meteorological type, unit topology, and time dynamics:
[0032] The meteorological type index is divided into layers based on wind speed range, temperature range, and pressure gradient.
[0033] The unit topology index generates a unique code based on the wind turbine model, location layout, and grid connection relationship;
[0034] The time-based dynamic index associates dynamic time tags with seasons, day / night cycles, and weather cycles.
[0035] As a further improvement to this method, the criteria for classifying the inclusion judgment into strong association, moderate association, weak association, and invalid association are as follows:
[0036] The core criterion for judgment is that the real-time encoding completely falls within the envelope boundary of the historical pattern library for a certain time segment and the feature density difference is ≤ the feature density threshold.
[0037] The core criteria for judging moderate correlation are that the offset between the real-time encoding and the historical pattern encoding group is ≤ the offset threshold and the difference in feature density is ≤ the feature density threshold.
[0038] The core criteria for weak association judgment are that the offset between the real-time encoding and the historical pattern encoding group is ≤ the offset threshold and the difference in feature density is ≥ the feature density threshold.
[0039] The core criterion for invalid association judgment is that the offset between the real-time encoding and the historical pattern encoding group is greater than or equal to the offset threshold, and the data is then removed.
[0040] As a further improvement to this method, the strong correlation is used to generate a trend basis vector through principal component conformal mapping. The final trend basis vector is:
[0041] T pred =T std +λ·U;
[0042] Among them, T pred T is the final trend basis vector. std The standard trend template is defined by λ, which is the real-time weight coefficient, and U is the initial trend basis vector.
[0043] The moderate correlation uses the feature density parameter to generate a probability perturbation term, and the final perturbation term generation formula is:
[0044]
[0045] Where Δx is the perturbation term vector, K is the total number of clusters, and w j Let c be the weight of the j-th cluster. j Let ε be the center vector of the j-th cluster.j For the j-th cluster, there is the Gaussian noise term;
[0046] The weak correlation is based on the energy conservation constraint to calculate the residual compensation, and the final optimal compensation amount is calculated using the following formula:
[0047]
[0048] Wherein, ΔP t For residual compensation at time t, R t Let ΔE be the residual at time t, ΔE be the total energy error, and T be the time window length.
[0049] As a further improvement to this method, the composite prediction model uses the trend basis vector T pred Disturbance term vector Δx, residual compensation ΔP t The component matrix M is constructed and orthogonalized into an orthogonal matrix Q and an upper triangular matrix R through QR decomposition. A weight vector W is constructed based on the weights of the association pattern classification. The contributions of each component are weighted and fused using the orthogonal projection formula to finally generate a composite prediction model, expressed as:
[0050] P F =QRW;
[0051] Among them, P F The final predicted power vector contains the corrected power values at all time points, and Q and R are the Q and R decomposition results from the component matrix M.
[0052] As a further improvement to this method, after the adaptive learning is triggered, a multi-level weight reallocation is performed, including:
[0053] The failure codes are sorted by sensitivity based on feature contribution, and the features with the highest contribution are split into sub-code groups;
[0054] Based on recent data, the wind speed range of the meteorological type index, the priority weight of the unit topology label, and the period division threshold of the time dynamic index are dynamically adjusted.
[0055] The beneficial effects of the technical solutions provided in this application include at least the following:
[0056] Beneficial effects:
[0057] This invention proposes a wind power prediction method based on data pattern prediction, achieving a technological breakthrough through the following innovative points:
[0058] Historical data is divided into core and peripheral groups, and a differentiated coding strategy is adopted. Combined with adaptive weight balancing of coding density, the frequency differences of multidimensional feature data are effectively handled.
[0059] A three-level index is constructed based on weather type, generator topology, and time dynamics to enable rapid multi-dimensional retrieval of historical patterns; differentiated prediction strategies are adopted based on different data characteristics through correlation pattern classification.
[0060] By combining principal component conformal mapping, probability distribution model and physical correction term, the system integrates statistical characteristics, probability fluctuations and physical constraints; and dynamically adjusts model parameters and index structure by triggering core group encoding decomposition through real-time error feedback.
[0061] Compared with the prior art, the present invention has significant advantages in the following aspects:
[0062] By using refined spatiotemporal coding and hierarchical weight fusion, the prediction error is reduced by 30%-50%, and it performs better in complex terrain and sudden weather scenarios.
[0063] Based on a three-level index-based fast retrieval and parallel computing architecture, prediction time is reduced to minutes, meeting the real-time requirements of power grid dispatch.
[0064] Adaptive learning mechanisms enable models to dynamically adapt to changes in data patterns and maintain stable predictive performance over the long term.
[0065] Through the visualization of intensive and envelope patterns, users can intuitively understand the basis for predictions, facilitating model debugging and optimization. Attached Figure Description
[0066] Figure 1 This is a schematic diagram of the overall process of a short-term wind power prediction method for wind farms based on data pattern prediction.
[0067] Figure 2 This is a flowchart of each sub-step of S100 in the present invention;
[0068] Figure 3 This is a flowchart of each sub-step of S200 in the present invention;
[0069] Figure 4 This is a flowchart of each sub-step of S300 in the present invention;
[0070] Figure 5 This is a flowchart of each sub-step of S400 in the present invention;
[0071] Figure 6 This is a flowchart of each sub-step of S500 in the present invention. Detailed Implementation
[0072] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0073] Example 1
[0074] In the field of wind power generation, accurate short-term wind power forecasting is crucial for grid dispatching, optimized turbine control, and improving wind energy utilization efficiency. Existing wind power forecasting methods are typically based on historical data statistics or physical models, but they have limitations in handling complex spatiotemporal data characteristics and dynamically adapting to weather changes and turbine operating conditions. Please refer to [link / reference]. Figure 1 This illustrates an embodiment of the present invention providing a method for short-term wind power prediction of wind farms based on data pattern prediction. The method includes:
[0075] S100: Collect historical multidimensional feature data and divide it into core group and edge group according to the frequency of occurrence of historical multidimensional feature data; the core group generates fine-grained coding by statistical extreme values through sliding window, and the edge group generates sparse coding based on event duration and intensity. Combine the adaptive weight algorithm to balance the density of the two types of coding and construct a frequency-balanced spatiotemporal coding matrix.
[0076] S200: Divide the spatiotemporal coding matrix into segments according to time windows, extract the condensed pattern and envelope pattern of each segment, and construct a three-level index based on the envelope pattern: meteorological type, unit topology and time dynamics. Store the two types of patterns in the historical pattern library, and establish a fast retrieval structure based on the three-level index.
[0077] S300: Extracts intensive and envelope patterns from real-time data, matches them against the historical pattern library using a three-level index, performs inclusive judgment, and classifies them into four types of association patterns: strong association, moderate association, weak association, and invalid association.
[0078] S400: Analyzes strong, moderate, and weak associations, and makes predictions based on hierarchical weighted fusion of association patterns.
[0079] S500: It makes predictions through a composite prediction model, compares the predictions with the actual results in real time, and if a certain type of pattern fails continuously, it triggers the core group coding decomposition, redistributes weights according to feature contribution and updates the three-level index to achieve adaptive learning.
[0080] This invention effectively overcomes the shortcomings of traditional physical and statistical models by using data pattern prediction and adaptive learning mechanisms, providing a high-precision and robust solution for short-term wind power prediction in wind farms.
[0081] The following are the specific implementation steps for each stage:
[0082] In the short-term wind power prediction method for wind farms based on data pattern prediction, S100 is the core step in constructing a frequency-balanced spatiotemporal coding matrix. Its core objective is to provide structured input for subsequent pattern matching by frequency partitioning, coding generation, and density balancing of historical multidimensional feature data. Please refer to [reference needed]. Figure 2 The diagram illustrates a flowchart of an exemplary medium-to-long-term energy determination method S100 based on state transition analysis, which includes:
[0083] S110: Collect historical multidimensional feature data and divide it into core group and edge group.
[0084] In short-term power forecasting of wind farms, the collection and preprocessing of historical multidimensional feature data is the cornerstone of the entire forecasting process.
[0085] In one possible implementation, historical multidimensional feature data includes wind farm operation data, meteorological data, unit topology information, and precise timestamp data.
[0086] For example, wind farm operation data includes rotor speed, generator power, pitch angle, gearbox temperature, etc.; meteorological data includes wind speed, wind direction, temperature, air pressure, etc.; and unit topology information includes wind turbine attributes, on-site layout data, etc.
[0087] The core group is defined as a high-frequency, dense region, and the edge group is defined as a sparse, low-frequency region. This division aims to distinguish between high-frequency, typical data patterns and low-frequency but high-impact special events. The division method is as follows:
[0088] Statistical analysis of the joint frequency of multidimensional feature combinations to identify high-frequency and low-frequency feature combinations.
[0089] For example, the frequency of the "wind speed-wind direction-temperature" combination in historical multidimensional feature data reflects the importance of the pattern. Frequency statistics can be used to identify which feature combinations are frequent and recurring, and which are infrequent and rare.
[0090] The feature combinations are grouped. In one possible implementation, a density-based clustering algorithm (DBSCAN) is used to group the feature combinations. DBSCAN can automatically identify dense and sparse regions without pre-specifying the number of clusters, making it suitable for the complex distribution characteristics of wind farm data.
[0091] S120: Generate fine-grained coding for the core group and sparse coding for the edge group.
[0092] To capture dynamic changes in the high-frequency characteristics of the core group, a variable-step-size sliding window technique is used. The window length is adaptively adjusted according to the wind speed fluctuation characteristics. Specific encoding methods include:
[0093] The core group data is segmented using a sliding window technique. For example, the window length is defined as 1 hour, the sliding step is 10 minutes, and each window covers 6 consecutive data points.
[0094] Within the window, key statistics are extracted, forming a fine-grained encoding of the window. In one possible implementation, key statistics include: maximum and minimum values to reflect extreme value fluctuations, corresponding timestamps marking the specific locations where extreme values occurred, and trend features characterizing the overall trend of changes in the data within the window.
[0095] Sparse coding for edge groups focuses on capturing low-frequency but high-impact events, such as extreme wind speed changes or sudden generator failures. Specific coding methods include:
[0096] Define the detection criteria for an "event". For example, a sudden and drastic fluctuation in wind speed within a short period of time, such as a change of more than 5 m / s within 10 minutes, or a sudden drop in power output exceeding a threshold, is considered an event.
[0097] A dual-threshold hysteresis comparator is used to detect event occurrence. For example, an event is determined to begin when the feature value exceeds the trigger threshold and lasts for at least 10 minutes, and to end when it falls back below the recovery threshold.
[0098] Extract event parameters. In one possible implementation, the extracted event parameters include an intensity index reflecting the degree of deviation, duration, and normalized occurrence time, which are then subjected to logarithmic transformation to compress the dynamic range and sine / cosine position encoding to simulate time periodicity.
[0099] Sparse coding employs a three-dimensional design. One-hot coding identifies event types, intensity-duration coding quantifies the impact of events, and time-location coding preserves the temporal information of event occurrence. When no event occurs, an all-zero code is generated.
[0100] S130: Design an adaptive weighting algorithm to construct a frequency-balanced spatiotemporal coding matrix.
[0101] The difference in coding density between the core group and the edge group may cause the model to favor high-frequency patterns and ignore important but low-frequency events. To address this, an adaptive weighting algorithm is designed to dynamically balance the contributions of both groups. First, the coding densities of the two groups are calculated. The core group density is the number of codes per unit time, reflecting the coverage of high-frequency patterns; the edge group density measures the frequency of low-frequency events. If the core group density is significantly higher than the edge group density, the model may over-rely on regular patterns and fail to respond to sudden events. Second, weights are dynamically allocated based on the coding density ratio. For example, if the edge group density is lower, it is given a higher weight to enhance the influence of low-frequency events; conversely, its weight is reduced to avoid noise interference.
[0102] The spatiotemporal coding matrix is constructed based on the weight allocation of the core group and the edge group, integrating the codes of the core group and the edge group to ensure frequency balance between the two. When constructing the spatiotemporal coding matrix, the time window is used as the row and the feature code is used as the column. The codes of the core group and the edge group are dynamically filled according to their weights, ultimately forming a frequency-balanced spatiotemporal coding matrix that can cover high-frequency patterns and capture key events.
[0103] In the short-term wind power prediction method for wind farms based on data pattern prediction, S200 transforms the original spatiotemporal encoding into physically meaningful structured knowledge through pattern extraction and three-level index construction. This provides efficient feature retrieval and state mapping capabilities for subsequent pattern matching and prediction models. Please refer to [reference needed]. Figure 3 The diagram illustrates a flowchart of an exemplary medium-to-long-term power determination method S200 based on state transition analysis, the contents of which include:
[0104] S210: Temporal window segmentation and fragment extraction of the spatiotemporal coding matrix.
[0105] The spatiotemporal coding matrix is the core data structure for short-term power prediction of wind farms, and it organizes feature encoding according to the time dimension. In order to extract effective pattern information, the matrix needs to be divided into several time segments to ensure that each segment can reflect local feature changes while maintaining temporal continuity.
[0106] In one possible implementation, a variable overlap sliding window technique is used to segment the spatiotemporal coding matrix. The window length is set to 1 hour, corresponding to 6 time intervals of 10 minutes, and the sliding step size is set to 10 minutes, so that each new window moves forward by 1 time interval relative to the previous window, forming a 50% overlap region.
[0107] For example, the first window covers 00:00-01:00, the second window covers 00:10-01:10, and so on. Each time point is contained in 6 consecutive windows to ensure that the data is fully utilized in the time dimension.
[0108] The spatiotemporal coding matrix is segmented into equal-length segments according to the window parameter, and each segment corresponds to a set of feature codes within a time window.
[0109] S220: Intensive pattern extraction—redundant feature removal and similar feature aggregation.
[0110] The core objective of the intensive model is to extract the most representative feature combinations from high-dimensional spatiotemporal encoding, reduce data complexity while retaining key information, and provide simplified input for subsequent pattern matching.
[0111] In one possible implementation, the intensive pattern extraction step includes:
[0112] Develop a redundancy removal strategy.
[0113] In one possible implementation, the redundancy removal strategy is selected based on variance threshold screening, that is, the variance of each feature column is calculated, columns with variance below the threshold are removed, and features sensitive to power fluctuations are retained.
[0114] In one possible implementation, the redundancy elimination strategy is selected based on correlation analysis, that is, by identifying highly linearly correlated feature groups through Pearson correlation coefficient, retaining only representative features, and reducing redundancy.
[0115] After removing redundant features, dynamic clustering is performed on the feature vectors within each time segment.
[0116] In one possible implementation, the K-means algorithm is used to cluster similar feature codes. For example, time windows with similar "wind speed-power" characteristics are clustered into one class, and the cluster centers are generated as representative vectors of the intensive pattern. The number and distribution density of codes within each cluster are counted to generate feature density parameters.
[0117] S230: Envelope Pattern Extraction—Extremum Boundary Construction and Encoding Optimization.
[0118] Envelope patterns are designed to capture extreme states and ranges of change within a time segment, providing boundary constraints for subsequent pattern matching and ensuring that there are clear physical threshold references for compatibility judgments between real-time data and historical patterns.
[0119] In one possible implementation, the envelope pattern extraction step includes:
[0120] For the fine-grained coding of the core group, the extreme values at the beginning and end of the time window are extracted, including the maximum value, minimum value and corresponding spatial distribution. For example, at the beginning of the window, the maximum wind speed measured by Unit 15 is 20 m / s, and at the end of the window, the wind speed of Unit 8 drops to 15 m / s. These two extreme points constitute the temporal evolution boundary of wind speed changes within this segment.
[0121] The initial and final extreme values are expanded to interval boundaries, and a four-dimensional boundary vector [min] is generated for each core feature dimension. start, max start ,min end, max end ] represents the minimum and maximum values of the features at the start and end times of the window, respectively.
[0122] As a further improvement to this method, for the event encoding of edge groups, a binary tag vector is used to record whether a preset type of event has occurred within the window, such as "turbulence event occurred once" or "device temperature exceeded limit 0 times".
[0123] S240: Construct a three-level index of weather type, unit topology, and time dynamics.
[0124] The three-level index is the core retrieval structure of the historical pattern database, accelerating the pattern matching process through hierarchical indexing.
[0125] Level 1: Weather Type Index
[0126] Meteorological types are classified based on the combination of wind speed, temperature, and air pressure. For example, "strong wind-low temperature-high pressure" corresponds to cold waves in winter, while "light breeze-high temperature-low pressure" corresponds to hot and humid weather in summer.
[0127] The index hierarchy is designed in three layers: the first layer is classified by wind speed range, such as 0-5m / s for low wind speed and 5-10m / s for medium wind speed; the second layer is subdivided by temperature range, such as -10℃ to 0℃ for low temperature and 20℃ to 30℃ for high temperature; and the third layer is sorted by air pressure gradient, such as high pressure zone and low pressure zone.
[0128] Level 2: Unit Topology Index:
[0129] The turbines are grouped according to their model, location, and grid connection. For example, the front-row turbines have significantly different power patterns than the rear-row turbines due to wake effects and need to be classified separately. Each turbine group is assigned a unique topology code, such as "T01" for front-row doubly-fed induction generator (DFIG) turbines and "T02" for rear-row direct-drive turbines. The topology code is embedded in the envelope pattern for easy retrieval based on turbine characteristics.
[0130] Level 3: Time-based dynamic index.
[0131] The time dimension is divided into layers based on season, day / night cycle, and weather cycle. For example, wind speed patterns during the day in summer differ significantly from those at night in winter. Dynamic timestamp association adds time labels to each envelope pattern, supporting quick filtering by time range. For instance, to retrieve all matching patterns during the "2023 typhoon season," the system only needs to iterate through the data corresponding to the time label, greatly improving efficiency.
[0132] S250: Stores historical pattern libraries of intensive and envelope patterns, and establishes fast retrieval based on a three-level index.
[0133] The feature density parameters of the intensive pattern and the boundary codes of the envelope pattern are stored separately. A fast retrieval structure is constructed based on a three-level index of weather type, unit topology, and time dynamics.
[0134] In one possible implementation, the weather type index achieves exact matching through a hash table, the unit topology index embeds a unique code to support group filtering, and the time dynamic index utilizes a B-tree to efficiently handle time interval queries.
[0135] In the short-term wind power prediction method for wind farms based on data pattern prediction, the S300 focuses on the efficient matching and classification of real-time data and historical patterns, solving the efficiency problem of high-dimensional data matching. Furthermore, it ensures the reliability of the matching results through physical constraints and dynamic weights, laying a solid foundation for subsequent hierarchical fusion prediction. Please refer to [reference needed]. Figure 4 The diagram illustrates a flowchart of an exemplary medium-to-long-term energy determination method S300 based on state transition analysis, the contents of which include:
[0136] S310: Extract intensive / envelope mode for real-time data.
[0137] Real-time data is segmented using a sliding window technique, with window parameters consistent with historical data. Within each window, core feature statistics are extracted, including extreme values, trend slope, and feature density. Redundant features are removed, and similar features are aggregated using a clustering algorithm to generate real-time intensive patterns. For example, if the "wind speed-power" feature distribution of the real-time window is highly similar to a cluster center in the historical pattern library, a corresponding intensive code is generated.
[0138] This involves capturing the extreme boundaries and potential events of the current time window from real-time data. First, the core feature extreme values at the beginning and end of the window are extracted to form [min...]. start, max start ,min end, max end A four-dimensional boundary vector. If an edge group event is detected, the start and end times of the event are determined using a dual-threshold hysteresis comparator, and an event code is generated. When no event occurs, the envelope pattern only contains the core extremum boundaries.
[0139] S320: Three-level index for fast matching of candidate historical patterns.
[0140] The three-level index of weather type, unit topology, and time dynamics is the core retrieval framework connecting real-time mode and historical data. It filters step by step according to the hierarchy of "time dynamics → weather type → unit topology" to ensure retrieval efficiency and accuracy.
[0141] Weather type matching: Weather tags are generated based on real-time weather data and matched with weather type indexes in the historical model database. For example, if the real-time wind speed is 8 m / s, the temperature is 25℃, and the air pressure is 1010 hPa, then the "medium wind speed - high temperature - normal pressure" type is matched in the historical database.
[0142] Unit topology filtering: Based on real-time unit topology information, it matches the topology codes in the historical pattern library. For example, if the real-time data comes from a front-row doubly-fed induction generator (DFIG) wind turbine with the code T01, then only the unit patterns marked as T01 in the historical database are retrieved.
[0143] Dynamic Time Filtering: The system filters the historical pattern library's dynamic time indexes based on the time labels of real-time data. The B-tree structure supports efficient time interval queries; for example, it retrieves all relevant patterns for "summer daytime over the past three years." If the real-time data includes specific weather cycles, the time range is further narrowed to improve matching accuracy.
[0144] Multithreading is employed to perform three-level index matching in parallel. In one possible implementation, one thread handles weather type hash table queries, another thread filters unit topology codes, and a third thread traverses a time-dynamic B-tree. Finally, the candidate pattern sets from the three threads are merged to form a preliminary matching result.
[0145] S330: Inclusive Judgment and Association Pattern Classification.
[0146] Based on the feature differences between real-time patterns and historical fragments, a four-layer inclusive judgment logic is designed to strictly distinguish between strong correlation, moderate correlation, weak correlation and invalid correlation.
[0147] The core criterion for strong correlation judgment is that the real-time encoding completely falls within the envelope boundary of the historical pattern library for a certain time segment and the feature density difference is ≤ the feature density threshold.
[0148] The core criteria for judging moderate correlation are that the offset between the real-time encoding and the historical pattern encoding group is ≤ the offset threshold and the difference in feature density is ≤ the feature density threshold.
[0149] The core criteria for weak association judgment are that the offset between the real-time encoding and the historical pattern encoding group is ≤ the offset threshold and the difference in feature density is ≥ the feature density threshold.
[0150] The core criterion for invalid association judgment is that the offset between the real-time encoding and the historical pattern encoding group is greater than or equal to the offset threshold, and the data is then removed. Such data is usually caused by sensor failure or extreme abnormal events and is directly removed to avoid interfering with the prediction model.
[0151] In the short-term wind power prediction method for wind farms based on data pattern prediction, the S400 uses a hierarchical weighted fusion prediction based on correlation patterns. Strong correlation generates trend basis vectors through principal component conformal mapping; moderate correlation generates probability perturbation terms using feature density parameters; and weak correlation calculates residual compensation based on energy conservation constraints. Finally, these components are fused through orthogonal projection to form a composite prediction model. Please refer to [reference needed]. Figure 5 The diagram illustrates a flowchart of an exemplary medium-to-long-term energy determination method S400 based on state transition analysis, the contents of which include:
[0152] S410: Deterministic trend basis vector generation for strongly correlated patterns.
[0153] Strong correlation patterns represent a high degree of consistency between real-time data and historical patterns, and their core value lies in providing a reliable deterministic prediction benchmark. Trend features are extracted and trend basis vectors are generated using principal component conformal mapping technology.
[0154] Construct a covariance matrix Σ∈R for the intensive pattern feature density parameters of strongly correlated historical fragments. D×D D is the feature dimension. The first k principal components are extracted through eigenvalue decomposition to form an orthogonal basis vector set {u1, u2, ..., u...}. k The feature vector X of the real-time strong correlation pattern is projected onto the principal component space to obtain the projection coefficients c. i =X k u i Among them, u i Given the orthogonal basis vectors of the i-th principal component, generate the initial trend basis vector U = {c1, c2, ..., c...} k Furthermore, by employing conformal mapping techniques, the principal component trajectories of historically strongly correlated segments are time-aligned to obtain a standardized trend template T. std The final trend basis vector is:
[0155] T pred =T std +λ·U;
[0156] Among them, T pred T is the final trend basis vector. std For the standardized trend template, λ is the real-time weight coefficient, and U is the initial trend basis vector.
[0157] S420: Probability perturbation term for moderately correlated patterns.
[0158] The moderately correlated patterns show some deviation from historical patterns. Uncertainty in prediction is simulated using a probabilistic perturbation term. The feature density parameter reflects the concentration of the pattern distribution and is used to construct the amplitude and direction of the perturbation term. For each cluster in the historical moderately correlated pattern library, its density parameter is calculated, specifically the proportion of patterns within the cluster to the total number of patterns. The distance between the real-time moderately correlated patterns and the centers of each cluster is quantized using Euclidean distance. Based on the density parameter and distance, a weight distribution for the perturbation term is constructed, using a Gaussian kernel function for weighting. A Gaussian noise term is introduced when generating the perturbation vector to simulate intra-cluster fluctuations. The final perturbation term generation formula is:
[0159]
[0160] Where Δx is the perturbation term vector, K is the total number of clusters, i.e., the number of clusters divided in the historical moderate association pattern library, and w j Let c be the weight of the j-th cluster. j Let ε be the center vector of the j-th cluster. j Let be the Gaussian noise term of the j-th cluster.
[0161] S430: Energy conservation residual compensation for weakly correlated modes.
[0162] The weakly correlated model differs significantly from historical models, and prediction bias is corrected through physical constraints. Energy conservation constraints ensure consistency between predicted and measured power in terms of global energy. First, the residual sequences of predicted and measured power are calculated, defining the global energy error as the sum of squared residuals, and then a constraint is introduced requiring the corrected predicted power to satisfy energy balance. The compensation amount is optimized by constructing a Lagrangian function; the final optimal compensation amount formula is:
[0163]
[0164] Wherein, ΔP t For residual compensation at time t, R t ΔE is the residual at time t, which is the difference between the measured power and the preliminary predicted power. ΔE is the total energy error, which is the sum of the residuals at all time points. T is the length of the time window.
[0165] S440: Three-level orthogonal projections fuse deterministic components, probabilistic perturbation terms, and physical correction terms to form a composite prediction model.
[0166] Orthogonal projection is used to compensate for the trend basis vector of the strongly correlated mode, the probability perturbation term of the moderately correlated mode, and the energy conservation residual of the weakly correlated mode, forming a composite prediction model.
[0167] The three components, trend basis vector T pred Disturbance term vector Δx, residual compensation ΔP tThe component matrix M is constructed and orthogonalized into an orthogonal matrix Q and an upper triangular matrix R through QR decomposition, eliminating the correlation between components. Subsequently, a weight vector W is constructed based on the weights of the association pattern classification. The contributions of each component are weighted and fused using the orthogonal projection formula, ultimately generating a composite prediction model, expressed as:
[0168] P F =QRW;
[0169] Among them, P F The final predicted power vector contains the corrected power values at all time points. Q and R are the Q and R decomposition results from the component matrix M.
[0170] In the short-term wind power prediction method for wind farms based on data pattern prediction, the S500 triggers failure mode analysis by comparing real-time predictions with actual measurements, decomposes the core group coding, and reallocates weights according to feature contribution. It also dynamically updates the three-level indexes of meteorology, turbine, and time, achieving adaptive learning and continuous optimization of the model. Please refer to [reference needed]. Figure 6 It shows a flowchart of an exemplary medium-to-long-term power determination method S500 based on state transition analysis, the contents of which include:
[0171] S510: Real-time error monitoring and adaptive triggering.
[0172] The system continuously collects actual power output data from wind farms and compares it with the predicted values of the composite prediction model at each time point to calculate error indices. When the error of a certain associated mode continuously exceeds a preset threshold, that mode is deemed to have failed. The triggering mechanism requires dynamic adjustment of the threshold: during periods of drastic wind speed fluctuations or extreme weather, the threshold is relaxed to avoid false positives; during steady-state operation, the threshold is tightened to improve sensitivity. Simultaneously, a detailed error report is generated, recording the time period of the failure mode, the error type, and potential causes, providing data support for subsequent analysis.
[0173] S520: Once a failure is detected, the system traces the root cause from three levels: three-level index, coding features, and physical constraints, and locates the specific failure link.
[0174] Check real-time data for noise or missing data. Extract the feature distribution of the core group code within the failure period and compare it with historical similar patterns. If the feature density parameter deviates significantly, it indicates that the environment or equipment status has changed.
[0175] The failure codes are split into sub-code groups based on the contribution of features. For example, if the "wind speed-power" feature has the highest error contribution, it is removed from the core group and re-clustered to generate fine-grained sub-codes such as "high wind speed-low power" and "medium wind speed-high power".
[0176] S530: Based on the failure diagnosis results, the core coding, three-level index and pattern library are optimized in a hierarchical manner to ensure that the model can quickly adapt to changes and avoid prediction oscillations.
[0177] The updating and validation of the three-level index ensures that the model continuously adapts to changes in the environment:
[0178] Weather type adjustment: Wind speed and temperature ranges are reclassified based on recent data. For example, if the regional average wind speed increases, the "medium wind speed" will be adjusted from 5-10 m / s to 6-12 m / s.
[0179] Unit topology reconfiguration: For equipment updates or layout changes, add new topology tags and eliminate outdated codes.
[0180] Time tag expansion: Add special weather cycles and delete historical tags that have not been matched for a long time.
[0181] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.
[0182] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0183] It should also be noted that in the apparatus, equipment, and methods of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.
[0184] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features of the invention herein.
[0185] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications or equivalent substitutions made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for short-term wind power prediction in wind farms based on data pattern prediction, characterized in that, include: Collect historical multidimensional feature data and divide it into core group and edge group according to the frequency of occurrence of historical multidimensional feature data; The core group generates fine-grained codes by statistically analyzing extreme values through a sliding window, while the edge group generates sparse codes based on event duration and intensity. An adaptive weighting algorithm is then used to balance the density of the two types of codes, thus constructing a frequency-balanced spatiotemporal coding matrix. The spatiotemporal coding matrix is divided into segments according to time windows. The intensive pattern and envelope pattern of each segment are extracted. Based on the envelope pattern, a three-level index of weather type, unit topology and time dynamics is constructed. The two types of patterns are stored in the historical pattern library. A fast retrieval structure is established based on the three-level index. The system extracts intensive and envelope patterns from real-time data, matches them against a historical pattern library using a three-level index, performs inclusive judgment, and classifies them into four types of association patterns: strong association, moderate association, weak association, and invalid association. Analyze strong, moderate, and weak associations, and make predictions based on hierarchical weighted fusion of association patterns; Prediction is made through a composite prediction model, and the prediction results are compared with the actual results in real time. If a certain type of pattern fails continuously, the core group encoding decomposition is triggered, the weights are redistributed according to the feature contribution and the three-level index is updated to achieve adaptive learning.
2. The method for short-term wind power prediction of wind farms based on data pattern prediction according to claim 1, characterized in that, When dividing the core group and the edge group, a dynamic joint frequency statistics and density clustering fusion strategy is adopted: Statistically analyze the joint frequency of multidimensional feature combinations and identify high-frequency feature combinations as core group candidates; Candidate feature combinations are grouped a second time using a density-based clustering algorithm to remove noise points and generate the final core group and edge group division results.
3. The method for short-term wind power prediction of wind farms based on data pattern prediction according to claim 2, characterized in that, The core group generates fine-grained codes using a variable step-size sliding window technique. The window length is adaptively adjusted according to the wind speed fluctuation characteristics, with a sliding step of 10 minutes and a window length of 1 hour. Within each window, the maximum value, minimum value, extreme value timestamp, and trend slope are extracted as encoding features. The edge group generates sparse coding using event intensity-time coding technology. Events are detected by a dual-threshold hysteresis comparator. For the detected events, event parameters are extracted and combined into a three-dimensional sparse vector using a three-dimensional design.
4. The method for short-term wind power prediction of wind farms based on data pattern prediction according to claim 3, characterized in that, The adaptive weighting algorithm's strategy for balancing the two types of coding densities includes: Calculate the density of fine-grained coding and sparse coding; Among them, the density of fine-grained coding is the number of codes per unit time, and the density of sparse coding is the frequency of occurrence of low-frequency events. Weights are dynamically allocated based on the density ratio of the encoding. The dynamic weight allocation includes reducing the weight of the core group and increasing the weight of the edge group proportionally if the density of the core group is higher than the frequency of the edge group.
5. The method for short-term wind power prediction of wind farms based on data pattern prediction according to claim 1, characterized in that, The intensive pattern is to aggregate similar features after removing redundancy. Generate feature density parameters; the envelope pattern is formed by extracting the first and last extreme values of the time window to create a minimum-maximum encoding boundary; The intensive pattern extraction step includes: The first layer filters out high-volatility features by using a variance threshold, and removes feature columns with variances below a preset threshold. The second layer uses Pearson correlation coefficient analysis to merge highly linearly correlated feature groups while retaining representative features. K-means dynamic clustering is performed on the simplified feature vectors to generate cluster centers as the intensive pattern, and the density parameters within the clusters are statistically analyzed. The envelope pattern extraction step includes: Extract the core feature extrema at the beginning and end of the time window to generate a four-dimensional boundary vector [min start, max start ,min end, max end ]; For edge group events, add a binary marker vector to record the event type and the number of occurrences, and concatenate it with the core boundary vector to form a hybrid envelope code.
6. The method for short-term wind power prediction of wind farms based on data pattern prediction according to claim 1, characterized in that, When constructing the three-level index of meteorological type, unit topology, and time dynamics: The meteorological type index is divided into layers based on wind speed range, temperature range, and pressure gradient. The unit topology index generates a unique code based on the wind turbine model, location layout, and grid connection relationship; The time-based dynamic index associates dynamic time tags with seasons, day / night cycles, and weather cycles.
7. The method for short-term wind power prediction of wind farms based on data pattern prediction according to claim 1, characterized in that, The criteria for classifying the inclusion judgment into strong association, moderate association, weak association, and invalid association are as follows: The core criterion for judgment is that the real-time encoding completely falls within the envelope boundary of the historical pattern library for a certain time segment and the feature density difference is ≤ the feature density threshold. The core criteria for judging moderate correlation are that the offset between the real-time encoding and the historical pattern encoding group is ≤ the offset threshold and the difference in feature density is ≤ the feature density threshold. The core criteria for weak association judgment are that the offset between the real-time encoding and the historical pattern encoding group is ≤ the offset threshold and the difference in feature density is ≥ the feature density threshold. The core criterion for invalid association judgment is that the offset between the real-time encoding and the historical pattern encoding group is greater than or equal to the offset threshold, and the data is then removed.
8. The method for short-term wind power prediction of wind farms based on data pattern prediction according to claim 1, characterized in that, The strong correlation is used to generate trend basis vectors through principal component conformal mapping. The final trend basis vectors are: T pred =T std +λ·U; Among them, T pred T is the final trend basis vector. std The standard trend template is defined by λ, which is the real-time weight coefficient, and U is the initial trend basis vector. The moderate correlation uses the feature density parameter to generate a probability perturbation term, and the final perturbation term generation formula is: Where Δx is the perturbation term vector, K is the total number of clusters, and w j Let c be the weight of the j-th cluster. j Let ε be the center vector of the j-th cluster. j For the j-th cluster, there is the Gaussian noise term; The weak correlation is based on the energy conservation constraint to calculate the residual compensation, and the final optimal compensation amount is calculated using the following formula: Where, ΔP t For residual compensation at time t, R t Let ΔE be the residual at time t, ΔE be the total energy error, and T be the time window length.
9. A method for short-term wind power prediction of wind farms based on data pattern prediction according to claim 1 or 8, characterized in that, The composite prediction model uses the trend basis vector T pred Disturbance term vector Δx, residual compensation ΔP t The component matrix M is constructed and orthogonalized into an orthogonal matrix Q and an upper triangular matrix R through QR decomposition. A weight vector W is constructed based on the weights of the association pattern classification. The contributions of each component are weighted and fused using the orthogonal projection formula to finally generate a composite prediction model, expressed as: P F =QRW; Among them, P F The final predicted power vector contains the corrected power values at all time points, and Q and R are the Q and R decomposition results from the component matrix M.
10. A method for short-term wind power prediction of wind farms based on data pattern prediction according to claim 1, characterized in that, After the adaptive learning is triggered, a multi-level weight reallocation is performed, including: The failure codes are sorted by sensitivity based on feature contribution, and the features with the highest contribution are split into sub-code groups; Based on recent data, the wind speed range of the meteorological type index, the priority weight of the unit topology label, and the period division threshold of the time dynamic index are dynamically adjusted.
Citation Information
Cited By
Transformer operation state real-time analysis method and system facing edge calculation
CN122332832A