Medium and short term wind speed correction method and model based on regional weather forecast mode
Through the tower-type local hierarchical capturing attention mechanism and activity difference term loss function based on regional weather forecasting mode, the problem of physical characteristics neglected and smooth forecasting in medium and short-term wind speed forecasting is solved, and higher accuracy and reliable wind speed forecasting is achieved.
Patent Information
- Application Number
- CN202510410800.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-02
AI Technical Summary
The existing medium- and short-term wind speed forecasting methods ignore the physical characteristics of wind speed, resulting in insufficient physical consistency and reliability of the forecast results, weak model robustness, and traditional loss functions cause the forecast results to be too smooth and cannot effectively capture wind speed fluctuations.
The medium- and short-term wind speed correction method based on regional weather forecast mode is adopted, and the meteorological variable data is encoded through the tower-type local hierarchical capture attention mechanism, combining multi-dimensional information embedding and activity difference term loss function to improve the physical consistency and accuracy of wind speed forecasting.
It significantly improves the accuracy and reliability of medium- and short-term wind speed forecasting, can express the physical laws of wind speed more accurately, reduce forecast errors, and enhance the stability of wind power power prediction.
Smart Images

Figure CN120278026A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of wind power generation, and in particular, to a medium - and short - term wind speed correction method and model based on a regional weather forecasting mode. Background Art
[0002] The acquisition of wind energy depends on the wind speed within the atmospheric boundary layer. Its volatility and intermittency lead to unstable output of wind power, posing an impact on the operation of the power grid. Therefore, the development of high - precision medium - and short - term wind speed forecasting technologies (1 - 7 days) is crucial for power grid dispatching and power system optimization.
[0003] The current medium - and short - term wind speed correction methods have evolved from traditional statistical methods to machine learning and deep learning. Although early linear regression and autoregressive integrated moving average (ARIMA) models had certain effects, they were difficult to meet the requirements of long - term forecasting. Subsequently, machine learning methods such as gradient boosting decision tree (GBDT) and neural network models (such as recurrent neural network (RNN), long short - term memory network (LSTM)) were introduced, significantly improving the forecasting performance. In recent years, attention mechanisms and Transformer models have also been applied to wind speed forecasting, further promoting the development of the technology. In addition, post - processing correction methods that combine signal processing methods (such as empirical mode decomposition (EMD), ensemble empirical mode decomposition (EEMD), variational mode decomposition (VMD)) with artificial intelligence have also achieved good results.
[0004] However, in the related technologies, directly using artificial intelligence models or artificial intelligence models after signal processing for correction ignores the physical characteristics of wind speed, resulting in insufficient physical consistency and reliability of the forecasting results, weak model robustness, and thus affecting its reliability and credibility. Summary of the Invention
[0005] In view of this, the present disclosure proposes a medium - and short - term wind speed correction method and model based on a regional weather forecasting mode.
[0006] According to one aspect of the present disclosure, there is provided a medium - and short - term wind speed correction method based on a regional weather forecasting mode, the method comprising:
[0007] Obtain meteorological variable data, which is generated by a preset regional weather forecasting model and is used to indicate the temporal variation of multiple meteorological variables related to wind speed in a target area over a period of time in the future;
[0008] Perform information embedding on the meteorological variable data to generate a target embedding tensor;
[0009] Adopt a tower-shaped local hierarchical capture attention mechanism to encode the target embedding tensor to obtain a prediction feature tensor, where the tower-shaped local hierarchical capture attention mechanism is used to indicate that the attention calculation of each time node in the target embedding tensor is constrained to a finite number of time nodes associated with the time node;
[0010] Perform feature transformation on the prediction feature tensor to obtain corrected wind speed forecast data.
[0011] In a possible implementation manner, the performing information embedding on the meteorological variable data to generate a target embedding tensor includes:
[0012] Convert the position index of each time node in the meteorological variable data into a position encoding vector;
[0013] Convert the time stamp of each time node in the meteorological variable data into a time feature vector;
[0014] Extract meteorological variable data with different time resolutions from the meteorological variable data at the original time resolution, and splice the meteorological variable data with different time resolutions to obtain a meteorological variable tensor;
[0015] Splice the position encoding vector, the time feature vector, and the meteorological variable tensor to obtain the target embedding tensor.
[0016] In another possible implementation manner, the extracting meteorological variable data with different time resolutions from the meteorological variable data at the original time resolution includes:
[0017] According to a preset time stamp, extract meteorological variable data with different time resolutions from the meteorological variable data at the original time resolution; or,
[0018] Perform a convolution operation on the meteorological variable data at the original time resolution by using a Convolutional Neural Network (CNN) to extract meteorological variable data with different time resolutions.
[0019] In another possible implementation, encoding the target embedding tensor by using the tower-type local hierarchical attention mechanism to obtain a predicted feature tensor includes:
[0020] Determine a set of hierarchical relationships for each time node of the target embedding tensor, where the relationship set includes a set of a finite number of time nodes associated with the time node at multiple levels, and the levels are divided according to different time scales;
[0021] Perform multi-head attention calculation on the set of hierarchical relationships of each time node to obtain an attention result;
[0022] Perform a preset process on the attention result to obtain the predicted feature tensor.
[0023] In another possible implementation, the relationship set is the union of a child-level relationship set, a sibling-level relationship set, and a parent-level relationship set. The child-level relationship set is used to indicate one or more time nodes closest to the time node at the first level, the sibling-level relationship set is used to indicate one or more time nodes closest to the time node at the second level, and the parent-level relationship set is used to indicate one time node closest to the time node at the third level. The time scales corresponding to the first level, the second level, and the third level are increasing.
[0024] In another possible implementation, converting the predicted feature tensor to obtain corrected wind speed forecast data includes:
[0025] Extract target features from the predicted feature tensor through a gather layer;
[0026] Perform a linear transformation on the extracted target features through a linear layer to obtain the corrected wind speed forecast data.
[0027] In another possible implementation, the method is executed based on a pre-trained wind speed correction model, and the method further includes:
[0028] Obtain training data, where the training data includes a plurality of wind speed forecast values and corresponding wind speed observation values;
[0029] Calculate a loss function according to the plurality of wind speed forecast values and the corresponding plurality of wind speed observation values. The loss function includes an activity difference term, and the activity difference term is used to indicate the difference in the fluctuation degree between the plurality of wind speed forecast values and the corresponding plurality of wind speed observation values;
[0030] According to the loss function, the model parameters are iteratively updated through a preset algorithm to train the wind speed correction model.
[0031] According to another aspect of the present disclosure, a wind speed correction model is provided, and the model includes:
[0032] An input module, configured to obtain meteorological variable data, where the meteorological variable data is generated by a preset regional weather forecasting model, and the meteorological variable data is used to indicate the temporal variation of multiple meteorological variables related to wind speed in a target area within a future period of time;
[0033] An embedding module, configured to perform information embedding on the meteorological variable data to generate a target embedding tensor;
[0034] An encoder module, configured to encode the target embedding tensor by using a tower-shaped local hierarchical attention mechanism, where the tower-shaped local hierarchical attention mechanism is used to indicate that the attention calculation of each time node in the target embedding tensor is constrained to a finite number of time nodes associated with the time node;
[0035] A post-processing module, configured to perform feature transformation on the predicted feature tensor to obtain corrected wind speed forecast data.
[0036] In a possible implementation manner, the embedding module is further configured to:
[0037] Convert the position index of each time node in the meteorological variable data into a position encoding vector;
[0038] Convert the time stamp of each time node in the meteorological variable data into a time feature vector;
[0039] Extract meteorological variable data with different time resolutions from the meteorological variable data with the original time resolution, and splice the meteorological variable data with different time resolutions to obtain a meteorological variable tensor;
[0040] Splice the position encoding vector, the time feature vector, and the meteorological variable tensor to obtain the target embedding tensor.
[0041] In another possible implementation manner, the embedding module is further configured to:
[0042] Extract meteorological variable data with different time resolutions from the meteorological variable data with the original time resolution according to a preset time stamp; or,
[0043] Perform a convolution operation on the meteorological variable data with the original time resolution by using a convolutional neural network to extract meteorological variable data with different time resolutions.
[0044] In another possible implementation, the encoder module is further configured to:
[0045] Determine a hierarchical relationship set for each time node of the target embedding tensor, where the relationship set includes a set of a finite number of time nodes associated with the time node at multiple levels, and the levels are divided according to different time scales;
[0046] Perform multi-head attention calculation on the hierarchical relationship sets of each time node to obtain an attention result;
[0047] Perform a preset process on the attention result to obtain the predicted feature tensor.
[0048] In another possible implementation, the hierarchical relationship set is the union of a child-level relationship set, a sibling-level relationship set, and a parent-level relationship set. The child-level relationship set is used to indicate one or more time nodes closest to the time node at the first level, the sibling-level relationship set is used to indicate one or more time nodes closest to the time node at the second level, and the parent-level relationship set is used to indicate one time node closest to the time node at the third level. The time scales corresponding to the first level, the second level, and the third level are increasing.
[0049] In another possible implementation, the post-processing module is further configured to:
[0050] Extract target features from the predicted feature tensor through an aggregation layer;
[0051] Perform a linear transformation on the extracted target features through a linear layer to obtain the corrected wind speed forecast data.
[0052] In another possible implementation, the loss function of the model includes an activity difference term, and the activity difference term is used to indicate the difference in the degree of fluctuation between the multiple wind speed forecast values and the corresponding multiple wind speed observation values.
[0053] According to another aspect of the present disclosure, there is provided a computing device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the above method.
[0054] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0055] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0056] The embodiments of the present disclosure provide a medium- and short-term wind speed correction method based on a regional weather forecasting model. By obtaining meteorological variable data generated by a preset regional weather forecasting model, these data can reflect the temporal variation of multiple meteorological variables related to wind speed in a target area over a period of time in the future. Subsequently, information embedding is performed on the meteorological variable data to generate a target embedding tensor. The core lies in encoding the target embedding tensor using a tower-shaped local hierarchical capture attention mechanism to obtain a prediction feature tensor, and then performing feature transformation on the prediction feature tensor to obtain corrected wind speed forecast data. This tower-shaped local hierarchical capture attention mechanism constrains the attention calculation to a finite number of time nodes associated with time nodes, which means that the model can more accurately focus on local time series information that has a greater impact on wind speed changes, avoiding the key information being submerged due to treating all time points equally in traditional methods. For example, during a period when the wind speed changes violently, the model can focus on the changes in meteorological variables associated with this period and its vicinity, thereby more accurately capturing the change trend of the wind speed. This design not only makes the model output more physically consistent but also significantly improves the reliability and robustness of the forecast results. Compared with traditional methods, the method provided by the embodiments of the present disclosure avoids the limitations of directly using artificial intelligence models or relying only on signal processing, can more accurately express the physical laws of wind speed, and provides a more efficient and reliable solution for medium- and short-term wind speed forecasting.
[0057] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The drawings included in and constituting a part of the specification illustrate exemplary embodiments, features, and aspects of the present disclosure together with the specification and are used to explain the principles of the present disclosure.
[0059] Figure 1 Shows a schematic structural diagram of a wind speed correction model provided by an exemplary embodiment of the present disclosure.
[0060] Figure 2 Shows a flowchart of a medium- and short-term wind speed correction method based on a regional weather forecasting model provided by an exemplary embodiment of the present disclosure.
[0061] Figure 3 Shows a flowchart of an information embedding method provided by an exemplary embodiment of the present disclosure.
[0062] Figure 4 The flowchart of an encoding method provided by an exemplary embodiment of the present disclosure is shown.
[0063] Figure 5 The flowchart of a feature conversion method provided by an exemplary embodiment of the present disclosure is shown.
[0064] Figure 6 The schematic diagram of the principle of a medium- and short-term wind speed correction method based on a regional weather forecasting model provided by an exemplary embodiment of the present disclosure is shown.
[0065] Figure 7 It is a block diagram of a device shown according to an exemplary embodiment. Detailed implementation manners
[0066] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.
[0067] As used herein, the terms "comprising", "including", "having", or variations thereof are open-ended and include one or more stated features, wholes, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, wholes, elements, steps, components, functions, or groups thereof.
[0068] When an element is referred to as being "connected", "coupled", "responsive" or variations thereof to another element, it can be directly connected, coupled, or responsive to the other element, or there can be intervening elements.
[0069] Although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Thus, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.
[0070] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein need not be construed as being superior to or better than other embodiments.
[0071] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0072] Short-term and medium-term wind speed forecasting usually involves predicting the wind speed for the next 1 to 7 days (or even 1 to 10 days), and the results can be used to calculate wind power forecasts. On this time scale, it is difficult to achieve effective wind speed forecasting solely relying on data-driven artificial intelligence (AI) models because the volatility and intermittency of wind speed are significantly affected by meteorological conditions. Therefore, based on the original forecast data of numerical weather prediction models and combined with statistical methods or artificial intelligence technologies for error correction is the key means to improve the accuracy of wind speed forecasting at present.
[0073] However, there are obvious deficiencies in the current wind speed correction methods. On the one hand, many methods tend to directly use artificial intelligence models or only rely on artificial intelligence models for post-processing correction after signal processing. This approach ignores the physical characteristics of wind speed and lacks in-depth understanding of the physical characteristics of wind speed, resulting in insufficient reliability and physical consistency of the forecast results, and relatively weak robustness of the model.
[0074] On the other hand, the applicability of signal processing methods (such as empirical mode decomposition, ensemble empirical mode decomposition, and variational mode decomposition) is not yet clear because their frequency decomposition highly depends on the sequence length. Therefore, a sequence decomposition model with physical significance is needed to better capture and express the physical laws of wind speed, thereby enhancing the constraints of the model on physical laws and improving the robustness and credibility of the forecast.
[0075] On the other hand, in the training of artificial intelligence models, general evaluation indicators such as mean squared error (MSE) and mean absolute error (MAE) are usually directly used. Although these indicators are widely used in model training, in specific applications such as wind power prediction, they may cause the model output results to be too smooth to effectively capture the volatility of wind speed. At present, no research or industry application has proposed effective improvement or solutions to this problem.
[0076] To solve the above problems, the embodiments of the present disclosure provide an innovative medium - and short - term wind speed correction method. This method obtains various meteorological variable data related to wind speed generated by a preset regional weather forecasting model to comprehensively understand the changes in meteorological elements in the target area over a period of time in the future. These data cover multi - dimensional information such as wind speed, wind direction, temperature, humidity, and air pressure, providing rich basic data for wind speed correction. The core innovation lies in using a tower - type local hierarchical attention mechanism to encode the target embedding tensor. This mechanism constrains the attention calculation to a limited number of time nodes associated with each time node by determining the set of hierarchical relationships at each time node. This design can better capture the physical characteristics of wind speed at different time scales, making the model output more physically consistent. Compared with traditional methods, this method can more accurately express the physical laws of wind speed, significantly improving the reliability and robustness of the prediction results.
[0077] In the data pre - processing stage, the embodiments of the present disclosure perform convolution operations on the meteorological variable data with the original time resolution through a convolutional neural network to extract meteorological variable data with different time resolutions. In addition, positional encoding, time features, and meteorological variable tensors are concatenated to generate the target embedding tensor. The convolutional neural network can automatically extract the spatial features of the data, avoiding the signal processing problems caused by the sequence length in traditional methods. At the same time, the multi - dimensional information embedding enhances the model's ability to understand the input data, further improving the accuracy of wind speed prediction.
[0078] In the model training stage, the embodiments of the present disclosure design a loss function that includes an activity difference term, which is used to measure the difference in the degree of fluctuation between the predicted wind speed value and the true value. This loss function can effectively avoid the problem that traditional loss functions (such as MSE, MAE) may cause the output results to be too smooth. This loss function with an activity difference term can effectively avoid this problem, enabling the model to better capture the volatility of wind speed, thereby improving the accuracy of wind power prediction.
[0079] In summary, through the innovative attention mechanism, multi - dimensional information embedding, and targeted loss function design, the embodiments of the present disclosure provide an efficient and reliable medium - and short - term wind speed correction method, significantly improving the accuracy and reliability of wind speed prediction, and providing strong technical support for wind power prediction and the stable operation of new power systems.
[0080] Embodiments of the present disclosure provide a wind speed correction model, which can be used to execute the short- and medium-term wind speed correction method based on the regional weather forecasting model provided by the embodiments of the present disclosure. This model adopts an encoder-decoder architecture, with the focus on the construction of the encoder. The reason for determining the main architecture of the model as the encoder part and discarding the decoder part is that the wind speed correction task is essentially a typical regression mapping problem rather than a generative prediction problem, so there is no need to introduce the structural design of the decoder.
[0081] Please refer to Figure 1 , which shows a schematic structural diagram of the wind speed correction model provided by an exemplary embodiment of the present disclosure. The wind speed correction model includes: an input module 11, an embedding module 12, an encoder module 13, and a post-processing module 14.
[0082] The input module 11 is used to obtain meteorological variable data, which are generated by a preset regional weather forecasting model and are used to indicate the temporal variation of multiple meteorological variables related to wind speed in a target area over a period of time in the future.
[0083] The embedding module 12 is used to perform an information embedding operation on the meteorological variable data to generate a target embedding tensor.
[0084] The embedding module 12 can also be referred to as an Embedding module. Essentially, it is an embedding layer. Through the embedding operation, the model can map complex meteorological variable data to a low-dimensional embedding space, thereby providing a more easily processed feature representation for the subsequent encoding process.
[0085] The encoder module 13, that is, the encoder, is the core part of the model. The encoder module 13 is used to encode the target embedding tensor by adopting an innovative tower-shaped local hierarchical attention mechanism to obtain a predicted feature tensor. The tower-shaped local hierarchical attention mechanism means that the attention calculation for each time node in the target embedding tensor is restricted to a limited number of time nodes associated with that time node. This mechanism can effectively capture key information in the local time series while avoiding the computational complexity problem that may be brought by the global attention mechanism.
[0086] The post-processing module 14 is used to perform feature transformation on the predicted feature tensor to obtain the corrected wind speed forecast data.
[0087] Next, several exemplary embodiments are used to introduce the short- and medium-term wind speed correction method based on the regional weather forecasting model provided by the embodiments of the present disclosure.
[0088] Please refer to Figure 2, which shows a flowchart of a medium- and short-term wind speed correction method provided by an exemplary embodiment of the present disclosure. The method includes the following steps.
[0089] Step 201, obtain meteorological variable data, which is generated by a preset regional weather forecasting model and is used to indicate the temporal variation of multiple wind speed-related meteorological variables in a target area over a period of time in the future.
[0090] In some embodiments, the computing device can extract meteorological variable data related to wind speed from a preset regional weather forecasting model. This is usually achieved by calling the application programming interface (API) of the regional weather forecasting model to read the data.
[0091] The meteorological variable data is sourced from a regional weather forecasting model, which is a forecasting system established based on a large amount of meteorological data, physical models, and statistical analysis. Its working principle is to collect observational data of numerous meteorological elements including temperature, pressure, humidity, etc., and through complex meteorological physical equations (such as the equations of atmospheric motion, thermodynamics equations, etc.) and statistical relationships, simulate and predict the state of the atmosphere, thereby generating meteorological variable data. For example, the regional weather forecasting model is the Weather Research and Forecasting Model (WRF), which simulates the weather state of the target area by inputting initial conditions (such as atmospheric pressure, temperature, humidity, wind speed, and direction, etc.) and boundary conditions (usually obtained from global models such as the Global Forecast System (GFS) or the European Centre for Medium-Range Weather Forecasts (ECMWF)).
[0092] Meteorological variable data includes a data set used to describe the time-series changes of multiple meteorological variables related to wind speed in a target area over a period of time in the future. In some embodiments, the meteorological variable data includes the following: 1. Meteorological variables: Also known as regional weather forecast model variables, these variables are the results output by the regional weather forecast model and are used to describe the weather conditions in the target area. Variables related to wind speed may include wind speed, wind direction, temperature, humidity, air pressure, etc. These variables usually exist in the form of a time series, and each time point has a corresponding variable value. They can be scalars (such as air temperature and air pressure) or vectors (such as wind speed and wind direction). 2. Time index: Each meteorological variable at a moment corresponds to a time index, which is used to identify the time order of the meteorological variables. The time index can be a specific date and time or a relative time (such as the forecast moment). For example, the time index is a timestamp, and the format can be "year-month-day hour:minute:second", such as "2025-03-09 14:00:00". These meteorological variables are usually collected or calculated simultaneously, so their time indexes are the same. For example: At a specific time point, meteorological variables such as wind speed, wind direction, temperature, humidity, and air pressure all have corresponding values.
[0093] Meteorological variable data can have the following data characteristics: 1. Temporality: Meteorological variable data is arranged in chronological order and can reflect the change trend of meteorological variables over a period of time in the future. 2. Multivariate nature: In addition to wind speed, it can also include other meteorological variables, such as air pressure, temperature, humidity, etc. There may be relationships between these variables. 3. Spatiality: The data is generated for a specific target area and has a clear spatial range.
[0094] In some embodiments, meteorological variable data can be presented in two main formats: spatial distribution maps and time series graphs, which respectively reveal the distribution of meteorological variables in the spatial and temporal dimensions and their changing trends. Spatial distribution maps can be in the form of raster maps or vector maps, which are used to visually display the distribution of meteorological variables in the geographical space. Such maps are usually based on a geographic coordinate system, that is, longitude and latitude are used to ensure the accuracy of the geographical location of the data. Through color coding or vector arrows, spatial distribution maps can clearly represent the intensity and direction of various meteorological variables, such as the magnitude and direction of wind speed. The depth of color or the density of arrows may correspond to the strength of the wind speed, and the direction of the arrows intuitively indicates the direction of wind flow. A time series graph is a data format that shows the change of meteorological variables over time. In this graph, the horizontal axis (X-axis) represents time, and its unit can be hours, days, months, etc., depending on the research needs and data resolution. The vertical axis (Y-axis) represents the specific values of meteorological variables, such as key meteorological parameters like temperature, wind speed, and air pressure. The time series graph depicts the evolution trend of these meteorological variables over time in the form of curves or broken lines, enabling observers to clearly identify the increase or decrease, periodic fluctuations, or other time-related characteristics of the variables at a glance. The embodiments of the present disclosure do not limit this.
[0095] Step 202: Embed information into the meteorological variable data to generate a target embedding tensor.
[0096] The target embedding tensor is a tensor generated by embedding information into the meteorological variable data. Information embedding is to map the original data into a low-dimensional vector space so that the embedded tensor can retain the important information and features of the original data, facilitating subsequent processing and analysis.
[0097] In some embodiments, the computing device can integrate the time information, location information, and physical feature information in the meteorological variable data into a high-dimensional tensor, namely the target embedding tensor, for subsequent model processing and feature extraction.
[0098] It should be noted that the information embedding method can refer to the relevant descriptions in the following embodiments and will not be introduced here first.
[0099] Step 203: Use a tower-shaped local hierarchical attention mechanism to encode the target embedding tensor to obtain a predicted feature tensor. The tower-shaped local hierarchical attention mechanism is used to indicate that the attention calculation for each time node in the target embedding tensor is constrained to a finite number of time nodes associated with the time node.
[0100] The tower - type local hierarchical attention mechanism is an improved attention mechanism used to encode the target embedding tensor. By means of layering, the attention calculation at each time node is constrained to a finite number of time nodes associated with that time node, so as to better capture the physical characteristics of wind speed at different time scales. This mechanism can capture local information more effectively and avoid the problems of high computational complexity and overly dispersed attention distribution in the global attention mechanism.
[0101] A time node refers to a time step in the target embedding tensor, which represents the characteristic representation of meteorological variables at a specific time point. The target embedding tensor is a multi - dimensional array, where each row (or column, depending on the shape of the tensor) usually corresponds to a time step and contains the characteristics of all relevant meteorological variables at that time step.
[0102] For each time node, the finite number of associated time nodes are at least one time node adjacent at different time scales. That is, the associated nodes of each time node include not only the adjacent nodes at the same time scale, but may also include the adjacent nodes at finer or coarser time scales. This multi - level association definition enables the model to capture both short - term dynamics and long - term trends simultaneously. The time scale refers to the granularity or resolution of time data. For example, a fine time scale may correspond to minute - level or hour - level data for capturing short - term dynamic changes. The same time scale may correspond to daily - level data for capturing the change trends within a medium term. A coarse time scale may correspond to weekly - level or monthly - level data for capturing long - term trends.
[0103] The predicted feature tensor is the feature representation obtained by encoding the target embedding tensor through the tower - type local hierarchical attention mechanism. It is a multi - dimensional tensor that contains the feature information of the meteorological variable data after encoding processing and is used for subsequent wind speed prediction.
[0104] It should be noted that the encoding method using the tower - type local hierarchical attention mechanism can refer to the relevant description in the following embodiments and will not be introduced here first.
[0105] Step 204: Perform feature transformation on the predicted feature tensor to obtain the corrected wind speed prediction data.
[0106] In some embodiments, the computing device can further process the predicted feature tensor, perform feature transformation on it, and convert it into the corrected wind speed prediction data. Feature transformation can include some linear or non - linear transformation operations.
[0107] The corrected wind speed forecast data is the final wind speed forecast result obtained by performing feature transformation on the predicted feature tensor, and is used to indicate the wind speed change in the target area over a period of time in the future. The corrected wind speed forecast data includes multiple wind speed values in the target area over a period of time after model adjustment. That is, the final output is the wind speed forecast data processed and corrected by the model. These data are more accurate than the wind speed data generated by the original weather forecasting model because the model utilizes the complex relationships between meteorological variables and the attention mechanism to improve the forecast results.
[0108] After the above series of processes, the corrected wind speed forecast data can be further applied to fields such as meteorological services, aviation and maritime safety guarantee, and wind power generation. The meteorological service department can issue more accurate weather warnings based on the accurate wind speed forecast. The aviation and maritime fields can reasonably plan routes and arrange flights according to the wind speed forecast. Wind power generation enterprises can optimize the power generation plan based on the wind speed forecast to improve power generation efficiency.
[0109] The following is a further explanation of the three key links of information embedding, encoding, and feature transformation in the above steps.
[0110] I. Information embedding.
[0111] In the information embedding stage, the goal is to convert the meteorological variable data into a comprehensive target embedding tensor. In some embodiments, the information embedding method, i.e., step 202 above, can be replaced and implemented as the following steps, such as Figure 3 shown below:
[0112] Step 301, convert the position index of each time node in the meteorological variable data into a position encoding vector.
[0113] The meteorological variable data belongs to time series data, and each time node has an index value, i.e., the position index, which is used to indicate the position of the time node in the sequence. For example, in a sequence containing 96 time nodes, the position index of the first time node is 0, the position index of the second time node is 1, and so on. The position index of the last time node is 95.
[0114] Position encoding is a method of converting the position index into a high-dimensional vector, and its purpose is to enable the model to better understand the position information of each time node in the time series. The position encoding vector is a vector used to represent position information. The position encoding vector can provide the model with the relative and absolute position information of each time node, thereby enhancing the model's understanding of the time series structure.
[0115] Position encoding can use trigonometric functions (usually sine and cosine functions) to convert the position index into a high-dimensional vector. This encoding method can capture the periodicity and relativity of position information.
[0116] Therefore, when processing meteorological variable data, the position index of each time node (e.g., which time point in the sequence) can be converted into a position encoding vector by means of trigonometric functions (e.g., sine function). This encoding method can not only provide the model with the position information of each time node in the time series, but also help the model capture the periodicity and relative position information of the time series, thus significantly improving the model's ability to understand the time series structure.
[0117] Step 302: Convert the time stamps of each time node in the meteorological variable data into time feature vectors.
[0118] When processing meteorological variable data, each time node is attached with a time stamp, which is also called a time mark and is a specific time point, such as "14:30, March 26, 2025".
[0119] The time feature vector is a numerical vector containing multiple time features. It maps different time attributes (such as hours, day of the week, date, season, etc.) in the time stamp to different dimensions, thus providing richer time series information for the model. This vector conversion can help the model capture the periodic patterns in the time series and the mutual relationships between time attributes.
[0120] The generation process of the time feature vector is as follows:
[0121] 1. Extract time attributes: Extract different time attributes from the time stamp of each time node, such as the hour of the day, the day of the week, the day of the month, the day of the year, etc.
[0122] 2. Vector conversion: Map these time attributes to different dimensions respectively to form a high-dimensional vector. For example, the hour of the day can be mapped to the first dimension, the day of the week to the second dimension, the day of the month to the third dimension, and the day of the year to the fourth dimension.
[0123] Through this vector conversion, the obtained time feature vector contains time information in four dimensions, namely the hour of the day, the day of the week, the day of the month, and the day of the year. The embedding of this time feature vector can significantly improve the model's ability to understand the time series structure and help the model better capture the periodicity and trend patterns in the time series.
[0124] Therefore, the purpose of time feature extraction is to convert the time attributes (such as hours, dates, seasons, etc.) in the time stamps into time feature vectors, providing richer temporal information for the model. This vector conversion can help the model better understand the periodic patterns in the time series and the interrelationships between time attributes, thereby improving the model's prediction performance.
[0125] Step 303: Extract meteorological variable data with different time resolutions from the meteorological variable data at the original time resolution, and splice the meteorological variable data with different time resolutions to obtain a meteorological variable tensor.
[0126] Time resolution refers to the time interval between two adjacent data points in time series data. It reflects the density of data in the time dimension. The higher the time resolution, the shorter the interval between data points, and the more detailed the time changes that can be captured. For example, in meteorological data, if the time resolution is 15 minutes, it means that data is recorded every 15 minutes. In meteorological data analysis, multi-time resolution data extraction refers to extracting data with different time resolutions from the data at the original time resolution. For example, extracting 1-hour and 4-hour resolution data from 15-minute resolution data. These data with different resolutions can provide characteristic information at different time scales, helping the model better capture long-term trends and short-term changes in the time series.
[0127] Among them, the data extraction methods can include but are not limited to the following two methods:
[0128] One method is direct extraction according to the timestamp: Based on the preset timestamp, extract data from the meteorological variable data at the original time resolution at a fixed time interval, so as to extract data with different time resolutions. For example, when extracting 1-hour resolution data, directly select every 4 time points of data; when extracting 4-hour resolution data, directly select every 16 time points of data.
[0129] Another method is to use a convolutional neural network:
[0130] 1. Define the convolutional kernel: According to the target time resolution, design a convolutional kernel for the corresponding time interval. The size of the convolutional kernel determines the time range covered by each convolution operation in the original time series. For example, if the goal is to extract 1-hour resolution data, the size of the convolutional kernel can be set to 4 (corresponding to 60 minutes / 15 minutes).
[0131] 2. Perform convolution operations: Use a one-dimensional convolutional neural network to perform convolution operations on the original time series data. The convolution operation slides the convolutional kernel on the original time series and calculates the eigenvalue within each time window. These eigenvalues can capture the local patterns and trends within the time window.
[0132] 3. Extract representative features: The result of the convolution operation is the representative features at each time node. These features not only contain the information of the original data but also extract the local features within the time window through the convolution operation, thus enhancing the data representation ability. For example, assume the original data has a 15-minute resolution and the time series length is 96 time points (24 hours). If we want to extract data with a 1-hour resolution, we perform a convolution operation using a convolution kernel of size 4, and finally obtain the representative features of 24 time points.
[0133] In a convolutional neural network, an automatic segmentation and fusion technique based on sparse matrices can be used. This is mainly to improve the computational efficiency and optimize the storage when dealing with sparse data, while extracting meteorological variable data with different time resolutions. In a convolutional neural network, after the input data is processed by the convolutional layer, a feature map with enhanced sparsity may be generated. By using the storage format of the sparse matrix, only the non-zero elements and their position information need to be stored, thus significantly reducing the memory occupancy. Further, through the automatic segmentation of the sparse matrix, the feature map can be divided into multiple sub-blocks, and each sub-block only contains the non-zero elements and their neighborhood information. Then, through the fusion operation, these sub-blocks are recombined to form a new feature map. This method can effectively reduce the computational amount while retaining the key features.
[0134] With the help of the automatic segmentation and fusion technique of sparse matrices, a convolutional neural network can efficiently process meteorological variable data with different time resolutions. For example, small-sized convolution kernels can be used to perform convolution operations on the original time-resolution data to extract local features. Subsequently, through the segmentation and fusion operations of the sparse matrix, these local features are combined into feature maps with different time resolutions. This process not only makes full use of the storage advantages of the sparse matrix but also realizes the efficient recombination of features through the segmentation and fusion operations, providing strong technical support for the multi-time-scale analysis of meteorological variables.
[0135] Use a one-dimensional convolutional neural network to perform convolution operations on the meteorological variable data with the original time resolution, and extract the representative features of each time node through the convolution kernel to extract meteorological variable data with different time resolutions. This method can more flexibly capture the local features in the time series rather than simply sampling.
[0136] Schematically, assume that the original meteorological variable data is recorded at a 15 - minute resolution, that is, the data is recorded every 15 minutes. For a 24 - hour time series, there will be 96 time points for such data (24 hours × 4 time points / hour = 96 time points). From the original 15 - minute resolution data, data is extracted every 4 time points to obtain 1 - hour resolution data. In this way, the length of the 24 - hour time series is reduced from 96 time points to 24 time points (24 hours × 1 time point / hour = 24 time points). Similarly, from the original 15 - minute resolution data, data is extracted every 16 time points to obtain 4 - hour resolution data. In this way, the length of the 24 - hour time series is reduced from 96 time points to 6 time points (24 hours ÷ 4 hours / time point = 6 time points). Finally, the data of these three resolutions are concatenated together to form a new three - dimensional tensor. The shape of the concatenated meteorological variable tensor is: [time series length, time resolution dimension, feature dimension], for example, [96, 3, feature dimension], where 96 represents the original time series length, 3 represents the three time resolutions, and the feature dimension represents the number of features at each time point.
[0137] Therefore, data with different time resolutions can be extracted from the meteorological variable data of the original time resolution, which can be achieved in two ways: directly extracting instantaneous data according to the timestamp or using a one - dimensional convolutional neural network to extract representative features. Finally, the data with different time resolutions are concatenated along a new data dimension to form a multi - scale spatio - temporal information tensor, that is, the meteorological variable tensor. The meteorological variable tensor is a multi - dimensional array used to represent meteorological - related data. The fusion of this multi - scale information can provide richer temporal features for the model and enhance the model's ability to understand the time series structure.
[0138] It should be noted that the above steps 301, 302, and 303 can be executed in parallel or sequentially in a certain order. The present disclosure does not limit the specific execution order of these steps.
[0139] Step 304, concatenate the position encoding vector, the time feature vector, and the meteorological variable tensor to obtain the target embedding tensor.
[0140] By combining the position information (position encoding vector), the time information (time feature vector), and the physical feature information (meteorological variable tensor), a tensor (target embedding tensor) containing multiple features is formed. The target embedding tensor can comprehensively reflect the multi - dimensional features of each time node, providing extremely rich input data for subsequent encoding and feature transformation steps, thus laying a solid foundation for the efficient learning and accurate prediction of the model.
[0141] II. Encoding.
[0142] The core of the encoding stage is to adopt a tower-shaped local hierarchical attention mechanism to encode the target embedding tensor and obtain a predicted feature tensor. The tower-shaped local hierarchical attention mechanism aims to make the attention only focus on the limited number of node information related to each node around it, so as to avoid the interference and computational burden that may be brought by global attention. In some embodiments, the encoding method, i.e., step 203 above, can be replaced and implemented as the following steps, as Figure 4 shown:
[0143] Step 401, determine the set of hierarchical relationships for each time node of the target embedding tensor. The set of relationships includes a set of limited time nodes associated with the time node at multiple levels, and the levels are divided according to different time scales.
[0144] For each time node of the target embedding tensor, determine the limited number of time nodes associated with it at multiple levels. The levels are divided according to different time scales, and there is an interaction relationship between the time node and the associated time nodes in each level. Different levels correspond to different time scales. For example, the first level corresponds to "1 hour", the second level corresponds to "1 day", and the third level corresponds to "1 week". "At a certain level" means within the target time period range before and after the current node, and the target time period is defined by the level, such as 1 hour, 1 day, or 1 week. The interaction relationship refers to the interaction, influence, or connection that exists between different time nodes. In time series analysis, this relationship can be reflected as the dependence, correlation, or causal relationship between time nodes.
[0145] The limited number of associated time nodes are selected according to the time scale and the predefined hierarchical relationship, so as to limit the attention calculation within a local range, reduce the computational burden and avoid global noise interference. Further, for the current time node, the process of determining the set of hierarchical relationships corresponding to this time node is as follows: First, determine one or more time nodes associated with it in each level, so as to determine the relationship subset of each level. The relationship subset of each level is used to indicate one or more time nodes associated with the current time node in that level. Then, merge the relationship subsets of multiple levels to obtain the set of hierarchical relationships of the time node.
[0146] Optionally, the set of hierarchical relationships of the time node includes the union of the relationship subsets of multiple levels. The relationship subsets of multiple levels can include the following types:
[0147] (1) A set of child relationships, used to indicate one or more (e.g., two) time nodes that are closest to the time node in the first level (if any). For example, if the current time node is at the daily level, the set of child relationships may include the time nodes of the previous hour and the next hour.
[0148] (2) A set of sibling relationships, used to indicate one or more (e.g., two) time nodes that are closest to the time node in the second level. For example, if the current time node is at the daily level, the set of sibling relationships may include the time nodes of the previous day and the next day.
[0149] (3) A set of parent relationships, used to indicate one time node that is closest to the time node in the third level (if any). For example, if the current time node is at the daily level, the set of parent relationships may include the time nodes of the previous week and the next week.
[0150] Wherein, the distance between time nodes is the time difference. In the set of hierarchical relationships of time nodes, the relationship subsets of different levels represent the closest time nodes within different time ranges. The time scales corresponding to the first level, the second level, and the third level are increasing respectively. The time scale of the first level is defined as a fine time scale, which focuses on capturing the subtle changes and dynamics in the short term within the time series; the time scale of the second level is the same time scale, which focuses on analyzing the interactions and relationships within the same time period as the current time node; the time scale of the third level is a coarse time scale, which is used to identify and understand the long-term trends and background changes in the time series. Through this hierarchical time scale division, the data characteristics within different time ranges can be analyzed and processed more precisely.
[0151] Schematically, the local information of each time node (the l-th time node in the s-th level, where s = 1, ……, S represents the hierarchical progression from the finest to the coarsest level) is divided into three levels according to the time granularity: the set of child relationships the set of sibling relationships and the set of parent relationships Wherein, the local information of the time node refers to the set of hierarchical relationships associated with the time node. The set of child relationships corresponds to the child level (i.e., a finer time scale), and focuses on the two time nodes that are closest to the current time node at this time scale; the set of sibling relationships corresponds to the same time scale as the current time node, and focuses on the two time nodes that are closest to the current time node at this time scale; while the set of parent relationships For the corresponding parent level (i.e., coarser time scale), the focus is on the time node closest to the current time node at that time scale, so as to capture long-term trends or overall background changes. Among them, the set of hierarchical relationships can be expressed as:
[0152]
[0153] Among them, represents the set of all time nodes associated with the time node l in the s-th level. represents the set of adjacent time nodes at the same level as the time node l in the s-th level. represents the set of adjacent time nodes at the sub-level (i.e., finer time scale) of the time node l in the s-th level. represents the set of adjacent time nodes at the parent level (i.e., coarser time scale) of the time node l in the s-th level. represents the j-th time node in the s-th level. A represents the number of adjacent time nodes that each time node in the set of same-level relationships can focus on, generally taking 3. L represents the length of the historical sequence, that is, the total number of time nodes. M represents the branching factor of the multi-way tree, which determines the number of fine-level nodes included in the coarse-level node. S represents the total number of levels. M s-1 represents the product of the branching factors from the finest level to the s-th level, which is used to calculate the index range of the time nodes in the s-th level. represents an empty set. When the condition is not met, the corresponding relationship set is empty.
[0154] Step 402, perform multi-head attention calculation on the hierarchical relationship sets of each time node to obtain the attention result.
[0155] Perform multi-head attention calculation on the hierarchical relationship sets of each time node in the target embedding tensor to obtain the attention result of each time node. The multi-head attention mechanism divides the input query, key, and value into multiple heads, calculates the attention result of each head respectively, and then splices and projects these results into a common dimension. The multi-head attention mechanism can capture the interaction relationships between time nodes from multiple perspectives and enhance the model's ability to understand complex time series patterns.
[0156] In the embodiments of the present disclosure, in order to reduce the computational complexity and improve the model's ability to capture local features, an innovative local attention mechanism is introduced. Through this mechanism, after the hierarchical relationship sets of each time node are processed in parallel by the multi-head attention mechanism, the complex multi-dimensional time interaction features between each level can be effectively extracted and integrated.
[0157] Specifically, this local attention mechanism calculates the similarity between the query vector and the key vector only within a predefined adjacency set, thus avoiding full-scale calculations for the entire sequence. This approach not only reduces the computational load but also enhances the model's ability to extract local features. For each target position, the system filters out multiple candidate positions from its adjacency set and calculates the corresponding attention weights. Among them, the target position is also called the target node, such as time node i. The candidate position is also called the candidate node, such as time node l. Schematically, the hierarchical relationship set of each time node The calculation formula for input multi-head attention is as follows:
[0158]
[0159] Where is the attention retrieval vector corresponding to the target node. represents the feature index vector of the candidate node, which is used to calculate the attention matching degree with the query vector. is the actual information corresponding to the position of the key vector, which is used as the weighted content of the attention output. The attention calculation is only performed on a finite adjacency set i.e., the target node i only calculates the similarity with the keys of the relevant nodes in the set and performs weighted aggregation based on their values. d K is the dimensional constant of the key vector, which is used to scale the dot product result to prevent the gradient from vanishing due to excessive numerical values. exp is the exponential function, which is used to calculate the attention weights. q i is the query vector of time node i, is the transpose of the key vector of time node l, v l is the value vector of time node l.
[0160] This design enables the model to focus more on the local context information of the target position, thus significantly improving the efficiency and representation ability of information aggregation. In this way, the model can more accurately capture the subtle features in time series data while maintaining computational efficiency, providing a more effective method for time series analysis.
[0161] Step 403, perform preset processing on the attention result to obtain the predicted feature tensor.
[0162] Perform preset processing on the attention result of each time node to obtain the final output of each time node. Combining the final outputs of all time nodes, the complete predicted feature tensor is obtained. The predicted feature tensor is the output of the encoding stage, which integrates the feature information of each time node in the target embedding tensor and is constrained by the tower-shaped local hierarchical attention mechanism, enabling the model to more effectively capture local and global temporal features.
[0163] The preset processing is a standardized and modular operation process. Its core purpose is to perform in-depth feature extraction and stabilization processing on the attention results at all time nodes, thereby significantly enhancing the overall performance and stability of the model. The process of preset processing may include but is not limited to the following steps: 1. Perform residual connection and normalization through the residual and normalization layer: Connect the attention result obtained by the multi-head attention calculation with the input in a residual manner. Here, the "input" refers to the input of the multi-head attention module, that is, the data before entering the multi-head attention module. The purpose of the residual connection is to directly add the input to the output of the attention module to maintain the integrity of the input information and avoid information loss or gradient disappearance in the deep network. Perform layer normalization on the result after the residual connection. Layer normalization is a normalization technique that normalizes all features of each sample, making the features of each sample have a distribution with a mean of 0 and a standard deviation of 1. This helps to stabilize the training process and accelerate convergence. 2. Perform non-linear transformation through the feed-forward neural network: Pass the normalized result into the feed-forward neural network for non-linear transformation. The feed-forward neural network usually consists of two linear layers and a non-linear activation function (such as the ReLU function). This non-linear transformation can further extract features and increase the expressive power of the model. 3. Perform residual connection and normalization again through the residual and normalization layer: The output of the feed-forward neural network will be connected with the normalized attention output in a residual manner again, and then layer normalization is performed on the final output. This step further ensures the integrity of the information and stabilizes the training process. After the above processing, the final output of each time node is a part of the predicted feature tensor. Combining the final outputs of all time nodes, the complete predicted feature tensor is obtained. The predicted feature tensor integrates the local and global information in the attention results of all time nodes and is the final output of the encoding stage, which is used for subsequent feature transformation.
[0164] III. Feature transformation.
[0165] The goal of the feature transformation stage is to extract the target features from the predicted feature tensor and convert them into the corrected wind speed forecast data. In some embodiments, the feature transformation method, i.e., the above step 204, can be alternatively implemented as the following steps, such as Figure 5 shown as
[0166] Step 501, extract the target features from the predicted feature tensor through the aggregation layer.
[0167] The aggregation layer can aggregate and filter the features in the predicted feature tensor to extract the features related to the wind speed forecast.
[0168] The target features refer to one or more features related to wind speed extracted from the predicted feature tensor. After being processed by the aggregation layer, these features can better reflect the law of wind speed change, thus providing more valuable information for subsequent linear transformation and wind speed correction.
[0169] Step 502, perform a linear transformation on the extracted target features through a linear layer to obtain the corrected wind speed forecast data.
[0170] The linear transformation can map the extracted target features to the output space of the wind speed forecast, thereby generating the final wind speed forecast result.
[0171] The linear transformation is usually used to perform a weighted sum on the target features and add a bias term to achieve the transformation and adjustment of the target features. The output of the linear layer can be expressed as y = Wx + b, where x is the target feature, and W and b are learnable parameters.
[0172] Through the processing of the above three stages of information embedding, encoding, and feature transformation, this method can make full use of the multi-dimensional information in the meteorological variable data and combine the powerful modeling ability of the tower-shaped local hierarchical attention mechanism to generate high-precision corrected wind speed forecast data.
[0173] Please refer to Figure 6 , which shows a schematic diagram of the principle of a medium-term and short-term wind speed correction method based on a regional weather forecast model provided by an exemplary embodiment of the present disclosure. This method is executed based on a pre-trained wind speed correction model, and the wind speed correction model includes: an input module 11, an embedding module 12, an encoder module 13, and a post-processing module 14.
[0174] 1. The input module 11 is used to obtain the meteorological variable data output by the single-point WRF model, which includes multiple meteorological variables and the time index of each moment. This information will be input into the embedding module 12.
[0175] 2. The embedding module 12 is used to perform information embedding on the meteorological variable data, involving the embedding of three types of information: (1) Position encoding: Use trigonometric functions to convert the position index of each time node in the meteorological variable data into a position encoding vector. (2) Time feature annotation: Convert the time mark of each time node in the meteorological variable data into a time feature vector. (3) Multi-time resolution data extraction: Extract meteorological variable data with different time resolutions from the original data (i.e., the meteorological variable data with the original time resolution), and splice the meteorological variable data with different time resolutions to obtain a meteorological variable tensor.
[0176] 3. The encoder module 13 is the core part of the model. It is used to encode the target embedding tensor by adopting an innovative tower-shaped local hierarchical attention mechanism to obtain the predicted feature tensor. The tower-shaped local hierarchical attention mechanism aims to make the attention only focus on the limited node information related to each node around it, thus avoiding the interference and computational burden that may be brought by global attention.
[0177] The encoder module 13 not only includes the tower-shaped local hierarchical capture and multi-head attention mechanisms, but also integrates some structures of classical Transformer models such as residual and normalization layers, feed-forward neural networks, and residual and normalization layers, further improving the performance and stability of the model.
[0178] 4. The post-processing module 14 is used to perform feature transformation on the predicted feature tensor to obtain the corrected wind speed forecast data. The post-processing module 14 includes an aggregation layer and a linear layer. Through the operations of these layers, the model can transform the predicted feature tensor generated by the encoder module 13 into the final wind speed correction result. This process not only effectively transforms and adjusts the features, but also ensures the accuracy and reliability of the output result.
[0179] It should be noted that each module provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. Its specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be elaborated here.
[0180] In some embodiments, the computing device can obtain training data, where the training data includes multiple wind speed forecast values and corresponding multiple wind speed observation values; calculate the loss function according to the multiple wind speed forecast values and corresponding multiple wind speed observation values; and iteratively update the model parameters through a preset algorithm according to the loss function to train and obtain the wind speed correction model.
[0181] The multiple wind speed forecast values are the wind speed values predicted by the initial wind speed correction model. The corresponding multiple wind speed observation values are the actually observed wind speed values, which are usually measured by instruments (such as anemometers) at meteorological observation stations.
[0182] The loss function is a function used to measure the difference between the model prediction value (wind speed forecast value) and the true value (wind speed observation value). Its purpose is to quantify the accuracy of the model prediction. The smaller the value of the loss function, the closer the wind speed predicted by the model is to the actually observed wind speed, and the better the performance of the model.
[0183] A preset algorithm usually refers to an optimization algorithm, such as the gradient descent algorithm. The purpose of the optimization algorithm is to continuously reduce the value of the loss function by adjusting the parameters of the model, thereby improving the prediction accuracy of the model. Taking the gradient descent algorithm as an example, it calculates the gradient (i.e., partial derivative) of the loss function with respect to each model parameter, and then updates the parameters along the opposite direction of the gradient. After the above process, the model parameters are continuously optimized, and the finally trained model is the wind speed correction model. The role of this wind speed correction model is to be able to predict the wind speed more accurately. It can output the corrected wind speed prediction value according to the input (such as time, meteorological conditions, etc.), and this prediction value is closer to the actually observed wind speed value than the original wind speed forecast value.
[0184] However, traditional loss functions usually only use metrics such as MSE or Root Mean Square Error (RMSE) to measure the difference between the forecast value and the error value. Among them, the calculation formula of MSE is:
[0185]
[0186] where, is the i-th wind speed forecast value, y i is the i-th wind speed observation value, n is the total number of samples, that is, the number of wind speed observation values or wind speed forecast values, and both i and n are positive integers.
[0187] In fact, MSE imposes a heavier penalty on larger deviations from the observed values, thus inhibiting the model from making bold prediction adjustments. Therefore, although the output results perform very well in metrics such as RMSE and MSE, showing small error values, from the details of the forecast sequence, there is often an over-smoothing phenomenon. In view of the problem of overly smooth forecast results in the artificial intelligence correction process in related technologies, the embodiments of the present disclosure introduce an "activity" index and incorporate it into the loss function of the wind speed correction model, optimizing the deficiencies of the traditional loss function.
[0188] The new loss function provided by the embodiments of the present disclosure includes an activity difference term, and the activity difference term is used to indicate the difference in the degree of fluctuation between multiple wind speed forecast values and the corresponding multiple wind speed observation values. The activity difference term can be determined based on the difference between the observed activity and the forecast activity. The observed activity is used to indicate the degree of fluctuation of multiple wind speed observation values, and the forecast activity is used to indicate the degree of fluctuation of multiple wind speed forecast values. For example, the observed activity is the standard deviation of multiple wind speed observation values, the forecast activity is the standard deviation of multiple wind speed forecast values, and the activity difference term is the square of the difference between the observed activity and the forecast activity. That is, the definition of the forecast activity A forcast is:
[0189]
[0190] Observation activity A obs is defined as:
[0191]
[0192] wherein, can be the average of n wind speed observation values or the historical observation average. The historical observation average refers to the average of multiple wind speed observation values in the target area in history (for example, several years or decades). When the difference between the observation activity and the forecast activity is very small, it indicates that the model has a good effect.
[0193] Therefore, the new loss function provided by the embodiments of the present disclosure can be jointly composed of MSE and the activity index, which can not only reduce the difference between the forecast and the actual value but also make the active amplitude of the forecast result close to the observation. That is, the loss function includes MSE and the activity difference term. MSE is used to indicate the average of the squares of the differences between multiple wind speed forecast values and multiple wind speed observation values, and the activity difference term is the square of the difference between the observation activity and the forecast activity. Schematically, after multiplying the activity difference term by a weight factor and adding it to MSE, the final loss function is obtained.
[0194]
[0195] where λ is a preset weight factor used to balance the relative importance of MSE and the activity difference term in the final loss function, and A obs1 can be a preset value or the historical observation activity. The historical observation activity refers to the result obtained by analyzing and calculating the volatility of the wind speed observation values in the target area in history (for example, several years or decades).
[0196] In the related art, the deficiencies of the medium and short-term wind speed correction method and the loss function have been widely discussed. To overcome the defect of relying on cumbersome signal processing (such as variational mode decomposition) for feature engineering in the related art, the embodiments of the present disclosure propose a medium and short-term wind speed prediction sequence correction model based on an encoder architecture. The model realizes the efficient extraction of input data features through the attention coding mapping technology of multi-time scale information embedding and tower-type local level capture. At the same time, aiming at the problem that the forecast result is too smooth in the artificial intelligence correction process, the embodiments of the present disclosure introduce the "activity" index and incorporate it into the loss function to optimize the deficiencies of the traditional loss function. The medium and short-term wind speed correction method based on the new loss function proposed by the embodiments of the present disclosure significantly improves the accuracy of medium and short-term wind speed prediction in the field of wind power generation, establishes a more robust and reliable artificial intelligence model, and has broad application prospects.
[0197] Taking four actually operating wind farms in the Gansu region as an example, the embodiments of the present disclosure propose a wind speed correction method based on an attention encoding mapping technology with tower-type local hierarchical capture, aiming to correct the wind speed forecasts for the target wind farm for 1 to 3 days. The implementation steps of this method are as follows:
[0198] 1. Data preparation.
[0199] (1). Determine the correction scenario and forecast lead time: The wind speed correction scenario of this embodiment focuses on four operating wind farms in the Yumen area of Gansu. The goal is to correct the wind speeds of the 1-3 day regional numerical weather prediction models for these four wind farms.
[0200] (2). Obtain data: Obtain long-term observational data for the target area, and obtain the corresponding mesoscale weather forecast data and its timestamp sequence. The long-term observational data includes the hub height wind speed observational data for 3 years of the target wind farm and the total power generation data of each wind farm. The mesoscale weather forecast data includes the regional forecast interpolation data for 0-72 hours per day corresponding to the historical observation period obtained by using the WRF model through the historical model back-calculation method, with a time granularity of 15 minutes. In this process, the back-calculated meteorological variables cover multiple relevant meteorological variables such as wind speed, wind direction, air temperature, air pressure, geopotential height, absolute vorticity (avo), potential vorticity (pvo), and boundary layer meteorological variables at each layer, and the time record of each moment is retained.
[0201] 2. Model construction.
[0202] Construct a wind speed correction model including an input module, an embedding module, an encoder module, and a post-processing module. The specific modeling process is as follows:
[0203] (1). Model segmentation: For three different time intervals of 0-24 hours, 24-48 hours, and 48-72 hours, establish corresponding models respectively.
[0204] (2). Input module processing: Input all meteorological variables into the embedding module. The time series length of the input data is the first 96 time steps including the target forecast time step. On this basis, multi-time resolution extraction is performed. In this embodiment, instantaneous data is directly extracted using timestamps, and extraction is performed every 4 time steps, thus forming three time resolution data of 15 minutes, 1 hour, and 4 hours. Subsequently, by expanding a new data dimension, the two-dimensional data is concatenated to form a new three-dimensional tensor.
[0205] (3) Embedding module processing: Use trigonometric functions to encode the position of each group of variables to obtain a new set of vectors. At the same time, for the timestamp, calculate and retain the hour of the day, the day of the week, the day of the month, and the day of the year, respectively, and then obtain four columns of information vectors representing time features. Finally, the three data streams are combined and input into the encoder module.
[0206] (4) Encoder module processing: First, the local relationship related to each time node is processed. For example, at the 1-hour resolution layer, the 3rd hour is represented as the time node The node's hierarchical relationship set contains the sibling relationship set Child relationship collection And the parent relationship collection constitute
[0207] The complete hierarchical relationship set of the node. Each time node is deduced in this way, and the hierarchical relationship set of all time nodes is calculated.
[0208] All the tensors of the set are input into a layer of 8-head attention mechanism for attention solution. The attention results obtained are input into the normalization layer and feedforward neural network for subsequent calculations.
[0209] (5) Post-processing module processing: The output results of the encoder module are aggregated through the aggregation layer and input into the linear layer again for dimensionality reduction mapping to finally generate wind speed forecast data.
[0210] 3. Model training and model application.
[0211] (1) Combining data and models: Combine the data processed in step 1 with the wind speed correction model built in step 2.
[0212] (2) Loss function selection and training: Select the new loss function composed of MSE and activity index proposed in the embodiment of the present disclosure, and perform model training according to the AI modeling process. Input the wind speed forecast data together with the wind speed observation data and power observation data at the target forecast time into the new loss function for solution, and complete the training of the wind speed correction model by updating the gradient and iterative optimization.
[0213] (3) Model application: According to the above method, the meteorological variable data generated by the regional weather forecast model is input into the trained wind speed correction model to output the corrected wind speed forecast data.
[0214] In summary, the embodiments of the present disclosure propose an innovative medium- and short-term wind speed correction method and a loss function optimization method for addressing the shortcomings of artificial intelligence forecasting. Based on the regional numerical weather prediction model and the observed wind speed at the wind farm's anemometer tower, this method constructs an artificial intelligence correction model with a multi-time-scale feature capture pyramid hierarchical attention mechanism. Its specific application scenarios include medium- and short-term wind speed correction for regional power grids or individual wind farms, as well as the definition of a new loss function for the problem of data smoothing in artificial intelligence forecasting. Both aims to significantly improve the accuracy of medium- and short-term wind speed forecasting.
[0215] In the wind speed correction model provided by the embodiments of the present disclosure, the embedding module performs multi-time-scale information embedding and coding mapping on the input meteorological variable data, simplifying the complex, cumbersome, and data-length-dependent frequency decomposition method in traditional technologies to efficiently extract data features. The core technologies of the embedding module include fixed time interval extraction fusion and automatic segmentation fusion based on sparse matrices. These two information fusion methods can effectively integrate meteorological data at different time scales, laying a foundation for subsequent feature extraction. The encoder module introduces a tower-type local hierarchical capture attention mechanism improved based on the full attention mechanism. By hierarchically constraining the attention calculation, it focuses on the key features associated with the current time node, thereby more accurately capturing the physical laws of wind speed changes. Compared with directly using the wind speed output by the regional numerical weather prediction model, the wind speed correction method provided by the embodiments of the present disclosure significantly reduces the RMSE of the predicted wind speed; compared with the current correction technologies, the embodiments of the present disclosure further continuously reduce the RMSE of the predicted wind speed, improve the accuracy of wind speed forecasting, and enhance the performance of the wind speed forecasting correction technology.
[0216] In terms of loss function optimization, the embodiments of the present disclosure not only retain the traditional mean square error method but also introduce a weight index reflecting the activity degree of the data. This activity index can quantify the intensity of wind speed fluctuations, prevent the model output results from being too smooth, and thus significantly improve the accuracy and reliability of wind speed forecasting.
[0217] The embodiments of the present disclosure also provide a computing device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.
[0218] The embodiments of the present disclosure also provide a non-volatile computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, the steps of the above method are implemented.
[0219] The embodiments of the present disclosure also provide a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program. When the computer program is executed by the processor, the steps of the above method are implemented.
[0220] Figure 7 is a block diagram of a device 1900 shown according to an exemplary embodiment. For example, the device 1900 is used to execute the above method and may be provided as a server or a terminal device. Referring to Figure 7 , the device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0221] The device 1900 may also include a power supply component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958 (I / O interface). The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0222] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the device 1900 to complete the above method.
[0223] A computer-readable storage medium can be a tangible device that can hold and store programs / instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0224] The computer programs (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0225] A computer program (or computer program instructions) for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages - such as Smalltalk, C++, etc., and conventional procedural programming languages - such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, executed as a stand-alone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit may execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0226] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0227] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. The computer-readable program instructions can also be stored in a computer-readable storage medium, which causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0228] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0229] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0230] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A method for correcting medium- and short-term wind speeds based on a regional weather forecasting model, characterized in that, The method includes: Obtaining meteorological variable data, which is generated by a preset regional weather forecasting model and is used to indicate the temporal variation of multiple meteorological variables related to wind speed in a target area over a period of time in the future; Performing information embedding on the meteorological variable data to generate a target embedding tensor; Using a tower-shaped local hierarchical capture attention mechanism to encode the target embedding tensor to obtain a prediction feature tensor, where the tower-shaped local hierarchical capture attention mechanism is used to indicate that the attention calculation for each time node in the target embedding tensor is constrained to a finite number of time nodes associated with the time node; Performing feature transformation on the prediction feature tensor to obtain corrected wind speed forecast data.
2. The method according to claim 1, characterized in that, The performing information embedding on the meteorological variable data to generate a target embedding tensor includes: Converting the position index of each time node in the meteorological variable data into a position encoding vector; Converting the time stamp of each time node in the meteorological variable data into a time feature vector; Extracting meteorological variable data with different time resolutions from the meteorological variable data at the original time resolution, and splicing the meteorological variable data with different time resolutions to obtain a meteorological variable tensor; Splicing the position encoding vector, the time feature vector, and the meteorological variable tensor to obtain the target embedding tensor.
3. The method according to claim 2, wherein The extracting meteorological variable data with different time resolutions from the meteorological variable data at the original time resolution includes: According to a preset time stamp, extracting meteorological variable data with different time resolutions from the meteorological variable data at the original time resolution; or, Performing a convolution operation on the meteorological variable data at the original time resolution using a convolutional neural network to extract meteorological variable data with different time resolutions.
4. The method according to any one of claims 1 to 3, characterized in that The using a tower-shaped local hierarchical capture attention mechanism to encode the target embedding tensor to obtain a prediction feature tensor includes: Determining a hierarchical relationship set for each time node of the target embedding tensor, where the relationship set includes a set of a finite number of time nodes associated with the time node in multiple levels, and the levels are divided according to different time scales; Performing multi-head attention calculation on the hierarchical relationship sets of each time node to obtain an attention result; Performing a preset process on the attention result to obtain the prediction feature tensor.
5. The method according to claim 4, wherein The hierarchical relationship set is the union of a child-level relationship set, a same-level relationship set, and a parent-level relationship set. The child-level relationship set is used to indicate one or more time nodes closest to the time node in the first level, the same-level relationship set is used to indicate one or more time nodes closest to the time node in the second level, and the parent-level relationship set is used to indicate one time node closest to the time node in the third level. The time scales corresponding to the first level, the second level, and the third level are increasing.
6. The method according to any one of claims 1 to 5, characterized in that, The performing feature transformation on the prediction feature tensor to obtain corrected wind speed forecast data includes: Extracting target features from the prediction feature tensor through an aggregation layer; The linear transformation is performed on the extracted target features through a linear layer to obtain the corrected wind speed prediction data.
7. The method according to any one of claims 1 to 6, characterized in that, The method is executed based on a pre-trained wind speed correction model, and the method further includes: Obtaining training data, where the training data includes a plurality of wind speed prediction values and corresponding plurality of wind speed observation values; Calculating a loss function according to the plurality of wind speed prediction values and the corresponding plurality of wind speed observation values, where the loss function includes an activity difference term, and the activity difference term is used to indicate the difference in the fluctuation degree between the plurality of wind speed prediction values and the corresponding plurality of wind speed observation values; According to the loss function, iteratively updating the model parameters through a preset algorithm to train the wind speed correction model.
8. A wind speed correction model, characterized in that, The model includes: An input module, configured to obtain meteorological variable data, where the meteorological variable data is generated by a preset regional weather forecasting model, and the meteorological variable data is used to indicate the temporal variation of a plurality of meteorological variables related to wind speed in a target area within a future period of time; An embedding module, configured to perform information embedding on the meteorological variable data to generate a target embedding tensor; An encoder module, configured to encode the target embedding tensor by using a tower-shaped local hierarchical attention mechanism to obtain a prediction feature tensor, and the tower-shaped local hierarchical attention mechanism is used to indicate that the attention calculation of each time node in the target embedding tensor is constrained to a finite number of time nodes associated with the time node; A post-processing module, configured to perform feature transformation on the prediction feature tensor to obtain the corrected wind speed prediction data.
9. A computing device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
10. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Mesoscale numerical weather forecast real-time correction method and device and storage medium
CN118536395A
Cloud air guide space-time error correction method and system based on deep learning
CN118839303A
Probabilistic wind speed forecasting method and system based on multi-scale information
US20230169230A1