A two-stage real-time wind wave field prediction method and system based on synchronous stationary satellite
Patent Information
- Application Number
- CN202610904234.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2046-06-23
AI Technical Summary
这类非深度学习的传统方法需要反复求解极其复杂的流体力学和大气热力学偏微分方程组
[0027] (1) Breaking through the time delay bottleneck of high-precision reanalysis data, a true "zero-time-difference" wind and wave field initial field reconstruction is achieved. Addressing the problem that existing deep learning extrapolation models overly rely on reanalysis data such as ERA5, which have delays of several hours, as initial fields, leading to severe forecast lags, this invention innovatively proposes a two-stage decoupling architecture. In the first stage, the model directly utilizes minute-level high-frequency synchronous geostationary meteorological satellite observation data. Through the SimVP model's cross-modal inversion network, the delay-free satellite multispectral data information is mapped in real time into a three-dimensional dynamic wind and wave field. This mechanism completely eliminates the "lagging data dependence" at the forecast start time, providing a "zero-time-difference" benchmark initial field that closely approximates the current real atmosphere for subsequent spatiotemporal extrapolation, effectively eliminating the cumulative effect of extrapolation errors.
Smart Images

Figure CN122430945B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of real-time wind and wave field forecasting technology, and in particular to a two-stage real-time wind and wave field forecasting method and system based on geostationary satellites. Background Technology
[0002] Accurate and real-time regional and local wind and wave field forecasts play a vital role in aviation and navigation, wind power generation, atmospheric pollutant dispersion, and disaster prevention and mitigation.
[0003] Currently, the mainstream methods for acquiring and forecasting high-precision wind and wave fields in the industry are mainly divided into two categories: one is the physical data assimilation method based on traditional numerical weather prediction (NWP); the other is the spatiotemporal extrapolation forecasting method that has emerged in recent years, based on deep learning and relying on reanalysis meteorological data such as ERA5. However, when facing the forecasting requirements of "high frequency and real time", both of these existing technologies have significant physical or engineering bottlenecks:
[0004] 1. Traditional non-deep learning assimilation methods consume huge amounts of computational resources and have poor timeliness:
[0005] Traditional wind and wave field generation often relies on data assimilation systems such as 3D / 4D Variational (3D / 4DVar) or ensemble Kalman filtering. These non-deep learning methods require repeatedly solving extremely complex systems of partial differential equations in fluid dynamics and atmospheric thermodynamics. This computational approach not only requires extremely large supercomputer clusters, but the assimilation and convergence process is also extremely time-consuming. The high computational cost and extremely low computational efficiency mean that it is fundamentally unable to meet the minute-level, high-frequency real-time wind and wave field updates required for localized sudden disasters or wind power dispatching.
[0006] 2. Existing deep learning extrapolation methods are limited by ERA5 data latency and cannot achieve true "real-time" forecasting:
[0007] To avoid the enormous computational overhead of traditional assimilation methods, many existing deep learning models have shifted to relying on reanalysis data such as ERA5 published by the European Centre for Medium-Range Weather Forecasts (ECMWF) for wind and wave field extrapolation. While ERA5 data can provide high-precision zonal wind (… ) and meridional wind ( While it contains component information, it is essentially historical data assimilated after the fact, inherently subject to a release delay of several hours. This practice of forcibly using "lagging data" as the "initial field" for deep learning models leads to the following significant drawbacks:
[0008] The extrapolation error accumulates significantly: the forecast model always starts with the atmospheric state several hours ago. Relying solely on deep neural networks to extrapolate forward will cause the prediction error to amplify rapidly and nonlinearly as the time window for compensating for the delay lengthens.
[0009] Slow response to sudden weather events: When faced with sudden and localized strong convection or wind shear and other dramatic changes in the wind and wave field, smooth extrapolation models trained on delayed historical data often react slowly and fail to capture the current real instantaneous dynamic changes.
[0010] In contrast, geostationary meteorological satellites can provide wide-area observations at the minute level, possessing extremely high temporal resolution and near real-time acquisition capabilities, making them an ideal data source to overcome the aforementioned time delay issues. However, geostationary satellites primarily provide multispectral radiation or cloud top brightness temperature observations, and lack the three-dimensional dynamic physical quantities of wind and wave fields (…). , There is a huge gap in the physical dimensions of the components, making it difficult to accurately deduce the underlying or upper-level wind and wave field structure using single-mode satellite data.
[0011] Therefore, how to break free from the computational power dependence of traditional assimilation systems, overcome the one-way dependence of existing deep learning methods on delayed ERA5 data, and effectively utilize the near real-time observation advantages of geostationary satellites to establish robust cross-modal wind and wave field mapping has become a pressing technical challenge in the field of wind and wave field forecasting. Summary of the Invention
[0012] To address the aforementioned problems, this invention provides a two-stage real-time wind and wave field forecasting method and system based on geostationary satellites.
[0013] In a first aspect, the present invention provides a two-stage real-time wind and wave field forecasting method based on geostationary satellites, which adopts the following technical solution:
[0014] A two-stage real-time wind and wave field forecasting method based on geostationary satellites includes:
[0015] Acquire multi-channel observation data from geostationary meteorological satellites;
[0016] Data preprocessing is performed based on the acquired observation data;
[0017] The process involves inverting the spatiotemporal sequence of cross-modal wind and wave fields based on preprocessed data, including spatial multi-scale information encoding based on grouping and normalization, spatiotemporal cross-modal conversion based on the principle of cross-correlation, and spatial decoding and reconstruction of physical wind and wave fields.
[0018] Based on 3DSwin-Transformer, multi-scale wind and wave field evolution forecasting is performed on the inversion results, including 3D spatiotemporal information embedding and local dynamic encapsulation, multi-scale spatiotemporal attention modeling, long-term wind and wave field reconstruction and forecast output.
[0019] Secondly, a two-stage real-time wind and wave field forecasting system based on geostationary satellites includes:
[0020] The data acquisition module is configured to acquire multi-channel observation data from geostationary meteorological satellites.
[0021] The data preprocessing module is configured to perform data preprocessing based on the acquired observation data;
[0022] The inversion module is configured to perform cross-modal wind and wave field spatiotemporal sequence inversion based on preprocessed data, including spatial multi-scale information encoding based on grouping normalization, spatiotemporal cross-modal conversion based on the principle of cross-correlation, and physical wind and wave field spatial decoding and reconstruction.
[0023] The forecasting module is configured to perform multi-scale wind and wave field evolution forecasting based on the inversion results using 3DSwin-Transformer, including three-dimensional spatiotemporal information embedding and local dynamic encapsulation, multi-scale spatiotemporal attention modeling, long-term wind and wave field reconstruction and forecast output.
[0024] Thirdly, the present invention provides a computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the two-stage real-time wind and wave field forecasting method based on geostationary satellites.
[0025] Fourthly, the present invention provides a terminal device, including a processor and a computer-readable storage medium, wherein the processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a two-stage real-time wind and wave field forecasting method based on geostationary satellites.
[0026] In summary, the present invention has the following beneficial technical effects:
[0027] (1) Breaking through the time delay bottleneck of high-precision reanalysis data, a true "zero-time-difference" wind and wave field initial field reconstruction is achieved. Addressing the problem that existing deep learning extrapolation models overly rely on reanalysis data such as ERA5, which have delays of several hours, as initial fields, leading to severe forecast lags, this invention innovatively proposes a two-stage decoupling architecture. In the first stage, the model directly utilizes minute-level high-frequency synchronous geostationary meteorological satellite observation data. Through the SimVP model's cross-modal inversion network, the delay-free satellite multispectral data information is mapped in real time into a three-dimensional dynamic wind and wave field. This mechanism completely eliminates the "lagging data dependence" at the forecast start time, providing a "zero-time-difference" benchmark initial field that closely approximates the current real atmosphere for subsequent spatiotemporal extrapolation, effectively eliminating the cumulative effect of extrapolation errors.
[0028] (2) An explicit three-dimensional spatiotemporal multi-scale modeling mechanism was established, which significantly improved the accuracy of long-term forecasts of complex wind and wave field evolution. Addressing the difficulty of decoupling atmospheric fluid advection in traditional two-dimensional time-series models (such as the ConvLSTM model), the second stage of this invention deeply integrates the 3DSwin-Transformer model architecture. Through three-dimensional block (3DPatchEmbedding) and shifted three-dimensional window multi-head self-attention (3DSW-MSA) mechanisms, the model can not only capture the overall migration trajectory of large-scale weather systems (such as cyclones and frontal zones), but also accurately track the formation and dissipation processes of local strong convective cells. The introduction of three-dimensional relative position offset enables it to physically possess the ability to accurately perceive atmospheric fluid displacement, thereby effectively suppressing the distortion of wind direction topology and the excessive smoothing of wind speed texture in extrapolation forecasts lasting up to 5 days.
[0029] (3) An innovative joint loss optimization strategy that balances global stability and extreme value sensitivity is proposed, significantly enhancing the response capability to extreme wind and wave fields. Addressing the engineering pain points of traditional extrapolation models, which are prone to "slow prediction" and "low magnitude" when facing sudden strong winds and wind shear, this invention designs a loss function jointly weighted by MAE and RMSE during the training phase of the forecast model. The MAE loss constrains the stability and coherence of the large-scale background circulation, while the square amplification characteristic within the RMSE loss imposes a severe exponential penalty on the prediction error of sudden changes in the wind and wave field or extreme high-wind-speed areas. This joint optimization mechanism forces the forecast model to break away from its "averaging" tendency, enabling it to possess extremely high sensitivity and magnitude recovery capability when facing extreme weather such as sudden strong winds in localized areas.
[0030] (4) It abandons the huge computing power overhead of traditional physical assimilation systems and meets the extremely high timeliness requirements of high-frequency operational deployment. Traditional numerical meteorological assimilation methods such as 3D / 4DVar or EnKF require repeated solving of huge partial differential equation systems, which consumes a lot of time. The two-stage deep learning model (SimVP+3DSwin-Transformer) of this invention exhibits extremely high forward computation efficiency in its online inference stage after completing offline training. It only needs to input real-time received satellite data to output the continuous wind and wave field evolution for the next few days in seconds. This extremely low computing power overhead and extremely high inference speed are extremely suitable for high-frequency operational deployment in industrial and disaster prevention and mitigation scenarios with extremely high real-time requirements, such as wind power generation scheduling and monitoring of sudden atmospheric pollutant diffusion. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of a two-stage real-time wind and wave field forecasting method based on geostationary satellites according to Embodiment 1 of the present invention;
[0032] Figure 2This is a schematic diagram of the first stage of Embodiment 1 of the present invention;
[0033] Figure 3 This is a schematic diagram of the second stage of Embodiment 1 of the present invention;
[0034] Figure 4 This is a comparative experimental diagram of MAE after 72 hours in Embodiment 1 of the present invention;
[0035] Figure 5 This is a comparative test chart of 24-hour extreme weather prediction errors in Embodiment 1 of the present invention;
[0036] Figure 6 This is a comparison chart of the 72-hour end-to-end prediction time in Embodiment 1 of the present invention.
[0037] Figure 7 This is a visual forecast result presented in the form of an image, as described in Embodiment 1 of the present invention. Detailed Implementation
[0038] The present invention will be further described in detail below with reference to the accompanying drawings.
[0039] Example 1
[0040] Reference Figure 1 This embodiment presents a two-stage real-time wind and wave field forecasting method based on geostationary satellites.
[0041] Specifically, the following steps are included:
[0042] S1 Data Preprocessing Stage
[0043] First, for the input multi-channel observation data from geostationary meteorological satellites, meteorologically sensitive channels are screened and target areas are selected. Specifically, considering the differences in the representation of meteorological physical quantities by different spectral channels, this invention selects specific channels that are highly sensitive to atmospheric motion and cloud water vapor changes as the initial input. Let the original wide-swath satellite data be... ,in and These represent spatial longitude and latitude, respectively. To eliminate redundant information and focus on the core prediction task, the coordinates are set according to preset longitude and latitude boundaries. Grid point selection is performed to obtain the observation tensor within the target region. The number of channels , and Indicates the size of the selected area.
[0044] S1.1 Sensitive Channel Filtering and Region Clipping
[0045] Satellite channels 7 and 13, which are sensitive to atmospheric motion, were selected. The original wide-swath satellite data is assumed to be... The observation tensor is obtained after cropping the target region. The number of channels , and Indicates the size of the selected area.
[0046] S1.2 Spatial Resolution Alignment
[0047] Next, the cropped satellite observation tensor Spatial resolution is unified with the ground truth ERA5 wind and wave field data. Since satellite observations typically have extremely high spatial resolution, while the spatial grid of ERA5 data is relatively coarse, this invention uses a downsampling method to unify the spatial resolution of the satellite data to the target grid system of ERA5 in order to effectively reduce the data size and ensure strict spatial correspondence between data pixels. Let the satellite image after unified projection onto the target grid be denoted as . For the center of each target pixel The calculation is performed using bilinear interpolation:
[0048] ,
[0049] in, These are the four nearest data points in the original satellite data. For distance-based interpolation weights, satisfying Aligned satellite data is obtained after downsampling. .
[0050] S1.3 Time Base Alignment
[0051] Finally, based on spatial grid alignment, a rigorous time reference determination was further implemented for the satellite data sequences and the ERA5 wind and wave field data. Let the time series set of the ERA5 wind and wave field true value sequences be denoted as... The corresponding two-dimensional physical wind and wave field matrix is denoted as (Including zonal wind component) Meridional wind component For any given moment... In a minute-level high-frequency satellite observation time series, the synchronous observation frame with the smallest time deviation is selected as the input information for that moment. Then, according to the preset time resolution of 6 hours, the satellite data at the corresponding time points are... With the true value of the storm field Composition of spatiotemporal paired samples This step effectively eliminates temporal sampling misalignments between multi-source heterogeneous data, establishing reliable training materials with consistent physical benchmarks for subsequent models.
[0052] S1.4 Data Dimensionlessness and Standardization Based on Z-score Algorithm
[0053] After spatiotemporal pairing, in order to eliminate the difference between satellite multispectral radiance and the underlying two-dimensional wind and wave field ( , The invention further performs numerical reconstruction and dimensionless processing based on Z-score (standard deviation standardization) on the paired sample set to accelerate the gradient convergence process of the subsequent deep neural network, which is due to the huge difference in dimensions and orders of magnitude between the components.
[0054] Specifically, within the constructed training set sample space, for the aligned satellite data tensor Each physical channel (For example, channels 7 and 13), calculate the global average of their historical time series data respectively. and standard deviation Subsequently, the Z-score operator is used to perform point-by-point normalization mapping on the observed pixels of each channel:
[0055] ,
[0056] in For initial data, and These are the mean and standard deviation, respectively. For the normalized data, similarly, for the ERA5 true wind and wave field data... Also targeting zonal winds and the wind Calculate the global mean for the two fluid dynamic components. with standard deviation (in Then perform the same normalization operation to obtain the standardized wind and wave field truth tensor. .
[0057] After this Z-score standardization step, the final cross-modal spatiotemporal dataset input to the model The information distribution of each physical dimension is strictly constrained to a standard normal distribution space with a mean of 0 and a variance of 1. This operation not only effectively suppresses the interference of local extreme anomalies in meteorological observations (such as sudden changes in brightness and temperature caused by strong convection) on the overall information, but also effectively avoids gradient diffusion or explosion problems caused by excessive numerical spans in cross-modal inversion and spatiotemporal extrapolation, laying a solid numerical computational foundation for establishing robust nonlinear mappings.
[0058] In this embodiment, the target study area for cross-modal inversion and subsequent spatiotemporal forecasting is strictly defined as the Northwest Pacific and surrounding East Asian region (latitude and longitude range: North Latitude). N, East longitude E). Since this geographical coordinate range constitutes the multi-scale physical constraints for the neural network design of this invention, its meteorological environment has the following typical characteristics and extreme wind field attributes:
[0059] Typical characteristics of the geographical and climatic system: This region spans the equator, the tropics, the subtropics, and even the mid-to-high latitude temperate zones, encompassing vast open ocean areas and extremely complex land-sea boundary zones. This enormous span and the differences in thermal properties between land and sea make it one of the regions with the most intense air-sea interactions and the most complex meteorological and hydrological conditions globally.
[0060] Characteristics of multi-scale wind and wave field topological evolution: Due to its unique geographical location, the wind field dynamics evolution in this region exhibits extremely high spatiotemporal variability and multi-scale coupling characteristics.
[0061] Macro-background wind field: It is constantly controlled by the East Asian monsoon (northwest wind in winter and southeast wind in summer), the subtropical high-pressure ridge and the upper-level jet stream of the mid-latitude westerly wind belt, and the basic wind field exhibits a large-scale directional advection characteristic.
[0062] Extreme sudden wind fields: This region is the sea area where tropical cyclones (typhoons / super typhoons) are most frequently generated and active globally. At the same time, squall lines, local wind shears, and cold waves caused by strong convective cloud clusters moving southward in winter cause extreme nonlinear fluid evolution, such as sudden increases in local wind speeds and abrupt changes in wind direction, to frequently occur in this sea area.
[0063] Wind field forecasting within the aforementioned latitude and longitude grid is a typical scenario for validating model performance. Traditional numerical weather prediction (NWP) data, or ERA5 data with delays of several hours, often react slowly to large wind speed jumps in the core area of sudden strong typhoons in this region. Therefore, the inversion forecasting system designed in this invention can provide zero-time-difference initial field support for real-time wind and wave field forecasting in this region.
[0064] S2 Phase 1: Cross-modal wind and wave field spatiotemporal sequence inversion based on SimVP architecture
[0065] This phase aims to completely eliminate the inherent time delay defect in high-precision reanalysis data such as ERA5. Since atmospheric wind fields are essentially continuous advection motions, static satellite data cannot characterize atmospheric motion vectors. Therefore, this invention innovatively constructs a cross-modal inversion model based on the deep fusion of a multi-temporal-series cloud-motionwind module (CMW) and the SimVP model architecture. The model takes the multi-channel satellite data sequence (continuous observation data sampled at a 6-hour time resolution) acquired in the S1 phase as input, and uses the spatial displacement of the continuous cloud image to directly deduce the current time-indelayed two-dimensional physical wind and wave field, thereby realizing the physical reconstruction from satellite spatiotemporal sequence to real-time wind and wave field data.
[0066] From an atmospheric dynamics perspective, the actual wind and wave field is essentially an open, highly nonlinear energy dissipation and exchange system. Within the forecast window, the target region will inevitably be continuously subjected to advection transport from beyond its spatial boundaries and the cross-boundary inflow of external energy (such as momentum flux and heat flux). If the spatial dimension of the input data is strictly limited to the target region, the forecast model will face a severe 'boundary information truncation' defect, unable to detect and calculate in advance the momentum carried by large-scale weather systems (such as cyclones and cold fronts) that are about to move into the target region. Therefore, in the model input construction at this stage, this invention deliberately sets the spatial coverage of the input information to be significantly larger than the final forecast target region, thereby constructing a 'spatial buffer' of a certain width around the target region. This expanded spatial input strategy provides the three-dimensional spatiotemporal forecast network with a wide-area receptive field containing the complete upstream and downstream energy cascade laws, enabling it to explicitly capture the trend of external energy inflow across the target region boundary, effectively eliminating the boundary distortion and error convergence problems common in extrapolation forecasts. After interpolation, the input data and the target data have the same height and width.
[0067] Let the input sequence tensor be... (Where the time step T=3 and the number of channels C=2 correspond to the water vapor and infrared sensitive channels of the satellite data, respectively). Structurally, this inversion model is decoupled into three cascaded subnetworks: a space information encoder, a spatiotemporal cross-modal converter, and a space information decoder. Specific implementation details are as follows:
[0068] S2.1 Spatial multi-scale information encoding based on group normalization (SpatialEncoder):
[0069] To learn the spatial physical properties of 18 consecutive hours of satellite data within a unified information space, the space encoder employs a two-dimensional convolutional network with shared weights to process the input sequence. Each frame in the process is individually learned for local spatial information.
[0070] The encoder is made by It consists of cascaded convolutional modules (EncoderBlocks). For any time step frame in the sequence... In the lth layer ( Its information mapping and propagation mechanism is defined as follows:
[0071] ,
[0072] ,
[0073] in, Pre-activation features of the l-th layer at time step k; and For learnable convolutional kernels and biases, where, and These represent the feature representations of the (l-1)th and lth layers of the encoder, respectively; k is the temporal sample index, and l is the network layer index; * represents the convolution operation, learning the cloud topology. For group normalization operations; It is a non-linear activation function.
[0074] Key implementation details: In normalization, this invention abandons the conventional layer normalization (LayerNorm) and innovatively adopts group normalization (GroupNorm), strictly setting the number of groups (num_groups) to 2. This setting effectively controls model overfitting while significantly reducing the computational cost of large-scale meteorological tensor operations, and maintains the independence of information representation from different physical channels (water vapor and infrared). Activation function Choose LeakyReLU.
[0075] After four layers of spatial dimensionality reduction and information abstraction, the information output of the 18-hour data is jointly represented along the time dimension to generate a spatiotemporal dynamic representation of the latent layer. Where T is the time step; D is the number of channels; and H and W are the sizes of the selected regions.
[0076] S2.2 Spatiotemporal cross-modal transformation based on the principle of cross-correlation (CMWBlock):
[0077] This module is a core component connecting the satellite's multi-temporal radiation modes and single-frame hydrodynamic modes. Its physical essence is the implicit implementation of cross-correlation template matching in cloud-guided wind (CMW) technology within a deep network. This module consists of... It consists of stacked Inception module variant units with residual shortcuts.
[0078] Before the data enters the CMW module, the time dimension T is first directly superimposed onto the axis containing the channel dimension C. This operation of stacking time onto the channel dimension allows subsequent single 2D convolutions to synchronously aggregate information across continuous time steps, enabling the network to directly infer the cloud's movement vector. To overcome the physical isolation of discrete time steps, without sacrificing spatial resolution, the time dimension is forcibly flattened and stitched along the channel dimension axis.
[0079] ,
[0080] ,
[0081] in, This is the input spatiotemporal information sequence tensor. It is the original latent temporal information output by the preceding spatial multiscale information encoder after learning local spatial information from 18 consecutive hours of geostationary satellite brightness temperature images. : Indicates channel splicing operation. It is a discrete-time step single-frame latent spatial information tensor. It corresponds to the input sequence respectively. In the past, the first Time, Number Time and the present Independent information slices of satellite data at any given time, after spatial encoding. For the data reshaping operation, B and T represent the batch size and time step, respectively. The size of the selected area.
[0082] This is the spatiotemporal composite information tensor after collapse and superposition. It is the final product of this algorithm operation. By forcibly integrating the time dimension into the channel dimension, the dynamic motion information of multiple time series is "laid out" in the spatial channel, serving as the standardized input for subsequent implicit template matching in CMWBlock (cloud wind guide block).
[0083] Within each Inception module unit, structural learning and optimization are performed based on the physical characteristics of wind field evolution, as represented by:
[0084] ,
[0085] in : Represents the dynamic spatiotemporal information representation output by the nth layer Inception unit. The subscript "dynamic" here is very important; it physically represents the cloud movement vector or dynamic evolution information deduced by the network through analysis of continuous time steps. : Indicates the channel concatenation operation. It merges the information learned by the three internal parallel multi-scale convolutional branches along the channel dimension, thereby fusing spatial information at different scales. : Represents the input to the current nth layer's latent spatiotemporal dynamics representation, i.e., the output of the previous layer (n-1th layer) of the network. : indicates use Two-dimensional convolution operations performed by convolution kernels. Due to their smallest receptive field, they are physically used to accurately frame and track local distortions of small-scale fluids, such as fine-grained eddies, in wind and wave fields. : indicates use The convolution kernel performs a two-dimensional convolution operation. Its receptive field is moderate, and physically it corresponds to capturing the motion trajectory and evolution process of intermediate convective cells. : indicates use The convolution kernel performs a two-dimensional convolution operation. As the largest receptive field in this module, it is physically responsible for estimating large-scale circulation patterns over a wide area. Furthermore, because the network is highly focused on capturing dynamic movement patterns, it is prone to losing "static background information" (such as the underlying stable framework structure of clouds), which is equally crucial for wind field reconstruction. To address this issue, this invention introduces a residual shortcuts mechanism in CMWBlock:
[0086]
[0087] in, This represents the dynamic spatiotemporal information output by the nth layer Inception unit. This represents the skip connection information mapping and dimension alignment function. Through forward and backward residual skip connections, deep networks can retrospectively re-introduce earlier static data information during dynamic motion estimation. This dual perspective of "dynamic and static" processing not only stabilizes deep gradient propagation but also endows the model with strong robustness to drastic wind field evolution.
[0088] S2.3 Spatial Decoding and Reconstruction of Physical Wind and Wave Fields (SpatialDecoder):
[0089] After obtaining the latent representation of a single frame after spatiotemporal evolution fusion, the decoder needs to upsample it and reconstruct it into a two-dimensional wind field on the target grid. The decoder is also composed of symmetrically stacked 4 layers of deconvolution (unConv2d) modules.
[0090] ,
[0091] in, This represents the transpose convolution operation. : indicates the first The learnable transposed convolution kernel weight tensor (Weight) in the layer transposed convolution module. : indicates the first Learnable bias vectors in the deconvolution module. : indicates input to the current number The pre-order low-resolution semantic information tensor of the layer-space decoder. Specifically, when... hour, That is, the input is a single-frame latent representation after being fused with macroscopic dynamics by a spatiotemporal cross-modal converter (CMWBlock). : Represents the grouping normalization operator, used to stabilize the numerical distribution of multi-channel brightness temperature and wind momentum in the recovered spatial dimension, and independently maintain the generalization ability of heterogeneous modes.
[0092] Finally, after dimensionality reduction using a linear projection layer, the initial two-dimensional physical wind and wave field is output, which precisely matches the target mesh. (The first channel is the zonal wind component) The second channel is the meridional wind component. ).
[0093] First stage of model parameter optimization:
[0094] In this stage of offline training, mean squared error (MSE) is used as the loss function for spatiotemporal reconstruction to measure the wind field at the current moment inferred by the model. ERA5 high-precision analysis field with strict time alignment Momentum deviation between:
[0095] ,
[0096] : Represents the global reconstruction loss of the first-stage cross-modal inversion network. Wherein, This represents the set of all learnable parameters in the entire deep neural network in the first stage (including the multi-scale information encoder, the spatiotemporal cross-modal converter, and the spatial information decoder), i.e., the weights of the multi-scale convolutional kernels mentioned earlier. With bias ). : Represents the batch size during a single model iteration training. This parameter is introduced to smooth out gradient jitter caused by extrema by calculating the average error of samples within a batch, ensuring stable Adam optimization convergence in complex high-dimensional spaces. : The number of physical channels representing physical forecast target information. Specifically, the first channel ( Strictly corresponds to the zonal wind component in atmospheric dynamics Second channel ( Strictly correspond to the meridional wind component . and : Height and width of the spatial grid for the target forecast area. : Indicates the output of the linear projection layer after network inference. One sample in Predicting the initial field of the two-dimensional physical wind and wave field at a given time. : Represents the Z-score normalized ERA5 high-precision wind and wave field ground truth tensor that is strictly aligned with satellite observation time during the data preprocessing stage (S1.3). : Represents the squared L2 norm expanded in the spatial and channel dimensions.
[0097] By continuously minimizing the loss function, the backpropagation and optimization of the weights of the multi-scale convolution kernel are completed, effectively realizing the "zero-time-difference" inversion of the initial field of the wind and wave field.
[0098] S3 Phase 2: Multi-scale wind and wave field evolution forecasting based on 3DSwin-Transformer
[0099] This phase aims to build a spatiotemporal prediction model with long-term memory and multi-scale dynamic evolution capture capabilities based on the high-precision real-time initial field obtained in the first phase. The core of this invention lies in using a 3D Shifted Window Attention mechanism to simulate the advection transport patterns and dynamic evolution of atmospheric circulation.
[0100] S3.1 Forecast System Input and Output Data Definitions
[0101] Before performing model simulations, the tensor dimensions of the spatiotemporal data must be strictly defined. Based on the timeliness requirements of meteorological forecasting services, this embodiment sets the time resolution to 6 hours.
[0102] Input data definition: Forecasting model The input is a continuous wind and wave field sequence over the past 5 days (120 hours). Let the input time step be... The input tensor is then represented as Where 2 represents zonal wind. and the wind Dual-channel and The spatial grid height and width for the target region are defined here. To ensure spatial consistency of dynamic characteristics throughout the entire 'inversion-prediction' process, this... and The value of is strictly consistent with the initial field specification reconstructed in the first stage (S2).
[0103] Output data definition: The model's output target is the wind and wave field evolution sequence for the next 3 days (72 hours). Let the output time step be... The predicted output tensor is then represented as .
[0104] S3.2 3D Spatiotemporal Information Embedding (3DPatchEmbedding) and Local Dynamics Encapsulation
[0105] Atmospheric wind and wave fields represent the evolution of continuous fluids in four-dimensional spacetime. Simple two-dimensional spatial convolution or one-dimensional temporal models (such as ConvLSTM) often sever the inherent temporal and spatial coupling within "advection." This step employs a three-dimensional convolution kernel to process the input tensor... Along the time dimension Space height and width Perform synchronous block partitioning (3DPatchPartition).
[0106] Spatio-temporal partitioning:
[0107] Let the block size of the 3D convolution kernel be... In order to capture short-term wind field evolution trends in the shallow layer of the network, this embodiment preferably uses... This means that, in the time dimension, the wind field evolution every 2 frames (i.e., a 12-hour physical time window) is encapsulated as a whole; in the spatial dimension, each The physical lattice points are divided into a local surface.
[0108] At this time, originally The continuous input is segmented into non-overlapping spatiotemporal data blocks.
[0109] Linear mapping and high-dimensional embedding:
[0110] Each segmented 3D data block is flattened and linearly projected into a one-dimensional token vector through a learnable 3D convolutional layer (3DConvolution), with its mapping dimension set to 1. .
[0111] ,
[0112] ,
[0113] in, : Represents the input continuous wind and wave field spatiotemporal tensor (containing the past 20 time steps, i.e., 5 days). , (Dual-channel evolution sequence). : Represents a three-dimensional convolutional layer. Here, it is not used as an operator for learning deep features, but rather as a linear projector that directly maps the momentum distribution of each local microscopic air mass to a high-dimensional latent space. and : These represent the receptive field size and stride size of the 3D convolution kernel, respectively. To ensure the independence of information from different spatiotemporal clusters and to ensure no mesh is missed, this embodiment employs overlapping segmentation, i.e., strictly setting... . : Represents the high-dimensional raw information tensor that has not yet been flattened after 3D convolution projection. : Represents layer normalization operation. Used to eliminate distribution offset caused by differences in the absolute baseline values of wind speed at different times or in different sea areas, ensuring that the information values input to the attention calculation module have good numerical stability. : Represents the high-dimensional mapping dimension (EmbeddingDimension). In this model, it is set to That is, each air mass with a physical size of 2×3×3 is ultimately condensed and characterized as a comprehensive dynamic information vector containing 96 channels. : This represents the initial token sequence tensor that is finally input into the 3DSwin-Transformer module after flattening and normalization. : Represents the total number of three-dimensional spatiotemporal blocks. As the formula shows, the originally complex continuous fluid evolution is discretized into... Each unit is an independent topological unit. This greatly preserves the local wind field vortex structure while exponentially reducing the time and space complexity of subsequent multi-head self-attention calculations.
[0114] This step discretizes the continuous wind and wave field over five days into high-dimensional potential topological information across ten time sections. This effectively encapsulates the microscopic three-dimensional flow field structure of the local temporal air mass (AirMass) at the network's front end, providing a standardized dynamic representation unit for the subsequent attention mechanism.
[0115] S3.3 Multi-scale Spatiotemporal Attention Modeling (3DSwin-TransformerBlocks)
[0116] To accurately simulate the nonlinear migration and evolution patterns of large-scale weather systems (such as cyclones, typhoons, and fronts) across grids and time periods (e.g., inputting the past 5 days and forecasting the next 3 days), this invention learns the initial spatiotemporal information sequence. Passing through in sequence by The 3DSwin-Transformer module is composed of stacked stages. Within each stage, in order to capture global fluid dependencies while significantly reducing computational complexity, the model alternately stacks pairs of network blocks containing regular 3D window multi-head self-attention (3DW-MSA) and shifted 3D window multi-head self-attention (3DSW-MSA).
[0117] (1) 3D Window Self-Attention Mechanism (3DW-MSA) - Local Weather System Evolution Capture and Complexity Reduction
[0118] Traditional Global Multi-Head Self-Attention (GlobalMSA) requires calculating the relationships between all spatiotemporal tokens, which Where T is the time step; H and W are the height and width of the spatial grid of the target prediction area. Where is the window size; D is the number of channels. Considering the enormous number of high-resolution meteorological grid points, such quadratic computational overhead is unacceptable in engineering. Therefore, this invention uses the 3DW-MSA mechanism to divide the massive spatiotemporal data into multiple non-overlapping three-dimensional windows.
[0119] Set the window size to (For example This means that attention computation is strictly limited to a single time step covering 5 time steps (equivalent to tens of hours of evolution in the latent layer) and Within the "Local Spatio-temporal Tube" of the spatial region, the computational complexity of a single layer of self-attention is significantly reduced to [value missing]. Where T is the time step; H and W are the height and width of the spatial grid of the target prediction area. is the window size; D is the number of channels; they grow linearly, thus enabling high-resolution, long-term wind and wave field calculations.
[0120] Within each local 3D window, the self-attention calculation formula is defined as:
[0121] ,
[0122] in : These represent the query matrix, key matrix, and value matrix generated within the current local 3D window by mapping the input spatiotemporal token sequence through a linear projection matrix. These represent the patch size (number of patches) of the local window on the time axis, spatial height axis, and spatial width axis, respectively. That is, the total number of spatiotemporal tokens contained within a single local window; This represents the mapping channel dimension under a single-head attention mechanism. : Represents the dot product of the query matrix and the transpose of the key matrix. Its core physical meaning lies in calculating the correlation and momentum similarity score between any two spatiotemporal air masses within the current local window. : Represents the channel dimension of the key vector. As a scaling factor, it is used to scale the dot product values in a high-dimensional feature space, avoiding excessively large dot product values that could cause the Softmax function to enter the gradient saturation region, thereby effectively preventing gradient vanishing during backpropagation. : Represents the three-dimensional relative position bias matrix directly superimposed on the self-attention scoring matrix. This matrix forces the model to perceive the spatial distance decay and temporal causality of fluid motion by explicitly introducing the prior physical constraints of relative position into the matrix operation. : Represents the normalized activation function, used to transform the composite scoring matrix into a nonlinear probability weight distribution that sums to 1 row by row. : Represents the spatiotemporal nonlinear coupling tensor of the final output of the current local window, which is expressed by the probability weights in the value matrix. By performing weighted aggregation, a digital simulation of the energy and momentum within a local weather system was achieved.
[0123] Then, by maintaining a learnable 3D parameter bias table To parameterize Within any local 3D window, since the number of blocks for each axis are respectively , and Then the relative coordinate difference between any two tokens in each dimension is strictly constrained to the following range: Time lag range: ,total Possible relative time differences. Spatial altitude distance range: ,total Local vertical / meridian relative displacement. Spatial width distance range: ,total Local horizontal / zonal relative displacement.
[0124] Therefore, in order to exhaustively enumerate all possible combinations of microscopic fluid relative displacements within the entire window, this invention constructs a learnable storage table composed of dimensionless scalars, denoted as By explicitly encoding the temporal lag and spatial relative distance between any two air masses within a local window, from the bias table The corresponding bias value is indexed and added to the attention score. This mechanism enables the model to have extremely accurate fluid distance decay perception and temporal causality of momentum transfer (i.e., clearly defined weights for the effects of time sequence).
[0125] (2) Shifted 3D Window (3DSW-MSA) – Cross-boundary simulation of large-scale advection effects
[0126] Isolated 3D windows lead to an absolute disconnect between information from different local air masses. For example, a typhoon system that crosses the window boundary would be artificially cut off, which clearly violates the fluid continuity equation and makes it impossible to simulate the cross-regional movement of large-scale systems such as high-pressure ridges. Therefore, this invention employs the 3DSW-MSA mechanism in the next consecutive Transformer layer.
[0127] 3D cyclic shift:
[0128] Before the data enters the attention calculation, the window is divided into time, height, and width directions respectively. A cyclical rolling operation of patch steps (Roll operation).
[0129] This shift forcibly breaks the fixed window boundary of the previous layer. In meteorological fluid dynamics, this is equivalent to a microscopic digital simulation of "Lagrange advection." The airflow characteristics at the left edge of the window in the previous moment are seamlessly integrated into the computational receptive field of the right window in the next moment through the shift mechanism. This allows deep networks to cross the boundaries of spatial grids and time steps, continuously tracking the movement trajectory of weather systems within a 5-day input cycle, thus providing accurate dynamic inertia for the forecast of the next 3 days.
[0130] The introduction of 3D attention masking mechanism:
[0131] While cyclic shifting solves the advection boundary problem, its "end-to-end" cyclical nature can piece together non-adjacent air masses that are physically and temporally far apart into a single new window. Unrestricted calculation of self-attention can lead to instantaneous non-physical coupling in the ocean wind field, disrupting the localized laws governing fluid flow.
[0132] To ensure the logical correctness of the shift calculation, this invention constructs an extremely fine three-dimensional mask matrix within the network. In the new window after shifting and reassembling, the tokens are divided into multiple unrelated sub-windows based on their origin from different regions of the original data. Before performing the Softmax operation, a mask matrix is overlaid on the attention score.
[0133] ,
[0134] in The query matrix, key matrix, and value matrix are generated respectively by mapping the input spatiotemporal data sequence through a linear projection matrix; b represents the channel dimension of the key vector, used to scale the dot product result to stabilize the training process; b is the three-dimensional relative position bias matrix, used to model the spatial or temporal relative relationships between tokens. This represents the attention mask matrix, used to constrain the scope of attention calculation. For data belonging to the same original adjacent region, its corresponding... The value is 0; for non-adjacent token data that are forcibly merged due to cyclic shifting, its corresponding value is 0. Value (or negative infinity). Due to The attention weights of these non-physically adjacent regions are perfectly blocked after Softmax. This design ensures the absolute rigor of atmospheric hydrodynamic topological constraints while significantly preserving the hardware efficiency of parallel computing.
[0135] (3) Stable transmission and nonlinear coupling of spatiotemporal information
[0136] Within each pair of 3DW-MSA and 3DSW-MSA blocks, the flow formula for latent layer information is defined as:
[0137] ,
[0138] ,
[0139] in, and They represent the first The intermediate information between layers l is represented by LayerNorm, which represents the layer normalization operation, and MLP represents the feedforward multilayer perceptron structure. Before each self-attention computation and multilayer perceptron (MLP) mapping, the information undergoes layer normalization (LayerNorm) to eliminate the information distribution shift (covariate shift) caused by differences in the absolute value of wind speed under different geographical latitudes or seasonal backgrounds. Simultaneously, cross-layer residual connections effectively alleviate the gradient vanishing phenomenon that may occur in deep networks with dozens of stacked layers. The MLP layer utilizes the GELU activation function to compress the data again after expanding the channel dimension, thereby enhancing cross-variable (such as zonal wind) performance. With the wind The nonlinear dynamic coupling between them (due to the interaction caused by the Coriolis force).
[0140] S3.4 Long-duration wind and wave field reconstruction and forecast output (3DPatchExpanding & Projection)
[0141] After learning through multiple layers of 3DSwin-Transformer, the abstract information sequence reaches the top layer of the network. It contains highly condensed future evolution patterns, and its time dimension corresponds to a highly compressed latent time step, which needs to be restored to the physical grid form of the next 3 days (12 frames).
[0142] Spatiotemporal block expansion (3DPatchExpanding):
[0143] This step employs the inverse structure of S3.1 downsampling. It uses a linear mapping to reduce the channel dimension from... Compress to This is combined with tensor reshaping operations to fold the one-dimensional token sequence back into a four-dimensional spatiotemporal coordinate system. To avoid the "chessboard effect" (i.e., abrupt changes in wind speed distribution that do not conform to fluid dynamics) caused by direct upsampling, this step uses three-dimensional sub-pixel convolution in the spatial dimension to smoothly amplify the probabilistic receptive field and resolution step by step. The channel mapping expansion formula is as follows:
[0144] ,
[0145] ,
[0146] in, These represent the height compression dimensions of the subsurface information in the time, height, and width dimensions, respectively, where D represents the number of channels. This is represented as a data reorganization operation. For learnable weight matrix, This is a learnable bias. Then, a 3D subpixel spatial rearrangement is performed using a periodic shuffling operator to interweave and rearrange the information in the dilation channel dimension onto the spatial height and width dimensions according to deterministic geometric rules:
[0147] ,
[0148] in, The underlying element index mapping relationship is defined as follows:
[0149] ,
[0150] in The rearranged high-resolution output tensor. Low-resolution, high-channel-count dilation tensor before rearrangement. This represents the target channel index on the high-resolution output data. These represent the batch size and time step, respectively. : Indicates the spatial upsampling rate. After three-dimensional subpixel rearrangement, the channel is equivalently converted into spatial data. This represents the spatial index of high-resolution coordinates in a low-resolution feature map, while This represents the rearrangement index of a subpixel in the channel dimension, used to map channel information back to the corresponding spatial location, thereby achieving subpixel-level spatial reconstruction.
[0151] Precise Temporal Projection:
[0152] Restoring to the target spatial resolution Subsequently, because the time step of the network's top-level output may differ from the target's... Not entirely consistent, this invention adds a three-dimensional anisotropic convolutional layer and a linear regression head at the end.
[0153] The regression head precisely maps the reconstructed latent information to physical channels over the next 3 days (12 time steps):
[0154] ,
[0155] in, This represents the final latent space feature representation of the encoder or spatiotemporal modeling module; (⋅) represents a three-dimensional subpixel rearrangement or patch expansion operation, used to map low-resolution features to a high-resolution space; and These represent the learnable convolutional kernel parameters and bias terms of the output layer, respectively, which are used to map the expanded features to the target physical variable space; Represents a three-dimensional convolution operation. This represents the final physical field prediction result output by the model. The final output tensor This precisely corresponds to the two-dimensional continuous fluid field distribution occurring every 6 hours over the next 72 hours. (Including) (V component and V component). Thus, by using a rigorous mathematical mapping of a 3-day sequence output from a 5-day sequence input to a 3D spatiotemporal attention advection simulation, a high-precision nonlinear extrapolation forecast of multi-scale wind and wave fields from historical evolution to the future is achieved.
[0156] Experimental verification
[0157] To systematically verify the performance advantages of the method of this invention in wide-area sea surface wind and wave field prediction, we constructed a multi-source spatiotemporally aligned dataset based on real meteorological observations. The dataset contains hourly satellite observations and hourly wind and wave field data from multiple years of continuous data, and is strictly divided into training, validation, and test sets according to the timeline to ensure that the test scenarios cover a variety of typical and extremely complex fluid evolution processes, such as tropical cyclones (typhoons), cold waves, and frontal systems.
[0158] To comprehensively evaluate the performance of the method of this invention, the following three mainstream comparative methods, which are representative of both academic and industrial meteorological forecasting, are set up:
[0159] Numerical weather prediction (NWP): The traditional numerical weather prediction method relies on supercomputers to solve extremely complex fluid dynamics and partial differential equations and assimilate massive amounts of multi-source observation data. It represents the highest accuracy benchmark in the current meteorological industry based on physical mechanisms.
[0160] ConvLSTM model: A classic deep learning spatiotemporal sequence prediction model. It uses a two-dimensional convolutional long short-term memory network to directly perform smooth extrapolation from historical wind field sequences, representing the traditional two-dimensional time series modeling approach that ignores the three-dimensional spatial advection effect and cross-modal initial field reconstruction.
[0161] UNet model: A basic deep visual mapping network. It directly translates single or short-sequence satellite data end-to-end into wind fields, representing a direct static mapping method that lacks long-term memory mechanisms and spatiotemporal dynamic evolution modeling.
[0162] All deep learning methods were evaluated under the same training and test sets. The evaluation metrics focused on the three core challenges of the model in weather forecasting: mean absolute error (MAE, m / s, assessing the stability of the global background wind field and overall error), root mean square error (RMSE, m / s, which, due to its square amplification characteristics, focuses on assessing the model's sensitivity to extreme high wind speeds and sudden weather events, as well as its positioning accuracy), and end-to-end 72-hour inference time (Time, assessing the timeliness of high-frequency operational deployment).
[0163] Table 1. Quantitative comparison of different methods under 24-hour extreme weather scenarios and inference timeliness.
[0164]
[0165] The experimental results are shown in Table 1 and Figure 4 , Figure 5 , Figure 6 , Figure 7 As shown.
[0166] Long-term forecast stability and error accumulation analysis, such as Figure 4 As shown:
[0167] The trend of 72-hour MAE change in core performance indicators ( Figure 5In terms of prediction accuracy, ConvLSTM and UNet methods are limited by their two-dimensional perspective, resulting in prediction errors that diverge significantly linearly or even exponentially over time (Lead Time), making them unsuitable for medium- to long-term forecasting tasks. Numerical weather prediction (NWP) demonstrates the long-term stability advantage of traditional physical assimilation. In contrast, the error curve of the method in this invention (SimVP+3D-Swin) is extremely flat throughout the entire 72-hour period, with its overall MAE significantly lower than ConvLSTM and UNet. Furthermore, in the initial 0-42 hours, its prediction accuracy even surpasses the industry gold standard, numerical weather prediction (NWP). Even at the end of the 72-hour period, its error closely matches the NWP results calculated by supercomputers. This fully demonstrates that the three-dimensional cyclic shift window and relative position offset introduced by 3DSwin-Transformer successfully establish dynamic constraints within the deep network that conform to the long-range advection laws of the atmosphere, effectively suppressing the accumulation of extrapolation errors.
[0168] Extreme weather sensitivity and dynamic attribution ability, such as Figure 5 As shown:
[0169] Extreme weather forecast error within the most challenging 24-hour period ( Figure 5 In comparison, the RMSE of traditional deep learning methods (ConvLSTM, UNet) exhibits extremely severe penalized amplification (reaching 4.31 and 3.59 respectively), revealing its serious "averaging sluggishness" defect when facing sudden weather changes such as typhoons and strong wind shear. In contrast, the method of this invention not only outperforms other deep learning models across the board in extreme scenarios, but its MAE (1.24) and RMSE (1.82) even significantly outperform numerical weather prediction (NWP) assimilation systems. This breakthrough demonstrates that the MAE and RMSE joint loss optimization strategy designed in this invention forces the model to break its smoothing tendency, endowing the network with extremely high sensitivity and accurate order-of-magnitude reconstruction capabilities for local extreme high wind speeds.
[0170] Extrapolating computational efficiency and operational deployment potential, such as Figure 6 As shown:
[0171] In the timeliness of reasoning ( Figure 6(Note the logarithmic scale in the chart) To obtain the same 72-hour high-precision global / regional forecast, traditional numerical weather prediction (NWP) relies on massive computing clusters, often taking 1 to 2 hours, resulting in inherent data time lag. In contrast, the method of this invention, after offline training, generates a 3-day three-dimensional spatiotemporal wind and wave field sequence through forward inference in only 14.5 seconds. Although it takes slightly longer than the extremely simple UNet (12 seconds) to capture three-dimensional advection patterns, it achieves far greater accuracy than UNet and reaches NWP-level forecast quality while compressing computation time from "hours" to "seconds." This demonstrates that the spatiotemporal decoupling architecture of this invention possesses extremely high parallel computing efficiency, perfectly meeting the stringent requirements of "high-precision, zero-time-difference, and high-frequency" forecasts in industrial scenarios such as real-time scheduling of offshore wind power generation and minute-level early warning of sudden disasters.
[0172] Qualitative visualization analysis, such as Figure 7 As shown:
[0173] Figure 7 The method of this invention visually demonstrates a two-dimensional grid comparison between the extrapolated zonal wind (U component) and meridional wind (V component) forecasts for the next 6 to 24 hours and the actual analysis field (true values). It can be seen that as the forecast step size increases (12h, 18h, 24h), the wind and wave fields predicted by the method of this invention not only maintain a high degree of spatial consistency with the true values in the macroscopic large-scale circulation zone, but also exhibit a strong topology preservation ability in local complex wind speed gradients and vortex structures (the dark extreme value areas in the figure), effectively avoiding the "texture smoothing" and "structural blurring" phenomena that occur in traditional extrapolation models over time. This is attributed to the "zero-time-difference" high-fidelity reconstruction of SimVP in the first stage of this invention, and the fidelity preservation mechanism of the second stage three-dimensional sub-pixel convolution when amplifying spatial resolution.
[0174] Example 2
[0175] This embodiment provides a two-stage real-time wind and wave field forecasting system based on geostationary satellites.
[0176] A computer-readable storage medium storing a plurality of instructions adapted for loading and execution by a processor of a terminal device of the two-stage real-time wind and wave field forecasting method based on geostationary satellites.
[0177] A terminal device includes a processor and a computer-readable storage medium, the processor being used to implement various instructions; the computer-readable storage medium being used to store multiple instructions, the instructions being adapted to be loaded and executed by the processor to provide a two-stage real-time wind and wave field forecasting method based on geostationary satellites.
[0178] The above are all preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape and principle of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A two-stage real-time wind and wave field forecasting method based on geostationary satellites, characterized in that, include: Acquire multi-channel observation data from geostationary meteorological satellites; Data preprocessing is performed based on the acquired observation data; The process involves inverting the spatiotemporal sequence of cross-modal wind and wave fields based on preprocessed data, including spatial multi-scale information encoding based on grouping and normalization, spatiotemporal cross-modal conversion based on the principle of cross-correlation, and spatial decoding and reconstruction of physical wind and wave fields. Multi-scale wind and wave field evolution forecasting is performed based on the 3DSwin-Transformer model, including three-dimensional spatiotemporal information embedding and local dynamic encapsulation, multi-scale spatiotemporal attention modeling, long-term wind and wave field reconstruction and forecast output; The data preprocessing based on the acquired observation data includes assuming the original wide-swath satellite data is... ,in and Representing spatial longitude and latitude respectively, according to preset longitude and latitude boundaries. Select data, acquire observation data within the target area, and organize it into... ,in Indicates the number of channels. and This indicates the height and width of the spatial grid representing the target forecast area; Then, sensitive channel filtering and region selection are performed, assuming the original wide-swath satellite data is... After selecting the target region, the observation tensor is obtained. ; The selected satellite observation tensor will then be... Spatial resolution and size adjustment were performed on the ERA5 wind and wave field ground truth data. Let the satellite data after unified projection onto the target grid be denoted as... For the center of each target pixel The calculation is performed using bilinear interpolation: ,in, This refers to the wind or wave field data at four neighboring latitude and longitude coordinates from the original satellite data. Aligned satellite projection data are obtained after downsampling, based on interpolation weights calculated from distance. Finally, based on spatial grid alignment, a time reference determination is performed on the satellite data sequence and the ERA5 wind and wave field data for any given time. ,in Given the time series set of true ERA5 wind and wave fields, the synchronous observation frame with the smallest time deviation is selected as the input information from the minute-level high-frequency satellite observation time series. Then, according to the preset time resolution, the satellite projection data at the corresponding time is... With the true value of the storm field Composition of spatiotemporal paired samples ; The spatial multi-scale information encoding based on grouping normalization includes constructing a cross-modal inversion model based on the deep fusion of a multi-temporal cloud wind-guiding module and a SimVP model architecture. Using multi-channel satellite data sequences as input, and leveraging the spatial displacement of continuous cloud images, the system deduces the current moment's delay-free two-dimensional physical wind and wave field, achieving physical reconstruction from a spatiotemporal sequence to a single frame. First, let the input sequence tensor be... Where the time step T, the number of channels C corresponds to the channels of the satellite data, and H and W represent the height and width of the spatial grid of the target prediction area. The spatial encoder uses a convolutional network with shared weights to process the input sequence. Each time step in the algorithm learns local spatial information independently, and the encoder is composed of... It consists of cascaded convolutional modules, for any time step in the sequence In layer l, the information mapping and propagation mechanism is defined as follows: , ,in The pre-activation features of the l-th layer at time step k; and For learnable convolution kernels and biases, and These represent the feature representations of the (l-1)th and lth layers of the encoder, respectively; k is the temporal sample index, and l is the network layer index; This represents the convolution operation; For group normalization operations; The activation function is nonlinear; after n layers of spatial dimensionality reduction and information abstraction, the information outputs from multiple time steps are jointly represented along the time dimension to generate a latent spatiotemporal dynamics representation. .
2. The two-stage real-time wind and wave field forecasting method based on geostationary satellites according to claim 1, characterized in that, The spatiotemporal cross-modal transformation based on the principle of cross-correlation includes implicitly implementing cross-correlation template matching in the cloud wind-guiding module (CMW) within a deep neural network. Before the data enters the CMW, the time dimension is first superimposed onto the axis containing the feature channel dimension C. This is used to forcibly flatten and stitch the time dimension along the channel axis without sacrificing spatial resolution, as shown below: , ,in, Input a spatiotemporal information sequence tensor. The discrete-time step single-frame latent spatial information tensors correspond to the input sequences respectively. In the past, the first Time, Number Time and the present The satellite cloud image at any given time is a spatially encoded independent information slice; B is the number of samples in the batch; T is the time step; D is the number of channels. The height is compressed in both the height and width dimensions; This is represented as a data reorganization operation; The spatiotemporal composite information tensor after collapse and superposition; within each Inception unit, the structure is pruned and optimized based on the physical characteristics of wind field evolution, and is represented as: , in, This represents the dynamic spatiotemporal information representation output by the nth layer Inception unit. This indicates a channel splicing operation. This represents the spatiotemporal dynamics of the potential layer input to the current nth layer. Indicates use The convolution kernel performs a two-dimensional convolution operation; then, an improved residual shortcut mechanism is introduced in CMWBlock, represented as: ,in, This represents the dynamic spatiotemporal information output by the nth layer Inception unit. This represents the function for mapping jump connection information and aligning dimensions.
3. The two-stage real-time wind and wave field forecasting method based on geostationary satellites according to claim 2, characterized in that, The physical wind and wave field spatial decoding and reconstruction includes obtaining the latent representation, upsampling using a decoder, and reconstructing the wind and wave field onto the target grid. The decoder consists of 4 layers of deconvolution modules symmetrically stacked, represented as follows: ,in, This represents the transpose convolution operation. Indicates the first The learnable transposed convolution kernel weight tensor in the layer transposed convolution module. Indicates the first Learnable bias vectors in the deconvolution module Indicates input to the current number The pre-order low-resolution semantic information tensor of the layer space decoder. The output feature tensor of the l-th layer decoder. The group normalization operation is then performed. Finally, after dimensionality reduction via a linear projection layer, the initial physical wind and wave field matching the target mesh is output. ,in, The height and width of the spatial grid for the target prediction region are determined. In offline training, the mean squared error (MSE) is used as the loss function for spatiotemporal reconstruction to measure the wind field at the current moment predicted by the model. ERA5 high-precision analysis field with strict time alignment Deviation between: , in, This represents the global reconstruction loss of the first-stage cross-modal inversion network. For a set of learnable parameters, This indicates the number of samples processed in a single model iteration training. This indicates the output of the linear projection layer after network inference. One sample in Predicting the initial field from the physical wind and wave field at a given moment. For high-precision wind and wave field measurements. It represents the squared L2 norm expanded in the spatial and channel dimensions.
4. The two-stage real-time wind and wave field forecasting method based on geostationary satellites according to claim 3, characterized in that, The embedding of the three-dimensional spatiotemporal information and the encapsulation of local dynamics include constructing a spatiotemporal prediction model with long-term memory and multi-scale dynamic evolution capture capabilities based on the acquired high-precision real-time initial field. Before performing model inference, define the tensor dimension of the spatiotemporal data and predict the model. The input is a continuous wind and wave field sequence over a past period, and the input time step is set to... The input tensor is then represented as Where 2 represents zonal wind and the wind The model has two channels, and its output target is the future wind and wave field evolution sequence. The output time step is set to... The predicted output tensor is then represented as Then, a three-dimensional convolution kernel is used to process the input tensor. Synchronous segmentation is performed along the time dimension, spatial height, and width. Let the segmentation size of the 3D convolutional kernel be... And capture short-term wind field evolution trends in the shallow layer of the network, assuming In the spatial dimension, each The physical grid points are divided into local surfaces; finally, each segmented 3D data block is flattened and linearly projected into a one-dimensional token vector through a learnable 3D convolutional layer, represented as: , , in, The input is the continuous wind and wave field spacetime tensor. It is a three-dimensional convolutional layer. and These represent the receptive field size and stride of the 3D convolution kernel, respectively. This refers to the high-dimensional raw data tensor that has undergone 3D convolutional projection but has not yet been flattened. For layer normalization operation, For high-dimensional mapping dimensions, This is the initial token sequence tensor that, after being flattened and normalized, is finally input into the 3DSwin-Transformer module. This represents the total number of three-dimensional spatiotemporal blocks.
5. The two-stage real-time wind and wave field forecasting method based on geostationary satellites according to claim 4, characterized in that, The multi-scale spatiotemporal attention modeling includes learning the initial spatiotemporal information sequence based on the nonlinear migration and evolution patterns of large-scale weather systems across grids and time periods. Passing through in sequence by A 3DSwin-Transformer module consisting of four stacked stages is configured. Within each stage, paired network blocks containing both conventional 3D window multi-head self-attention 3DW-MSA and shifted 3D window multi-head self-attention 3DSW-MSA are alternately stacked. First, the massive spatiotemporal data is divided into multiple non-overlapping 3D windows using the 3DW-MSA mechanism. The window size is set to... The computational complexity of a single-layer self-attention system is compressed to [number missing]. Where T is the time step, H×W represents the window size; H×W represents the height and width of the spatial grid for the target prediction area; D represents the channel dimension; within each local 3D window, the self-attention calculation formula is defined as: ,in Let represent the query matrix, key matrix, and value matrix generated by mapping the input spatiotemporal data sequence through a linear projection matrix within the current local 3D window, respectively. Let d represent the information mapping channel dimension under the single-head attention mechanism. This represents the dot product of the query matrix and the transpose of the key matrix. The channel dimension represents the key vector. This represents the 3D relative position bias matrix directly superimposed on the self-attention scoring matrix; then, a learnable 3D parameter bias table is maintained. To parameterize Time lag range: ,total Possible relative time differences; spatial height distance range: ,total Local vertical / meridian relative displacement, spatial width range: ,total Local horizontal / zonal relative displacements are obtained by explicitly encoding the temporal lag and spatial relative distance between any two air masses within a local window, from the offset table. The corresponding bias value is indexed and added to the attention score.
6. The two-stage real-time wind and wave field forecasting method based on geostationary satellites according to claim 5, characterized in that, The multi-scale spatiotemporal attention modeling also includes performing three-dimensional cyclic shifting using the 3DSW-MSA mechanism in subsequent Transformer layers. Before the data enters the attention calculation, the window is shifted in the time, height, and width directions respectively. The network iterates through patches in a loop, then introduces a 3D masking mechanism to construct a fine-grained 3D mask matrix within the network. In the new window for shifting and reassembling, the data is divided into multiple unrelated sub-windows based on the different grid regions from which it originates. Before performing the Softmax operation, a mask matrix is superimposed on the attention score, represented as: , in These represent the query, key, and value vectors, respectively, obtained from the input features through linear mapping; d represents the information mapping channel dimension under the single-head attention mechanism. represents the channel dimension of the key vector; b represents the three-dimensional relative position bias matrix directly superimposed on the self-attention scoring matrix; This represents the attention mask matrix; for data belonging to the same original adjacent region, its corresponding... The value is 0; for non-adjacent token data that are forcibly merged due to cyclic shifting, its corresponding value is 0. Value Within each pair of 3DW-MSA and 3DSW-MSA blocks, the flow formula for latent layer information is defined as: , in, and They represent the first The intermediate information between the first and second layers is represented by LayerNorm, which represents the layer normalization operation, and MLP represents the feedforward multilayer perceptron structure.
7. The two-stage real-time wind and wave field forecasting method based on geostationary satellites according to claim 6, characterized in that, The long-term wind and wave field reconstruction and forecast output includes the abstract information sequence that reaches the top layer of the network after being learned through multiple layers of 3DSwin-Transformer. Containing highly condensed patterns of future evolution, it employs a structure inverse of downsampling, using linear mapping to reduce the channel dimension from... Compress to In conjunction with tensor reshaping operations, the one-dimensional token sequence is folded back into a four-dimensional spatiotemporal coordinate system; in the spatial dimension, three-dimensional subpixel convolution is used to smoothly amplify the probabilistic receptive field and resolution at each level, wherein the subpixel channel mapping expansion formula is: , ,in, These represent the highly compressed dimensions of the latent information in the time, height, and width dimensions, respectively; B is the number of samples in the batch; and D is the channel dimension. This is represented as a data reorganization operation. For learnable weight matrix, For learnable bias, a 3D subpixel spatial rearrangement is then performed using a periodic shuffling operator to interweave and rearrange the elements in the dilation channel dimension according to deterministic geometric rules into the spatial height and width dimensions: ,in, The underlying element index mapping relationship is defined as follows: ,in The rearranged high-resolution output tensor Low-resolution, high-channel-count dilation tensor before rearrangement. This represents the target channel index on the high-resolution output data. These represent the time steps, This represents the spatial upsampling rate. After three-dimensional subpixel rearrangement, the channels are equivalently converted into spatial data. This represents the spatial index of high-resolution coordinates in a low-resolution feature map, while Represents the rearranged index of the subpixel in the channel dimension; the output tensor precisely evolves into After restoring to the target spatial resolution Then, a three-dimensional anisotropic convolutional layer and a linear regression head are connected at the end to map the reconstructed latent layer information to the physical channels for the next few days, as shown below: , in, This represents the final latent space feature representation of the encoder or spatiotemporal modeling module; (⋅) indicates a three-dimensional subpixel rearrangement or patch expansion operation; and These represent the learnable convolutional kernel parameters and bias terms of the output layer, respectively. Represents a three-dimensional convolution operation. This represents the final physical field prediction result output by the model.
8. A two-stage real-time wind and wave field forecasting system based on geostationary satellites, executing the two-stage real-time wind and wave field forecasting method based on geostationary satellites as described in claim 1, characterized in that, include: The data acquisition module is configured to acquire multi-channel observation data from geostationary meteorological satellites. The data preprocessing module is configured to perform data preprocessing based on the acquired observation data; The inversion module is configured to perform cross-modal wind and wave field spatiotemporal sequence inversion based on preprocessed data, including spatial multi-scale information encoding based on grouping normalization, spatiotemporal cross-modal conversion based on the principle of cross-correlation, and physical wind and wave field spatial decoding and reconstruction. The forecasting module is configured to perform multi-scale wind and wave field evolution forecasting based on the inversion results using 3DSwin-Transformer, including three-dimensional spatiotemporal information embedding and local dynamic encapsulation, multi-scale spatiotemporal attention modeling, long-term wind and wave field reconstruction and forecast output.
Citation Information
Patent Citations
Multi-modal deep learning fusion ocean remote sensing sea wave parameter inversion method and system
CN120995032A
Sea wave prediction method based on weather forecast field data fusion and space-time attention mechanism
CN121614778A