Marine-atmospheric Coupled Data Sampling Method Based on Asynchronous Cross-Iterative Random Sampling Strategy

Through asynchronous cross-iteration random sampling strategy and Swin Transformer V2 model, the problems of time lag and resolution in air-coupled modeling are solved, and efficient and accurate forecast of ocean state variables is achieved, reducing calculation costs.

CN119670585BActive Publication Date: 2025-07-11INST OF OCEANOLOGY - CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510191937.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-07-11
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The existing air-coupled modeling method fails to effectively simulate the time lag effect of the atmosphere on ocean variables, it is difficult to deal with the spatial resolution inconsistency between ocean and atmospheric data, and the computing resources are limited, making it difficult to take into account the needs of high efficiency and high resolution modeling, and the complex nonlinear air-coupled relationship modeling capabilities are insufficient.

Method used

The asynchronous cross-iteration random sampling strategy is adopted to simulate the lag effect of the atmosphere on ocean variables through sliding time windows and random sampling, and a multi-card parallel computing framework is used for asynchronous sampling, and a data processing model of Swin Transformer V2 structure is constructed to achieve data diversity and efficient fusion.

Benefits of technology

It significantly improves data sampling efficiency and model forecasting accuracy, reduces computing resource requirements, improves the accuracy and computing efficiency of ocean state variable forecasting, and is especially suitable for high-resolution modeling tasks with resource constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119670585B_ABST
    Figure CN119670585B_ABST
Patent Text Reader

Abstract

The present invention discloses an air-sea coupling data sampling method based on an asynchronous cross-iterative random sampling strategy, belonging to the technical field of marine meteorology. Through the asynchronous cross-iterative random sampling strategy, the present invention overcomes many deficiencies in the existing air-sea coupling modeling methods and shows significant advantages in terms of data sampling efficiency, model accuracy, and computational resource utilization. To improve the data sampling efficiency, the present invention adopts a sliding time window and a random sampling mechanism to simulate the lag effect and randomness between atmospheric and oceanic variables, and at the same time, the asynchronous cross-iterative strategy significantly improves the sampling efficiency. Through multi-card asynchronous scheduling and non-blocking sampling mechanisms, each computing node can complete the sampling task in parallel. Compared with the traditional serial sampling method, the sampling efficiency is increased by about 30%. The mechanism of dynamically adjusting the sampling weight focuses resources on high-error regions, avoiding redundant sampling of irrelevant or low-impact data, and further reducing the waste of computational resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of marine meteorology. Specifically, it particularly relates to a sea-air coupling data sampling method based on an asynchronous cross-iteration random sampling strategy. Background Art

[0002] The coupling between the ocean and the atmosphere is an important part of the Earth system, and is of great significance for climate change, extreme weather events, and high-precision forecasting of ocean state variables. The changes in ocean state variables (such as sea surface temperature, sea surface height, and ocean currents) do not occur in isolation, but are driven by complex sea-air interactions. These interactions involve multi-scale, multi-variable dynamic feedback mechanisms, and are of important value for the study of Earth system dynamics, disaster forecasting and prevention, and the sustainable utilization of ocean resources. Therefore, developing an efficient forecasting method that can comprehensively capture the sea-air coupling relationship is an important direction for promoting Earth system science and practical applications.

[0003] Existing sea-air coupling modeling methods mainly include numerical simulation models, data-driven methods, and data splicing and fusion techniques. Among them, numerical simulation models can simulate the dynamic process by solving the physical equations of the sea-air system, but have high computational costs and limited ability to capture small-scale changes; data-driven methods, such as spatio-temporal convolutional networks and long short-term memory networks (LSTM), perform well in processing a single subsystem of the ocean or the atmosphere, but often fail to fully capture the complex coupling relationship between the two; in addition, some studies use simple data splicing methods for modeling, but this method fails to reflect the causal relationship and time lag effect between variables, resulting in limited forecasting accuracy.

[0004] Although certain progress has been made in existing methods, there are still many problems. First, the difference in the time scales of ocean and atmospheric changes has not been fully considered, and existing methods are difficult to effectively simulate the time lag effect of the atmosphere on ocean variables. Second, the inconsistency in the spatial resolution of ocean and atmospheric data is likely to cause information loss or error accumulation during data fusion. In addition, in the case of limited computing resources, existing coupling models are difficult to balance the requirements of high efficiency and high-resolution modeling. More importantly, the current methods have insufficient modeling ability for complex, non-linear sea-air coupling relationships, and can often only capture local or shallow associations, limiting the improvement of forecasting performance. Summary of the Invention

[0005] The object of the present invention is to construct a sea-air coupling data sampling method based on an asynchronous cross-iteration random sampling strategy, aiming to improve the accuracy and computational efficiency of ocean state variable forecasting by simulating the complex interactions between the atmosphere and the ocean, so as to make up for the deficiencies of the existing technology.

[0006] To achieve the above object, the present invention is realized through the following specific technical solutions:

[0007] An air-sea coupling data sampling method based on an asynchronous cross-iteration random sampling strategy, comprising the following steps:

[0008] S1: Collect ocean variable data and atmospheric variable data, and perform data preprocessing;

[0009] S2: Use the sliding time window mechanism to simulate the lag effect of the atmosphere on ocean variables, and enhance the diversity of data through the random sampling method;

[0010] S3: The cross-iteration sampling strategy is carried out asynchronously, dynamically optimizing the sampling process under the multi-card parallel computing framework, and realizing the seamless connection between sampling and training;

[0011] S4: Construct a data processing model including: a data sampling module, an encoder, and a decoder; among them, both the encoder and the decoder adopt the Swin Transformer V2 structure;

[0012] S5: The data processing model is trained based on the sampled data, and the prediction performance and air-sea coupling modeling ability are verified.

[0013] Further, in the S1, by unifying the spatio-temporal resolution of ocean and atmospheric variables, a consistent input basis is provided for subsequent sliding window sampling and model training.

[0014] Furthermore, the S1 includes:

[0015] S1-1: The ocean variable data adopts GLORYS12 reanalysis data with a horizontal resolution of 1 / 12°, including 50 vertical layers; the atmospheric variable data adopts ERA5 reanalysis data with a horizontal resolution of 0.25°, including 37 vertical isobaric layers;

[0016] S1-2: Interpolation processing: Use the bilinear interpolation method to interpolate the atmospheric variables of ERA5 to the resolution of GLORYS12 (1 / 12°), and the interpolation formula is:

[0017] ;

[0018] Among them, represents the interpolated atmospheric variables, represents the original atmospheric variables, is the resolution of GLORYS12 ocean variables, is the interpolation function; the nearest neighbor interpolation method is adopted in the present invention;

[0019] S1-3: Time alignment: Align the two types of data to a daily time step to ensure consistency in the time dimension. Through this step, multi-source coupled data with unified resolution and time step is generated, providing a standardized input for the subsequent sliding window sampling mechanism.

[0020] Further, the said S2 includes:

[0021] S2-1: Set a time impact window for the ocean state variable forecast of each day, limit the sampling range of atmospheric variables, and define as follows:

[0022] ;

[0023] Wherein, represents the current time of the time window, and are the maximum and minimum lag times respectively (in the present invention, = 7, = 3) days); The reason for setting the time window is to characterize that the current ocean state variable may be affected by the atmospheric conditions in the past 3 to 7 days by setting the maximum and minimum lag times.

[0024] S2-2: Random sampling: Randomly select a time point within the time window , and sample the corresponding atmospheric variables. The sampling method is as follows:

[0025] ;

[0026] Wherein, represents the atmospheric variable corresponding to the selected time point , represents the sampling time range; The random sampling method reflects the randomness and complexity of the atmosphere on ocean variables.

[0027] S2-3: Training sample combination: Combine the atmospheric variables obtained by random sampling with the ocean variables at the current time point to form a complete training sample:

[0028] ;

[0029] Wherein, represents the input features for training, including the ocean state variable at the current time and the atmospheric variable at time . Through this step, training samples considering lag effects and randomness are generated, providing rich and diverse data support for subsequent modeling.

[0030] Furthermore, S3 includes:

[0031] S3-1: Sampling task allocation: Allocate random sampling tasks to multiple computing nodes to ensure that each node can independently complete the sampling operation and avoid resource conflicts;

[0032] S3-2: Cross-iteration update: Dynamically adjust the sampling weights of atmospheric and ocean variables according to the prediction error of the data processing model, and the update state is carried out using the following formula:

[0033] ;

[0034] Wherein, and respectively represent the sampling probabilities of atmospheric and ocean variables, and are the prediction errors of atmospheric and ocean variables respectively; in the sampling scheme, the sampling probability is inversely proportional to the prediction error, the larger the error, the higher the sampling probability, so as to enhance the attention to the high-error area;

[0035] S3-3: Asynchronous execution: Adopt a non-blocking scheduling mechanism to enable the sampling task and model training to proceed simultaneously, further improving the computing efficiency; through the asynchronous cross-iteration strategy, the dynamic optimization of the sampling process is realized, enhancing the modeling ability of the air-sea coupling relationship while ensuring the efficiency.

[0036] Furthermore, the data processing model in S4 specifically includes:

[0037] (1) Data sampling module

[0038] The data sampling module is responsible for processing the input original data and generating training samples according to a specific sampling strategy (such as an asynchronous cross-iteration random sampling strategy). The role of this module is to preprocess and sample the air-sea coupling data to ensure the quality and diversity of the input data. The specific steps include:

[0039] Data preprocessing: Interpolate and normalize the original data (such as atmospheric variables and ocean variables) to ensure the consistency and usability of the data.

[0040] Sampling strategy: Implement asynchronous cross-iteration random sampling according to the model requirements to ensure that the time-series data affecting the future state is extracted from different time windows. This strategy helps to capture the dynamic changes on long time scales, especially when dealing with the complexity of air-sea coupling.

[0041] Output: Transmit the sampled and processed data to the encoder.

[0042] (2) Encoder

[0043] The encoder is responsible for extracting high-level feature representations from the input data. Swin Transformer V2 is highly expressive when dealing with image and sequence data. It uses the window attention mechanism and hierarchical feature extraction strategy to efficiently capture local and global features.

[0044] Patch Embedding: First, the input two-dimensional data (such as the time series data of air-sea coupling) is divided into non-overlapping image patches. The size of each image patch is adjusted according to the resolution of the input data. These image patches are then mapped to high-dimensional embedding vectors through a linear transformation.

[0045] Window Attention: In Swin Transformer V2, the window self-attention mechanism (WindowAttention) is its core feature. Each Transformer layer captures local context information by calculating the relationships between data within local windows. To avoid the breaking effect between windows, Swin V2 adopts the Shifted Window mechanism, allowing information transfer across windows to better model global dependencies.

[0046] Multi-level structure: The encoder adopts a hierarchical Transformer structure. As the number of layers increases, the spatial resolution gradually decreases while the depth of the feature map increases. By aggregating information layer by layer, the encoder can effectively extract features at different scales.

[0047] Output: The output of the encoder is a high-dimensional feature representation that contains in-depth information of the input data, including both local features and global information across windows.

[0048] (3) Decoder

[0049] The decoder is responsible for mapping the features extracted by the encoder to the final output space, such as predicting ocean state variables, generating air-sea coupling forecast results, etc.

[0050] Reverse construction: The working mode of the decoder is similar to that of the encoder, but its goal is to recover the specific output from the high-dimensional features generated by the encoder. The decoder gradually refines the output of each layer through backpropagated information to generate high-quality prediction results.

[0051] Cross-layer feature fusion: In each layer of the decoder, through a specific fusion mechanism, low-level features and high-level features from the encoder are combined to ensure the accuracy of the output. The decoder not only relies on the features from the encoder but can also focus on different parts of the input data through the self-attention mechanism for effective feature restoration.

[0052] Generate the output: The final output of the decoder is the result of prediction and regression processing, such as the predicted values of variables such as sea surface temperature and ocean current velocity.

[0053] Furthermore, the S5 includes:

[0054] S5-1: Training sample construction, using the sampled data to generate input features and target outputs:

[0055] ;

[0056] Among them, is the set of input features, is the target output (i.e., the ocean state variable at the next time step);

[0057] S5-2: Model training, realizing the prediction of ocean state variables by optimizing the loss function to minimize the prediction error. The optimization formula is as follows:

[0058] ;

[0059] Among them, is the loss function, is the true value, is the model prediction value, is the total number of samples. By optimizing the loss function, focus on minimizing the prediction error to improve the model's prediction ability for the target variable.

[0060] S5-2: Result verification: Use the test data set to evaluate the prediction accuracy of the model and the ability to characterize the air-sea coupling relationship.

[0061] Compared with the prior art, the beneficial effects of the present invention are:

[0062] Through the asynchronous cross-iteration random sampling strategy, the present invention overcomes many deficiencies in the existing air-sea coupling modeling methods and shows significant advantages in terms of data sampling efficiency, model accuracy, and computing resource utilization. The specific manifestations are as follows: 1. Improve data sampling efficiency. The present invention adopts a sliding time window and a random sampling mechanism to simulate the lag effect and randomness between atmospheric and ocean variables. At the same time, the asynchronous cross-iteration strategy significantly improves the sampling efficiency. Through multi-card asynchronous scheduling and a non-blocking sampling mechanism, each computing node can complete the sampling task in parallel. Compared with the traditional serial sampling method, the sampling efficiency is increased by about 30%. The mechanism of dynamically adjusting the sampling weight focuses resources on high-error regions, avoiding redundant sampling of irrelevant or low-impact data, and further reducing the waste of computing resources.

[0063] In summary, through technological innovation, the present invention has achieved a double improvement in data sampling efficiency and model prediction accuracy. It is particularly applicable to high-resolution air-sea coupling modeling tasks, while reducing the usage cost of computing resources, and has important application value and practical significance. The present invention can provide technological breakthroughs in complex coupling relationship modeling, high-resolution data fusion, and computational efficiency optimization, and provide strong support for high-precision prediction of ocean state variables and earth system science research. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is a design concept diagram of air-sea coupling data for the asynchronous cross-iterative random sampling strategy of the present invention.

[0065] Figure 2 It is a basic architecture design diagram of the data processing model of the present invention.

[0066] Figure 3 It is a result comparison diagram between the present invention and existing methods; (a) is the RMSE comparison of the east-west component of ocean current between the present invention and existing methods; (b) is the RMSE comparison of the north-south component of ocean current between the present invention and existing methods; (c) is the RMSE comparison of sea temperature prediction between the present invention and existing methods; (d) is the RMSE comparison of salinity prediction between the present invention and existing methods. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] The following further describes and illustrates the technical solutions of the present invention in combination with embodiments and drawings.

[0068] Aiming at the deficiencies of existing air-sea coupling modeling methods, the present invention aims to solve the following technical problems: 1. Solve the problem of time scale difference: Existing methods do not fully consider the time scale difference between ocean and atmospheric changes and cannot accurately simulate the time lag effect of the atmosphere on ocean variables. The present invention designs a sliding time window and a random sampling mechanism to achieve dynamic modeling of the impact of atmospheric changes on ocean state variables, thereby improving the physical consistency and accuracy of prediction. 2. Achieve cross-resolution data fusion: Aiming at the information loss and error accumulation problems caused by the inconsistency of spatial resolutions of ocean and atmospheric data, the present invention proposes a cross-iterative sampling strategy to achieve effective fusion of data with different resolutions through flexible data selection and dynamic optimization mechanisms. 3. Improve computational efficiency and adaptability: Considering the high demand for computing resources of high-resolution air-sea coupling models, the present invention executes sampling operations asynchronously under a multi-level multi-card parallel computing framework, significantly improving computational efficiency while maintaining the accuracy of coupling modeling, and being applicable to practical application scenarios with limited resources.

[0069] The present invention proposes an air-sea coupling data sampling method of "asynchronous cross-iterative random sampling strategy" (the design concept is as Figure 1As shown, by simulating the time-lag effect of the atmosphere on the ocean and combining a sliding time window with a random sampling mechanism, the complex dynamic characteristics of the air-sea coupling system are effectively captured, which is especially suitable for high-resolution modeling requirements under resource-constrained conditions.

[0070] In the following embodiments, for the prediction of air-sea coupling data, a newly proposed asynchronous cross-iteration random sampling strategy is combined and compared with an existing numerical model (the PSY4 model of Mercator Ocean International). Among them, the data sets include:

[0071] (I) GLORYS12 data set:

[0072] GLORYS12 is a global ocean reanalysis data set designed under the framework of the European Copernicus Marine Environment Monitoring Service (CMEMS), with a spatial resolution of 1 / 12° and 50 vertical levels, covering data from January 1993 to December 2020. As the main training data set, GLORYS12 contains high-quality ocean state variable data, such as sea surface temperature (SST), salinity, ocean current, etc. These data have high temporal continuity and spatial resolution and are suitable for high-resolution ocean state prediction tasks.

[0073] (II) ERA5 atmospheric data set:

[0074] To perform coupled modeling by combining atmospheric and ocean data, we adopted the ERA5 data set, which is provided by the European Centre for Medium-Range Weather Forecasts (ECMWF), has a spatial resolution of 0.25°, and includes atmospheric variable data such as wind speed, air temperature, and humidity at 37 vertical levels.

[0075] (III) Numerical model (PSY4) data set:

[0076] PSY4 is an ocean numerical forecasting model released by Mercator Océan International in France, which provides predictions of state variables of the global ocean, including temperature, salinity, flow velocity, etc. PSY4 uses traditional numerical weather forecasting methods and numerically solves the air-sea coupling equations to predict the ocean state. This model has wide applications in short-term forecasting tasks of global ocean and atmosphere interactions.

[0077] Example 1

[0078] The sea surface temperature (SST) forecasting method based on GLORYS12 and ERA5 data includes the following steps:

[0079] 1. Data preprocessing and resolution matching

[0080] 1.1. Input data:

[0081] Ocean data: GLORYS12 reanalysis data, including variables such as sea surface temperature (SST) and sea surface salinity (SSS), with a horizontal resolution of 1 / 12°, and a time range from 1993 to 2020.

[0082] Atmospheric data: ERA5 reanalysis data, including variables such as wind speed (U10, V10), air pressure, and temperature, with a horizontal resolution of 0.25°, and a time range from 1993 to 2020.

[0083] 1.2. Interpolation processing:

[0084] Using the bilinear interpolation method, interpolate the ERA5 data to a resolution of 1 / 12° to match the spatial resolution of the GLORYS12 data. The specific interpolation formula is as follows:

[0085] ;

[0086] Where, represents the interpolated value, are the values of the adjacent 4 grid points, are the interpolation weights, calculated based on the distance.

[0087] 1.3. Time alignment:

[0088] Align the two datasets according to the daily time step to ensure the time consistency of the atmospheric and ocean variables.

[0089] 2. Definition of sliding time window and random sampling

[0090] 2.1. Definition of time window:

[0091] Assume the current time is the -th day. To predict the SST on the -th day, define the influence time window of the atmospheric variables as days, and generate the window range:

[0092] ;

[0093] 2.2. Random sampling:

[0094] Randomly select a time point within the window , and extract the corresponding atmospheric variables. The formula is as follows:

[0095] ;

[0096] For example, for the = 50 -th day, the randomly sampled If it is 47, then select the atmospheric variables on the 47th day as one of the input features to participate in the prediction of SST.

[0097] 2.3. Training sample construction:

[0098] Combine the sampled atmospheric variables with the ocean variables on the th day to form a training sample:

[0099] ;

[0100] Among them, is the input feature, is the target output, that is, the SST on the th day.

[0101] 3. Asynchronous cross-iterative sampling strategy

[0102] 3.1. Multi-card parallel sampling:

[0103] Through the multi-card parallel computing framework, distribute the data sampling tasks at different time steps to multiple computing nodes and execute them asynchronously:

[0104] Node 1: Process the data from = 1 to = 100 days;

[0105] Node 2: Process the data from = 101 to = 200 days;

[0106] And so on for other nodes.

[0107] 3.2. Sampling weight adjustment: Dynamically adjust the sampling weight according to the prediction error of the model to increase the sampling probability in the high-error area:

[0108] ;

[0109] Among them, is the prediction error of the ocean state variables on the th day, is the sampling probability.

[0110] 3.3. Sampling and training coordination:

[0111] Asynchronous sampling and model training are carried out simultaneously. Once a sampling task is completed, the result is immediately sent to the training queue.

[0112] 4. Model training and validation

[0113] 4.1. Train the model (the model architecture design is asFigure 2 as shown in

[0114] Based on the sampled data, train a deep learning model (such as STNet) and optimize the loss function:

[0115] ;

[0116] where is the mean squared error loss, is the true value of SST, is the predicted value of SST.

[0117] 4.2. Verify the performance:

[0118] Use the test data to evaluate the prediction accuracy of the model and compare it with the traditional method:

[0119] Mean squared error (MSE): Reduced by approximately 15%.

[0120] Extreme condition prediction: The error is reduced by approximately 18%.

[0121] Through the above steps, an efficient air-sea coupling data sampling method has been successfully constructed. Each step is closely connected. Data preprocessing lays the foundation for sliding window sampling, random sampling enhances data diversity, the asynchronous cross-iteration strategy optimizes sampling efficiency, and model training finally achieves high-precision prediction of complex air-sea coupling relationships.

[0122] Example 2: Prediction of ocean dynamic variables

[0123] In this example, the same technical solution is adopted to predict ocean dynamic variables (such as flow velocity and flow direction).

[0124] 1. Data preprocessing: Match the resolution and align the time of the flow velocity and flow direction data.

[0125] 2. Time window and random sampling: Define -day time window, randomly sample atmospheric variables and combine them into input samples.

[0126] 3. Asynchronous sampling and training: The same as in Example 1, adopt the multi-card parallel and weight adjustment mechanism.

[0127] Experiments show that the mean squared error in the prediction of ocean dynamic variables by the present invention is reduced by approximately 12%, and the calculation efficiency is increased by approximately 20%. Experimental comparison shows that as Figure 3 shown: Under the same hardware conditions, the time required for the sampling process of the present invention is shortened by approximately 25% - 40% compared with the traditional method.

[0128] The present invention can improve the model prediction accuracy. Since the present invention comprehensively considers the complexity of air-sea interaction, the generated training samples can better reflect the real air-sea coupling relationship, significantly enhancing the prediction performance of the model. The sliding time window mechanism effectively captures the lag effect of the atmosphere on ocean variables, making the prediction results closer to the actual situation. The diversity of random sampling enhances the generalization ability of the model, avoiding the overfitting problem caused by the single data distribution. Experimental results show that: in prediction tasks such as sea surface temperature (SST), the root mean square error (RMSE) of the method of the present invention is reduced by 15%-20% compared with the traditional method.

[0129] The present invention can save computing resources and support high-resolution modeling. The present invention is specifically optimized for the resource limitation problem of high-resolution modeling and has the following characteristics: i) By means of data resolution matching and interpolation methods, the spatio-temporal resolutions of different data sources are unified, reducing the inconsistency of model inputs and laying a foundation for high-resolution modeling. ii) The asynchronous cross-iteration strategy makes full use of computing resources, reducing the computing overhead while ensuring the model performance. Compared with the traditional method, on the 1 / 12° high-resolution global reanalysis data, the computing resource requirements of the present invention are reduced by about 20%-30%, but it can still achieve high-precision prediction results consistent with the original resolution.

[0130] The present invention realizes the simplicity and adaptability of operations. The method structure of the present invention is clear and easy to be integrated into the existing air-sea coupling model. The data sampling and model training processes are completely modular, applicable to various reanalysis data (such as GLORYS12 and ERA5) and different prediction tasks (such as sea surface temperature, sea current, sea surface height, etc.).

[0131] Through the above embodiments, the present invention proves its superior performance in air-sea coupling modeling, providing an effective solution for high-resolution prediction tasks under resource-constrained conditions.

[0132] Finally, although this specification is described according to the embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A sea-air coupling data sampling method based on an asynchronous cross-iteration random sampling strategy, characterized in that It includes the following steps: S1: Collect ocean variable data and atmospheric variable data, and perform data preprocessing; S2: Use the sliding time window mechanism to simulate the lag effect of the atmosphere on ocean variables, and enhance the data diversity through the random sampling method; S2 includes: S2-1: Set a time impact window for the ocean state variable forecast of each day, limit the sampling range of atmospheric variables, and define as follows: W t = [t - n max , t - n min ​ Among them, W t represents the time window at the current time t, n max and n min are the maximum and minimum lag times respectively; S2-2: Random Sampling: Randomly select a time point t within the time window W t and sample the corresponding atmospheric variables. The sampling method is as follows: s ​ V atm (t s )~W t Among them, V atm (t s ) represents the selected time point t s corresponding atmospheric variable, and W t represents the sampling time range; S2-3: Training sample combination: Combine the randomly sampled atmospheric variables with the ocean variables at the current time point to form a complete training sample: X t = {V ocean (t), V atm (t s )} Among them, X t represents the input features for training, including the ocean state variable V ocean (t) and the atmospheric variable V s at time t atm (t s ); S3: The cross-iterative sampling strategy is carried out asynchronously, and the sampling process is dynamically optimized under the multi-card parallel computing framework; S3 includes: S3-1: Sampling task assignment: Assign the random sampling tasks to multiple computing nodes to ensure that each node can complete the sampling operation independently; S3-2: Cross-iterative update: Dynamically adjust the sampling weights of atmospheric and ocean variables according to the forecast error of the data processing model, and the update state is carried out using the following formula: where P(V atm ) and P(V ocean ) denote the sampling probabilities of atmospheric and oceanic variables, respectively, and L atm and L ocean are the forecast errors of atmospheric and oceanic variables, respectively; S3-3: Asynchronous execution: Adopt a non-blocking scheduling mechanism to make the sampling task and model training proceed simultaneously; Through the asynchronous cross-iterative strategy, the dynamic optimization of the sampling process is realized; S4: Construct a data processing model including: a data sampling module, an encoder, and a decoder; Among them, both the encoder and the decoder adopt the Swin Transformer V2 structure; S5: The data processing model is trained based on the sampled data, and the forecast performance and air-sea coupling modeling ability are verified.

2. The sea-air coupling data sampling method according to claim 1, characterized in that, S1 includes: S1-1: The ocean variable data uses GLORYS12 reanalysis data with a horizontal resolution of 1 / 12°, including 50 vertical layers; The atmospheric variable data uses ERA5 reanalysis data with a horizontal resolution of 0.25°, including 37 vertical isobaric layers, and unify the spatio-temporal resolution of ocean and atmospheric variables; S1-2: Interpolation processing: Use the bilinear interpolation method to interpolate the atmospheric variables of ERA5 to the resolution of 1 / 12° of GLORYS12, and the interpolation formula is: V′ atm = f interp (V atm , R ocean ) Among them, V′ atm represents the interpolated atmospheric variable, V atm represents the original atmospheric variable, R ocean is the resolution of the GLORYS12 ocean variable, f interp is the interpolation function; S1-3: Time alignment: Align the two kinds of data to a daily time step to ensure the consistency of the time dimension. Through this step, multi-source coupled data with unified resolution and time step is generated, providing a standardized input for the subsequent sliding window sampling mechanism.

3. The sea-air coupling data sampling method according to claim 1, wherein The data processing model in S4 specifically includes: (1) Data sampling module The data sampling module is responsible for processing the input original data, and generating training samples according to the asynchronous cross-iterative random sampling strategy; Preprocess the original data; Implement asynchronous cross-iterative random sampling to ensure that the time series data that affects the future state is extracted from different time windows; Transmit the sampled and processed data to the encoder; (2) The encoder is responsible for extracting high-level feature representations from the input data. By using the window attention mechanism and hierarchical feature extraction strategy, it can efficiently capture local and global features; Patch Embedding: The input two-dimensional data is divided into non-overlapping image patches, and the size of each image patch is adjusted according to the resolution of the input data. These image patches are then mapped into high-dimensional embedding vectors through a linear transformation; Window Attention: Each Transformer layer captures local context information by calculating the relationships between the data within local windows; Multi-level structure: The encoder adopts a hierarchical Transformer structure. As the number of layers increases, the spatial resolution gradually decreases while the depth of the feature map increases; The output of the encoder is a high-dimensional feature representation that contains the deep information of the input data, including both local features and global information across windows; (3) The decoder maps the features extracted by the encoder to the final output space; The decoder gradually refines the output of each layer through the information of backpropagation to generate high-quality prediction results; At each layer of the decoder, through a specific fusion mechanism, the low-level features and high-level features from the encoder are combined; The decoder not only relies on the features from the encoder but also can focus on different parts of the input data through the self-attention mechanism for effective feature restoration; The final output of the decoder is the result after prediction and regression processing.

4. The sea-air coupling data sampling method according to claim 1, wherein (2) The S5 includes: S5-1: Training sample construction, generating input features and target outputs using sampled data: X = {X t | t = 1, 2, ..., T}, Y = {V ocean (t + 1)} Where X is the set of input features and Y is the target output; S5-2: Model training, realizing the prediction of ocean state variables by optimizing the loss function to minimize the prediction error. The optimization formula is as follows: Among them, L is the loss function, Y i is the true value, is the model prediction value, and N is the total number of samples; S5-3: Result verification: Using the test dataset to evaluate the prediction accuracy of the model and its ability to depict the air-sea coupling relationship.

Citation Information

Patent Citations

  • Grid processing method for improving efficiency and stability of ocean numerical forecasting model

    CN118886368A

  • Wave height prediction system and method for deep and far sea intelligent culture platform

    CN119005009A