Four-dimensional variation priori knowledge-based end-to-end data-driven weather forecasting method
Through the end-to-end data-driven method of four-dimensional variation prior knowledge, ERA5 reanalysis data and observation data fine-tune the weather forecast model, solving the problems of low accuracy and poor scalability of existing models, and achieving high-precision weather prediction and model expansion.
Patent Information
- Application Number
- CN202510788571.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The existing artificial intelligence-based weather forecasting model has low accuracy and poor scalability. It depends on the data assimilation module of the expensive numerical weather forecasting system, making it difficult to realize a completely end-to-end business-oriented weather forecasting system.
The end-to-end data-driven weather forecasting method based on four-dimensional variation prior knowledge is adopted. By obtaining the target background field, ERA5 reanalysis data, conventional and satellite observation data, a medium-term weather forecast model is constructed, and fine-tuning is performed through the four-dimensional variation cost function and data assimilation model to generate weather forecast results.
It improves the accuracy of weather forecasts and the scalability of the model, can assimilate conventional and satellite observation data, and enhances the accuracy and adaptability of weather forecasts.
Smart Images

Figure CN120294879A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of weather forecasting, and in particular, to an end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge. Background Art
[0002] In this era of the rapid development of artificial intelligence, end-to-end systems based on artificial intelligence technology have been developed, which can directly predict future weather conditions according to observation data. It is worth noting that some models have successfully generated forecasts directly from raw observation data, showing good performance. However, the natural sparsity of observation data poses a major challenge to achieving forecast accuracy comparable to that of operational numerical weather prediction (NWP) systems. In addition, directly taking all observation data as input without strict quality control may affect the forecast accuracy.
[0003] Recent research has also explored integrating pre-trained artificial intelligence forecasting models into traditional four-dimensional variational (4DVar) frameworks to create hybrid 4DVar systems using automatic differentiation. However, these efforts often rely on idealized models or simulated observations. Even when using real-world conventional observation data, these methods still ignore satellite observation data that is crucial for operational NWP systems. On the other hand, some researchers have also ambitiously pursued the development of end-to-end data assimilation (DA) systems, such as FengWu-Adas, FuXi-DA, DiffDA, and 4DVarFormer. These models use modern neural network architectures to absorb conventional and satellite observation data and combine them with artificial intelligence-based weather forecasting models to create accurate end-to-end forecasting systems. Despite these advancements, these methods still cannot achieve performance comparable to that of operational DA systems. In addition, they can only absorb the types of observations encountered during training and need to retrain the entire model to absorb new types of observations.
[0004] Therefore, existing artificial intelligence-based models have relatively low accuracy in weather prediction and rely on expensive data assimilation (DA) modules of numerical weather prediction (NWP) systems to construct the initial field, which limits the operationalization of fully end-to-end weather forecasting systems, i.e., the scalability is relatively poor. Summary of the Invention
[0005] This application aims to propose an end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge, which can improve the accuracy of weather prediction and has strong scalability.
[0006] An embodiment of this application provides an end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge, and the method includes: Obtain a target background field, a target forecast lead time, ERA5 reanalysis data, first conventional observation data, first satellite observation data, second conventional observation data, and second satellite observation data, where the conventional observation data is the observation data in the Global Data Assimilation System, and the satellite observation data includes the satellite observation data obtained by the Advanced Microwave Sounding Unit and the satellite observation data obtained by the Microwave Humidity Sensor; Pre-train the constructed weather forecasting model into a medium-term weather forecasting model using the ERA5 reanalysis data, and obtain the prediction result of the medium-term weather forecasting model; Fine-tune the medium-term weather forecasting model into a first observation operator using the first satellite observation data; Input the first observation operator, the prediction result, the first conventional observation data, and the first satellite observation data into a four-dimensional variational cost function to obtain the first gradient of the four-dimensional variational cost function; Use the prediction result as the first background field, and use the first gradient and the first background field to fine-tune the weather forecasting model into a data assimilation model; Input the second conventional observation data and the second satellite observation data into the data assimilation model, fine-tune the data assimilation model into a second observation operator using the second satellite observation data, and input the second observation operator and the second conventional observation data into the four-dimensional variational cost function to obtain the second gradient of the four-dimensional variational cost function; Input the second gradient and the target background field into the data assimilation model for fine-tuning to obtain the analysis field output by the fine-tuned data assimilation model; Input the analysis field and the target forecast lead time into the fine-tuned data assimilation model to obtain the target weather prediction result.
[0007] In some embodiments, the pre-training of the constructed weather forecasting model into a medium-term weather forecasting model using the ERA5 reanalysis data includes: Construct a pre-training model loss function as: ; where, represents the pre-training model loss function, represents the expected value, represents the total number of forecast steps, represents the total number of variables, represents the number of latitude coordinates, represents the number of longitude coordinates, represents the variable 's pressure weighting, represents the latitude weighting factor, represents the variable At time and the difference between the predicted state and the training label at the longitude, latitude and coordinate wherein, represents the forecast step number, represents the F1 norm; Based on the pre-training model loss function, the constructed weather forecasting model is pre-trained into a medium-term weather forecasting model by using the ERA5 reanalysis data.
[0008] In some embodiments, the step of fine-tuning the medium-term weather forecasting model into a first observation operator by using the first satellite observation data includes: Constructing a fine-tuning operator loss function as: ; wherein, represents the fine-tuning operator loss function, represents the expected value, represents the channel of satellite observation, represents the calculated observation operator of the channel, represents the state at time , represents the auxiliary observation information at the time, represents the first satellite observation data of the channel, represents the F1 norm; Based on the fine-tuning operator loss function, the medium-term weather forecasting model is fine-tuned into a first observation operator by using the first satellite observation data.
[0009] In some embodiments, before inputting the first observation operator, the prediction result, the first conventional observation data and the first satellite observation data into the four-dimensional variational cost function, the method further includes: Constructing a four-dimensional variational cost function: ; wherein, represents the four-dimensional variational cost function, represents the background field, represents the initial field to be optimized, represents the background error covariance matrix, represents the inverse transformation of the matrix , represents the total time, represents the time observation value of, represents the prediction result, represents the observation operator, represents the observation error covariance matrix, Denotes the inverse transformation of the matrix .
[0010] In some embodiments, the step of fine-tuning the weather forecasting model into a data assimilation model by using the first gradient and the first background field includes: Constructing an objective loss function: ; Wherein, denotes the objective loss function, denotes the expected value, denotes the total number of variables, denotes the number of latitude coordinates, denotes the number of longitude coordinates, denotes the variable weighted by pressure, denotes the latitude weighting factor, denotes the variable at time and the latitude and longitude coordinates at the initial field state, denotes the variable at time and the latitude and longitude coordinates at the state, denotes the F1 norm; Based on the objective loss function, using the first gradient and the first background field, fine-tuning the weather forecasting model into a data assimilation model.
[0011] In some embodiments, the step of pre-training the constructed weather forecasting model into a medium-term weather forecasting model by using the ERA5 reanalysis data includes: Constructing a conditional hybrid neural operator including a conditional feedforward network, a cross-attention mechanism, layer normalization, an attention mechanism, and a feedforward network; Constructing a hybrid neural operator including layer normalization, an attention mechanism, and a feedforward network; Constructing a weather forecasting model according to the conditional hybrid neural operator and the hybrid neural operator; Pre-training the constructed weather forecasting model into a medium-term weather forecasting model by using the ERA5 reanalysis data.
[0012] In some embodiments, the step of constructing a weather forecasting model according to the conditional hybrid neural operator and the hybrid neural operator includes: Constructing an encoder based on the conditional hybrid neural operator; Constructing a decoder based on the hybrid neural operator; Construct a weather forecasting model including the encoder, the decoder, patch embedding, a feature fusion module, and a fully connected layer, where the feature fusion module includes a multi-layer perceptron.
[0013] In some embodiments, constructing a conditional hybrid neural operator including a conditional feed-forward network, a cross-attention mechanism, layer normalization, an attention mechanism, and a feed-forward network includes: Input a conditional input item into the conditional feed-forward network to obtain an output result of the conditional feed-forward network; Input the output result of the conditional feed-forward network into layer normalization to obtain a first layer normalization result; Input the input item into layer normalization to obtain a second layer normalization result; Input the second layer normalization result into the attention mechanism to obtain an attention result; Perform a residual connection between the attention result and the input item to obtain a first addition result; Input the first addition result into layer normalization to obtain a third layer normalization result; Input the first layer normalization result and the third layer normalization result into the cross-attention mechanism to obtain a cross-attention result; Input the cross-attention result and the first addition result into layer normalization to obtain a fourth layer normalization result; Input the fourth layer normalization result into the feed-forward network to obtain an output result of the feed-forward network; Perform a residual connection between the output result of the feed-forward network and the cross-attention result to obtain an output result of the conditional hybrid neural operator, thereby constructing the conditional hybrid neural operator.
[0014] In some embodiments, constructing a hybrid neural operator including layer normalization, an attention mechanism, and a feed-forward network includes: Input an input item into layer normalization to obtain a first layer normalization result; Input the first layer normalization result into the attention mechanism to obtain an attention result; Perform a residual connection between the attention result and the input item to obtain a first addition result; Input the first addition result into layer normalization to obtain a second layer normalization result; Input the second layer normalization result into the feed-forward network to obtain an output result of the feed-forward network; Perform a residual connection between the output result of the feed-forward network and the first addition result to obtain an output result of the hybrid neural operator, thereby constructing the hybrid neural operator.
[0015] In some embodiments, the attention mechanism includes fast Fourier transform, inverse fast Fourier transform, frequency-domain convolution, and spatial-domain convolution. The step of inputting the first-layer normalization result into the attention mechanism to obtain an attention result includes: Processing the first-layer normalization result through the fast Fourier transform to obtain a first processing result; Inputting the first processing result into the frequency-domain convolution to obtain a first convolution result; Processing the first convolution result through the inverse fast Fourier transform to obtain a second processing result; Inputting the first-layer normalization result into the spatial-domain convolution to obtain a second convolution result; Performing residual connection on the second processing result, the second convolution result, and the first-layer normalization result to obtain an attention result.
[0016] Compared with the prior art, the present application has the following beneficial effects: In this method, the constructed weather forecasting model is pre-trained into a medium-term weather forecasting model by using ERA5 reanalysis data, and the prediction result of the medium-term weather forecasting model is obtained; the first satellite observation data is used to fine-tune the medium-term weather forecasting model into a first observation operator; the first observation operator, the prediction result, the first conventional observation data, and the first satellite observation data are input into the four-dimensional variational cost function to obtain the first gradient of the four-dimensional variational cost function; the prediction result is used as the first background field, and the weather forecasting model is fine-tuned into a data assimilation model by using the first gradient and the first background field; the second conventional observation data and the second satellite observation data are input into the data assimilation model, and the data assimilation model is fine-tuned into a second observation operator by the second satellite observation data, and the second observation operator and the second conventional observation data are input into the four-dimensional variational cost function to obtain the second gradient of the four-dimensional variational cost function; the second gradient and the target background field are input into the data assimilation model for fine-tuning to obtain the analysis field output by the fine-tuned data assimilation model; the analysis field and the target forecast lead time are input into the fine-tuned data assimilation model to obtain the target weather prediction result. In this way, by pre-training the weather forecasting model into a medium-term weather forecasting model, then fine-tuning it into an observation operator and a data assimilation model, and also fine-tuning the observation operator and the data assimilation model during weather prediction, and fine-tuning by combining four-dimensional variational prior knowledge, the accuracy of the fine-tuned model for weather prediction can be improved. This method can not only assimilate conventional observation data but also assimilate satellite observation data, enhancing the scalability of the weather forecasting model. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where: Figure 1It is a schematic flowchart of an embodiment of the end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge provided by this application; Figure 2 It is a schematic diagram of the Xichen system in the best embodiment of the end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge provided by this application; Figure 3 It is a schematic flowchart of the end-to-end weather forecasting system constructed by Xichen in the best embodiment of the end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge provided by this application; Figure 4 It is a schematic diagram of the Xichen model architecture in the best embodiment of the end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge provided by this application; Figure 5 It is a schematic diagram of calculating the 4DVar cost function during the process of assimilating conventional observation data and satellite observation data in the best embodiment of the end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge provided by this application. Detailed implementation manners
[0018] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application.
[0019] In the description of the present application, if the first, second, etc. are described, they are only for the purpose of distinguishing technical features and should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0020] In the description of the present application, it should be understood that for the orientation description, such as up, down, etc., the indicated orientation or positional relationship is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the indicated device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application.
[0021] In the description of the present application, it should be noted that unless otherwise clearly defined, words such as setting, installation, connection, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present application in combination with the specific content of the technical solution.
[0022] Since the existing artificial intelligence-based models have relatively low accuracy in weather forecasting and rely on expensive data assimilation (DA) modules of numerical weather prediction (NWP) systems to construct the initial field, this limits the operationalization of fully end-to-end weather forecasting systems, i.e., their scalability is relatively poor.
[0023] To solve the problems of relatively low accuracy in weather forecasting and relatively poor model scalability existing in the prior art, this application proposes an end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge.
[0024] Refer to Figure 1 , which is a schematic flowchart of the end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge provided by an embodiment of this application. The end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge is applied to an electronic device, which can be a server, a mobile terminal, etc. As Figure 1 shown, the end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge may include the following steps: Step S100: Obtain a target background field, a target forecast lead time, ERA5 reanalysis data, first conventional observation data, first satellite observation data, second conventional observation data, and second satellite observation data. Among them, the conventional observation data are the observation data in the global data assimilation system, and the satellite observation data include the satellite observation data obtained by the advanced microwave sounding unit and the satellite observation data obtained by the microwave humidity sensor; Step S200: Use the ERA5 reanalysis data to pre-train the constructed weather forecasting model into a medium-term weather forecasting model and obtain the prediction result of the medium-term weather forecasting model; Step S300: Use the first satellite observation data to fine-tune the medium-term weather forecasting model into a first observation operator; Step S400: Input the first observation operator, the prediction result, the first conventional observation data, and the first satellite observation data into the four-dimensional variational cost function to obtain the first gradient of the four-dimensional variational cost function; Step S500: Use the prediction result as the first background field, and use the first gradient and the first background field to fine-tune the weather forecasting model into a data assimilation model; Step S600: Input the second conventional observation data and the second satellite observation data into the data assimilation model, fine-tune the data assimilation model into a second observation operator through the second satellite observation data, and input the second observation operator and the second conventional observation data into the four-dimensional variational cost function to obtain the second gradient of the four-dimensional variational cost function; Step S700: Input the second gradient and the target background field into the data assimilation model for fine-tuning to obtain the analysis field output by the fine-tuned data assimilation model; Step S800: Input the analysis field and the target forecast lead time into the fine-tuned data assimilation model to obtain the target weather prediction result.
[0025] In this embodiment, the constructed weather forecasting model is pre-trained into a medium-term weather forecasting model by using ERA5 reanalysis data, and the prediction result of the medium-term weather forecasting model is obtained; the first satellite observation data is used to fine-tune the medium-term weather forecasting model into the first observation operator; the first observation operator, the prediction result, the first conventional observation data, and the first satellite observation data are input into the four-dimensional variational cost function to obtain the first gradient of the four-dimensional variational cost function; the prediction result is used as the first background field, and the weather forecasting model is fine-tuned into a data assimilation model by using the first gradient and the first background field; the second conventional observation data and the second satellite observation data are input into the data assimilation model, and the data assimilation model is fine-tuned into the second observation operator by the second satellite observation data, and the second observation operator and the second conventional observation data are input into the four-dimensional variational cost function to obtain the second gradient of the four-dimensional variational cost function; the second gradient and the target background field are input into the data assimilation model for fine-tuning to obtain the analysis field output by the fine-tuned data assimilation model; the analysis field and the target forecast lead time are input into the fine-tuned data assimilation model to obtain the target weather prediction result. In this way, by pre-training the weather forecasting model into a medium-term weather forecasting model, then fine-tuning it into an observation operator and a data assimilation model, and also fine-tuning the observation operator and the data assimilation model during weather prediction, and fine-tuning by combining four-dimensional variational prior knowledge, the accuracy of the fine-tuned model for weather prediction can be improved. This method can not only assimilate conventional observation data but also assimilate satellite observation data, enhancing the scalability of the weather forecasting model.
[0026] The above constructed weather forecasting model can be a weather forecasting model constructed based on an artificial intelligence model well-known to those skilled in the art, or a weather forecasting model constructed by using a conditional hybrid neural operator and a hybrid neural operator.
[0027] In some embodiments, pre-training the constructed weather forecasting model into a medium-term weather forecasting model by using ERA5 reanalysis data includes: Construct the pre-training model loss function as: ; where represents the pre-training model loss function, represents the expected value, represents the total number of forecast steps, represents the total number of variables, represents the number of latitude coordinates, represents the number of longitude coordinates, represents the variable 's pressure weighting, represents the latitude weighting factor, represents the variable at time and the difference between the predicted state and the training label at the longitude and latitude coordinates The number of prediction steps is represented by represents the forecast step; represents the F1 norm; Based on the pre-training model loss function, the constructed weather forecasting model is pre-trained into a medium-term weather forecasting model using ERA5 reanalysis data.
[0028] In this embodiment, based on the pre-training model loss function, the constructed weather forecasting model is pre-trained into a medium-term weather forecasting model using ERA5 reanalysis data. This loss function adopts a training strategy of multi-step prediction, that is, the model is executed times per batch, which can reduce the cumulative error in the prediction process and thus improve the accuracy of the prediction results of the medium-term weather forecasting model.
[0029] In some embodiments, the medium-term weather forecasting model is fine-tuned into a first observation operator using the first satellite observation data, including: Construct the fine-tuning operator loss function as: ; where represents the fine-tuning operator loss function, represents the expected value, represents the channel of satellite observation, represents the calculated observation operator of the channel, represents the state at time , represents the auxiliary observation information at time, represents the first satellite observation data of the channel, represents the F1 norm; Based on the fine-tuning operator loss function, the medium-term weather forecasting model is fine-tuned into a first observation operator using the first satellite observation data.
[0030] In this embodiment, by fine-tuning the medium-term weather forecasting model into a first observation operator based on the fine-tuning operator loss function and using the first satellite observation data, it helps to develop scalable plugins, that is, to assimilate the original satellite observation data.
[0031] In some embodiments, before inputting the first observation operator, the prediction result, the first conventional observation data, and the first satellite observation data into the four-dimensional variational cost function, the method further includes: Construct the four-dimensional variational cost function: ; Among them, represents the four-dimensional variational cost function, represents the background field, represents the initial field to be optimized, represents the background error covariance matrix, represents the inverse transformation of the matrix , represents the total time, represents the time observation value of, represents the prediction result, represents the observation operator, represents the observation error covariance matrix, represents the inverse transformation of the matrix .
[0032] In this embodiment, the constructed four-dimensional variational cost function can also seamlessly integrate various observation operators to promote the assimilation of satellite raw observation data.
[0033] In some embodiments, the weather forecasting model is fine-tuned to a data assimilation model by using the first gradient and the first background field, including: Construct a target loss function: ; Among them, represents the target loss function, represents the expected value, represents the total number of variables, represents the number of latitude coordinates, represents the number of longitude coordinates, represents the variable pressure weighting of, represents the latitude weighting factor, represents the variable at time and the latitude and longitude coordinates initial field state at, represents the variable at time and the latitude and longitude coordinates state at, represents the F1 norm; Based on the target loss function, the weather forecasting model is fine-tuned to a data assimilation model by using the first gradient and the first background field.
[0034] In this embodiment, by minimizing the target loss function, the weather forecasting model is fine-tuned into a data assimilation model, enabling the fine-tuned data assimilation model to improve the accuracy of weather prediction. It can not only assimilate conventional observation data but also satellite observation data, enhancing the scalability of the weather forecasting model.
[0035] In some embodiments, the pre-trained weather forecasting model is pre-trained into a medium-term weather forecasting model using ERA5 reanalysis data, including: Construct a conditional hybrid neural operator including a conditional feed-forward network, a cross-attention mechanism, layer normalization, an attention mechanism, and a feed-forward network; Construct a hybrid neural operator including layer normalization, an attention mechanism, and a feed-forward network; Construct a weather forecasting model based on the conditional hybrid neural operator and the hybrid neural operator; Pre-train the constructed weather forecasting model into a medium-term weather forecasting model using ERA5 reanalysis data.
[0036] In this embodiment, by constructing a conditional hybrid neural operator including a conditional feed-forward network, a cross-attention mechanism, layer normalization, an attention mechanism, and a feed-forward network; constructing a hybrid neural operator including layer normalization, an attention mechanism, and a feed-forward network; constructing a weather forecasting model based on the conditional hybrid neural operator and the hybrid neural operator; and pre-training the constructed weather forecasting model into a medium-term weather forecasting model using ERA5 reanalysis data. In this way, the information from conditional tokens and input tokens can be fully integrated through the conditional hybrid neural operator, enabling the weather forecasting model constructed based on the conditional hybrid neural operator and the hybrid neural operator to better learn the feature information in the input data, thereby improving the accuracy of the prediction results of the weather forecasting model.
[0037] In some embodiments, constructing a weather forecasting model based on the conditional hybrid neural operator and the hybrid neural operator includes: Construct an encoder based on the conditional hybrid neural operator; Construct a decoder based on the hybrid neural operator; Construct a weather forecasting model including an encoder, a decoder, patch embedding, a feature fusion module, and a fully connected layer, where the feature fusion module includes a multi-layer perceptron.
[0038] In this embodiment, a weather forecasting model including an encoder, a decoder, patch embedding, a feature fusion module, and a fully connected layer is constructed. Among them, the feature fusion module includes a multi-layer perceptron. The main purpose of patch embedding is to reduce the spatial dimension of the input data, thereby reducing redundancy and accelerating the training process. And based on the conditional hybrid neural operator and the hybrid neural operator, the constructed weather forecasting model can better learn the feature information in the input data, thereby improving the accuracy of the prediction results of the weather forecasting model.
[0039] In some embodiments, constructing a conditional hybrid neural operator including a conditional feedforward network, a cross-attention mechanism, layer normalization, an attention mechanism, and a feedforward network includes: Input the conditional input item into the conditional feedforward network to obtain the output result of the conditional feedforward network; Input the output result of the conditional feedforward network into layer normalization to obtain the first layer normalization result; Input the input item into layer normalization to obtain the second layer normalization result; Input the second layer normalization result into the attention mechanism to obtain the attention result; Perform a residual connection on the attention result and the input item to obtain the first addition result; Input the first addition result into layer normalization to obtain the third layer normalization result; Input the first layer normalization result and the third layer normalization result into the cross-attention mechanism to obtain the cross-attention result; Input the cross-attention result and the first addition result into layer normalization to obtain the fourth layer normalization result; Input the fourth layer normalization result into the feedforward network to obtain the output result of the feedforward network; Perform a residual connection on the output result of the feedforward network and the cross-attention result to obtain the output result of the conditional hybrid neural operator, so as to construct the conditional hybrid neural operator.
[0040] In this embodiment, by constructing a conditional hybrid neural operator including a conditional feedforward network, a cross-attention mechanism, layer normalization, an attention mechanism, and a feedforward network, the information from the conditional tokens and the input tokens can be fully integrated through the conditional hybrid neural operator, thereby improving the accuracy of the model prediction results.
[0041] In some embodiments, constructing a hybrid neural operator including layer normalization, an attention mechanism, and a feedforward network includes: Input the input item into layer normalization to obtain the first layer normalization result; Input the first layer normalization result into the attention mechanism to obtain the attention result; Perform a residual connection on the attention result and the input item to obtain the first addition result; Input the first addition result into layer normalization to obtain the second layer normalization result; Input the second layer normalization result into the feed-forward network to obtain the feed-forward network output result; Perform a residual connection between the feed-forward network output result and the first addition result to obtain the hybrid neural operator output result, thereby constructing a good hybrid neural operator.
[0042] In this embodiment, the attention mechanism can enhance the local feature extraction ability, lay a good foundation for constructing a weather forecasting model, and thus improve the accuracy of the model prediction result.
[0043] In some embodiments, the attention mechanism includes fast Fourier transform, inverse fast Fourier transform, frequency domain convolution, and spatial domain convolution. Input the first layer normalization result into the attention mechanism to obtain the attention result, including: Process the first layer normalization result through fast Fourier transform to obtain the first processing result; Input the first processing result into the frequency domain convolution to obtain the first convolution result; Process the first convolution result through inverse fast Fourier transform to obtain the second processing result; Input the first layer normalization result into the spatial domain convolution to obtain the second convolution result; Perform a residual connection on the second processing result, the second convolution result, and the first layer normalization result to obtain the attention result.
[0044] In this embodiment, by combining frequency domain convolution and spatial domain convolution, the attention mechanism can enhance the local feature extraction ability, lay a good foundation for constructing a weather forecasting model, and thus improve the accuracy of the model prediction result.
[0045] For the convenience of those skilled in the art to understand, the following provides a set of best embodiments: A century ago, weather forecasting was conceptualized as an initial value problem, that is, starting from the initial conditions estimated by observations and applying physical laws to obtain forecasts. At that time, atmospheric observations were extremely limited, computers did not exist, and the predictability of the weather was still largely unknown. Nowadays, the continuous progress of people's understanding of atmospheric mechanisms, the continuous accumulation of observational data, and the significant improvement of computing power have all together improved the accuracy of numerical weather prediction (NWP) and made it an indispensable part of daily life.
[0046] Numerical weather prediction consists of two core parts. The first part is data assimilation (DA), which estimates the most likely state of the atmosphere based on the latest observational data from satellites, weather stations, ships, and other sources. The second part is the forecast model, which traditionally simulates the spatio-temporal evolution of atmospheric variables by solving a set of complex partial differential equations (PDEs). These two parts are essentially interdependent. DA relies on the forecast model to provide the background field and combines observational data to estimate the initial field for the next forecast. Although near-earth astronomical forecasting has undergone a technological revolution, some challenges have emerged. The slowdown of Moore's Law, the increasing complexity of modern NWP systems, and the significant increase in resolution have made it very difficult to provide forecast results in a timely manner within limited computing time. Therefore, these limitations have hindered the further development of NWP.
[0047] In recent years, the rapid development of artificial intelligence (AI) technology has triggered a transformative revolution in the field of weather forecasting. On dedicated computing hardware such as graphics processing units (GPUs), AI-based models, such as FourCastNet, Pangu-Weather, GraphCast, FengWu, Fuxi, and AIFS, have achieved an acceleration three to four orders of magnitude faster than operational NWP systems, while maintaining comparable prediction performance. However, these AI models still rely on the (re)analysis fields generated by NWP systems to produce forecasts. For example, the current AI-based forecast models (such as AIFS of ECMWF) use the analysis fields that are products generated by the four-dimensional variational (4DVar) method. The 4DVar method adopts the optimal control theory and assimilates more than 10 million observational data points within each 12-hour data assimilation window (DAW). It is worth noting that this method effectively integrates indirect satellite observational data, filling the information gap of the atmospheric field that is difficult to capture due to sparse traditional observational data. Nevertheless, the computational cost of 4DVar is very high, limiting the number of observational data that can be assimilated within a limited time, and only 5% to 10% of the available satellite data can be used. In addition, 4DVar requires meticulous adjustment and precise formulation of the complex background error covariance matrix. It also depends on complex observation operators, tangent models, and adjoint models, all of which require a large amount of development time and expertise from trained experts.
[0048] In this era of the rapid development of artificial intelligence, a bold claim has emerged: to develop an end-to-end system based on artificial intelligence technology to directly predict future weather states from observational data. It is worth noting that models such as Aardvark and GraphDOP have successfully generated predictions directly from raw observational data, showing good performance. However, the natural sparsity of observational data poses a significant challenge to achieving prediction accuracy comparable to that of operational NWP systems. In addition, directly taking all observational data as input without strict quality control may affect prediction accuracy. Recent research has also explored integrating pre-trained artificial intelligence prediction models into traditional 4DVar frameworks to create hybrid 4DVar systems using automatic differentiation. However, these efforts often rely on idealized models or simulated observations. Even when using real-world conventional observational data, these methods still ignore satellite observational data that is crucial for operational NWP systems. On the other hand, some researchers have also ambitiously pursued the development of end-to-end DA systems, such as FengWu-Adas, FuXi-DA, DiffDA, and 4DVarFormer. These models use modern neural network architectures to absorb conventional and satellite observational data and combine them with artificial intelligence-based weather forecasting models to create accurate end-to-end prediction systems. Despite these advances, these methods still cannot achieve performance comparable to that of operational DA systems. In addition, they can only absorb the types of observations encountered during training and require the entire model to be retrained to absorb new types of observations. Compared with traditional data assimilation systems, this limitation restricts their scalability. Therefore, developing an end-to-end weather forecasting system suitable for operational implementation remains an unsolved challenge.
[0049] Therefore, this embodiment proposes a Xichen model (i.e., a weather forecasting model) that integrates 4DVar prior knowledge for end-to-end weather forecasting. The Xichen model was initially pre-trained on ERA5 data to learn basic atmospheric dynamics representations as a medium-range weather forecasting model. Subsequently, the medium-range weather forecasting model was fine-tuned to transfer to the tasks of observation operators and data assimilation (i.e., using four-dimensional variational assimilation). This embodiment proposes a cascaded sequential data assimilation framework to assimilate conventional observational data from the Global Data Assimilation System (GDAS) (i.e., GDAS prepbufr observational data) and raw satellite observational data from the Advanced Microwave Sounding Unit (i.e., AMSU-A observational data) and raw satellite observational data from the Microwave Humidity Sensor (MHS) (i.e., MHS observational data). By combining 4DVar prior knowledge, the Xichen model can estimate the analysis field with high quality, thus initializing accurate medium-range forecasts.
[0050] This framework theoretically addresses the compatibility and robustness challenges associated with integrating more indirect satellite observations in the future. Therefore, the Xichen model is expected to completely replace the traditional NWP system workflow and operate independently using only observational data. Although Xichen currently assimilates only GDAS prepbufr, AMSU-A, and MHS observational data, it can also assimilate other satellite observational data, which is not specifically limited in this embodiment, and operates on a global grid with a resolution of only 1.40625 degrees. However, its accuracy on multiple variables in DA and medium-range weather forecasting tasks is comparable to that of operational NWP systems. It should be noted that this embodiment can be not limited to a global grid with a resolution of 1.40625 degrees and can also be used for other resolutions, and this embodiment does not specifically limit the resolution. The Xichen model achieves a forecast lead time of 8.5 days, exceeding that of NOAA's GFS operational system, Aardvark, and GraphDOP.
[0051] The technical solution of this embodiment specifically includes the following content: 1. Symbols and problem description.
[0052] This embodiment represents the complete state of the atmosphere at a specific time as a tensor with dimensions , where represents the total number of variables, represents the number of latitude coordinates, represents the number of longitude coordinates. The index tensor refers to the state of variable at time and latitude and longitude coordinates . In addition, this embodiment also uses the tensor to represent the traditional atmospheric observational data of GDAS prepbufr at . For the unobserved variable at coordinates , this embodiment sets , where NAN represents non-numeric, and for the observed variable , , where represents the observation error. For satellite observations, they are represented as , where is the satellite observation error, and is the observation operator that maps atmospheric variables to satellite observation variables.
[0053] The objective of this embodiment is to develop an end-to-end weather forecasting system based on artificial intelligence that integrates a weather forecasting model and a data assimilation model. For the weather forecasting model, given the system state , the objective of this embodiment is to predict the future time status To achieve this goal, a weather prediction model is trained to map the current atmospheric state to a predicted state at a future time status Regarding the data assimilation model, given the background field at a given time field , GDASprepbufr observation data , and satellite observation data During DAW , the goal of this embodiment is to learn to estimate the optimal initial field by fusing these data . To assimilate satellite observation data, this embodiment needs to learn an observation operator , where represents the auxiliary observation information at time . Subsequently, this embodiment needs to learn a DA model , where reanalysis is used as a reference
[0054] 2. Dataset introduction
[0055] In this study, this embodiment uses the DABench dataset and combines AMSU-A and MHS satellite observation data to train and evaluate the potential of the Xichen model in satellite assimilation tasks. DABench includes 14-year ERA5 reanalysis data from 2010 to 2023 and GDAS prepbufr observation data, and the grid resolution of these data is 1.40625° ( ). This embodiment studies 5 atmospheric variables (each variable has 9 pressure levels) and 4 surface variables, for a total of 49 variables. The atmospheric variables include geopotential (Z), temperature (T), specific humidity (Q), the longitudinal component of wind (U), and the meridional component of wind (V). The 9 sub-variables at different vertical levels are represented by their abbreviations and respective pressure levels (for example, Z500 represents the geopotential at 500 hPa). The 4 surface variables include 2-meter air temperature (T2M), 10-meter wind zonal component (U10), 10-meter wind meridional component (V10), and mean sea level pressure (MSLP).
[0056] For the observation data of AMSU-A and MHS, the nearest neighbor interpolation method is used to map the original satellite data and auxiliary information to a resolution of 1.40625°. In this study, this embodiment excludes the observation data in the regions above 60° north and south latitudes to avoid complex problems caused by sea ice. In addition, this embodiment also adopts a direct quality control procedure to eliminate the data with brightness temperature values exceeding 350 K or lower than 150 K
[0057] All models were trained on data collected before 2021, and 2022 was designated as the validation set. Subsequently, the data from 2023 was used for testing.
[0058] 3. Build an end-to-end weather forecasting system.
[0059] The Xichen model (i.e., the weather forecasting model) is based on the Conditional Hybrid Neural Operator (CHNO) proposed in this embodiment. This operator integrates cross-attention and conditional information into the Hybrid Neural Operator (HNO) of this embodiment. This design achieves a unified model architecture that can adapt to various downstream tasks. In addition, since the Adaptive Fourier Neural Operator (AFNO) can only perform global feature extraction through convolution operations in the Fourier space, a convolution-based branch is introduced in this embodiment to enhance the local feature extraction ability, thus forming the HNO.
[0060] The end-to-end weather forecasting system of this embodiment uses the Xichen model to cover the most important workflows in NWP, from the observation operator, DA model to the weather forecasting model. Figure 2 Figure (a) is a schematic diagram of the overall framework of the Xichen model of the end-to-end weather forecasting system. The Xichen model integrates background field data, conventional observation data, and satellite observation data within the assimilation window to perform the DA task. This process first calculates the 4DVar cost function and then determines the gradient relative to the background field. By combining this gradient with the background field, Xichen can generate the analysis field. Using the analysis field and the forecast step as inputs, Xichen will generate the forecast field or update the background field. Among them, the data assimilation task of this embodiment includes sequentially using data assimilation models for different observation data. For example, in this embodiment, the Xichen model first assimilates GDAS prepbufr data, then assimilates AMSU-A observation data, and finally assimilates MHS observation data. This configuration enhances the scalability of the Xichen model. When new observation data is introduced in the future, only the respective observation operators and DA model components need to be fine-tuned, without retraining the entire DA system.
[0061] 4. Pretrain the medium-term weather forecasting model.
[0062] The goal of this embodiment is to design a base model (i.e., the Xichen model) that can be pre-trained on reanalysis data and then fine-tuned to solve various downstream weather tasks, thus building an end-to-end weather forecasting system. Figure 2 Figure (b) is a schematic diagram of the pre-training of the medium-term weather forecasting model. During the pre-training process, the Xichen model takes the reanalysis field as the input and the lead time Output the predicted state based on the condition In this embodiment, this embodiment extracts from the hourly set . This process enables the Xichen model to effectively capture the relationships between various weather variables, taking into account the dynamic and thermodynamic characteristics of the atmosphere. To generate predictions for any time increment, the hierarchical time aggregation method proposed by Pangu-Weather can be used to merge the predictions of the learned forecasting model. Therefore, the Xichen model is pre-trained to minimize the following losses: ; where represents the F1 norm, represents the pressure weighting of the variable , is the difference between the predicted state and the training label (i.e., the true weather state) at time and the latitude and longitude coordinates , is the latitude weighting factor: ; where is the latitude of the th row of the grid. This term is commonly used in the training process of previous artificial intelligence-based weather forecasting models.
[0063] To reduce the cumulative error during the prediction process, this embodiment adopts a multi-step prediction training strategy to fine-tune Xichen. Specifically, this embodiment executes the model times per batch and calculates the average loss of these step predictions to optimize the model parameters. The loss function for multi-step prediction (i.e., the pre-trained model loss function) is as follows: ; where represents the number of forecast steps, .
[0064] In practice, this embodiment first fine-tunes the pre-trained model, ranging from to , to establish the Xichen-Short-Term model. Then, the Xichen-Medium-Term model is initialized with the weights from the Xichen-Short-Term model, and this model is then fine-tuned to achieve the best prediction performance during 5 to 10 forecast steps. This embodiment uses the same in all Sampled values. The prediction model of this embodiment is similar to FuXi during the prediction process. Since the analysis field error provided by the Xichen model is comparable to the error of the 2-day forecast initialized with ERA5 for short-term Xichen, the 10-day medium-range weather forecast generated using the initial field estimated by the Xichen model has the first 3 days generated by Xichen - short-term and the remaining 7 days generated by Xichen - medium-term.
[0065] 5. Fine-tuning of the observation operator.
[0066] To compare the atmospheric state with satellite observations, the observation operator is essential. The simulated radiance at each observation point is calculated through the atmospheric variable curve. Figure 2 In (c) is a schematic diagram of the fine-tuning of the data assimilation model. In this embodiment, the observation operator is constructed conditional on the satellite attitude parameters and scanning positions to fine-tune the Xichen model: , represents the mapping relationship, where represents time the auxiliary observation information at time. For training the loss function (i.e., the fine-tuning operator loss function) is as follows: ; where, represents the channel of satellite observation, and the subscript represents the channel to be calculated. This process helps to develop scalable plugins to absorb the original satellite observation data.
[0067] This embodiment selects channels 5 to 10 of AMSU-A and channels 3 to 5 of MHS to train the observation operator and conduct assimilation experiments. This selection is because channels 5 to 10 of AMSU-A are very sensitive to the atmospheric temperature profile from 50 hPa to 1000 hPa, while channels 3 to 5 of MHS are sensitive to the atmospheric water vapor in the troposphere.
[0068] 6. Fine-tuning of the assimilation model.
[0069] This embodiment makes full use of the atmospheric dynamics information provided by the Xichen model by combining 4D-Var prior knowledge. This method can effectively propagate observations in the spatial and variable dimensions of the atmospheric field. In addition, the 4D-Var cost function (i.e., the four-dimensional variational cost function) can also seamlessly integrate various observation operators to promote the assimilation of the original satellite observation data.
[0070] Figure 2 In (d) is a schematic diagram of the fine-tuning of the data assimilation model. In this embodiment, the gradient of the 4DVar cost function with respect to the background field is expressed as , thereby fine-tuning the Xichen model to make it function as a DA model. The output of the Xichen model is the representation of the analysis increment in the latent space, rather than the analysis increment itself. This representation is added to the embedding of the background field and then decoded by the HNO block to generate the final analysis field (the obtained analysis field can be used as the initial field for weather prediction at the next moment). Here, the background field is the conditional input. Here, represents the 4DVar cost function as follows: ; where represents the initial field to be optimized, represents the background field, represents time of the observation value. The prediction model is represented by , which maps the initial field to the field at time . In addition, and represent the background error covariance matrix and the observation error covariance matrix respectively. The observation covariance matrix of GDAS prepbufr observations is represented as . The observation operator error is used to construct the error covariance matrix of satellite observations. Since only the gradient of the initial value of the 4DVar cost function with respect to needs to be calculated in this embodiment, there is no need to estimate . This simplifies the process and avoids a difficult problem usually encountered in traditional 4DVar methods. In addition, different from the methods used in Aardvark, GraphDOP, FuXi-DA, and FengWu-Adas, a quality control process is added in this embodiment when calculating the 4DVar cost function, that is, the observations whose distance from the background field exceeds 5 times the standard deviation of the variable are excluded. This method can prevent the over-adjustment of local regions in the background field due to the large difference between the observation data and the background.
[0071] Therefore, this embodiment takes as the input, takes the background field as the condition, and fine-tunes the Xichen model to construct a DA model . The training objective of this model is to minimize the following loss function (i.e., the target loss function): .
[0072] In this embodiment, the Xichen model was first fine-tuned using GDAS prepbufr observation data to establish a GDAS prepbufr assimilation model. Subsequently, the Xichen model was further fine-tuned to an AMSU-A assimilation model, and the analysis field generated after assimilating GDAS prepbufr was used as the background field. Finally, the Xichen model was fine-tuned in sequence to assimilate MHS satellite observation data, and the assimilation order of AMSU-A and MHS was random.
[0073] The flowchart of the end-to-end weather forecasting system constructed by Xichen is as Figure 3 shown. Figure 3 In (a), it is a schematic diagram of the cascaded method for assimilating multi-source observations, and the assimilation system is implemented using the cascaded mode method. When executing the assimilation model, Xichen first assimilates GDAS prepbufr data, then assimilates AMSU-A observation data, and finally assimilates MHS observation data. In future implementations, by simply adding the assimilation steps of other satellite observation data to the data assimilation workflow, these observation data can be integrated. In the DA cycle task, the DA mode and the forecasting mode are alternately executed to obtain an assimilation forecasting cycle system. Figure 3 In (b), it is a schematic diagram of the assimilation forecasting process of the Xichen model. The data assimilation cycle starts from the initial background field. The Xichen model first assimilates the observation data to generate an analysis field. Subsequently, a forecast is generated based on the analysis field to generate a background field for subsequent time steps. This iterative process continues to determine the analysis field at each time point, which serves as the initial condition for subsequent medium-term forecasting. Figure 3 In (c), it is a schematic diagram of the medium-term weather forecasting experiment process of the Xichen model. The analysis field obtained from the DA cycle is used as the initial condition for initializing the Xichen forecasting model, and the Xichen forecasting model performs a 10-day medium-term weather forecasting task.
[0074] In the pre-training stage, the Xichen model (i.e., the weather forecasting model) of this embodiment is characterized as a weather forecasting simulator. It takes the current atmospheric state and the forecast lead time as inputs, and then predicts the atmospheric state at the corresponding moment , as defined in the main text: ; To achieve this goal, the model of this embodiment must utilize the information of the forecast step length to effectively guide the prediction process.
[0075] 7. Introduction to the Xichen model structure.
[0076] The structural diagram of the Hybrid Neural Operator (HNO) of the Xichen model (i.e., the weather forecasting model) is as Figure 4As shown in (a), specifically, the hybrid neural operator includes layer normalization, attention mechanisms (including Fast Fourier Transform (FFT), Inverse Fast Fourier Transform (IFFT), frequency-domain convolution (Multiply Frequency), and spatial-domain convolution (Group CON, i.e., a group)), and a feed-forward network. The input item (i.e., input data) is input into layer normalization, and the resulting normalized result is input into the attention mechanism to obtain an attention result; the attention result and the input item are residually connected to obtain a first addition result; the first addition result is input into layer normalization, and the resulting normalized result is input into the feed-forward network to obtain a feed-forward network output result. The feed-forward network output result and the first addition result are residually connected to obtain a hybrid neural operator output result. Among them, the attention mechanism operation includes: processing the layer normalization result through the Fast Fourier Transform (FFT) to obtain a first processing result, inputting the first processing result into the frequency-domain convolution to obtain a first convolution result, then processing the first convolution result through the Inverse Fast Fourier Transform (IFFT) to obtain a second processing result, inputting the layer normalization result into the spatial-domain convolution to obtain a second convolution result, and residually connecting the second processing result, the second convolution result, and the layer normalization result to obtain an attention result.
[0077] Recently, the Adaptive Fourier Neural Operator (AFNO) has shown excellent performance in medium-range weather forecasting tasks, which is attributed to its reduced GPU memory usage and computational cost. FourCastNet is the first forecasting model that can be directly compared with IFS HRES at the resolution level and is developed using AFNO. However, since AFNO, as an implementation of global attention, has limitations in capturing local features, this embodiment introduces an additional convolutional branch that uses circular padding and grouped convolution with a kernel size of 3x3. This combination of AFNO and convolutional blocks is called the hybrid neural operator (HNO).
[0078] Subsequently, the structural diagram of the Conditional Hybrid Neural Operator (CHNO) is shown in Figure 4As shown in Figure (b), specifically, the conditional hybrid neural operator includes a conditional feed-forward network, a cross-attention mechanism, layer normalization, an attention mechanism (including Fast Fourier Transform (FFT), Inverse Fast Fourier Transform (IFFT), frequency-domain convolution (Multiply Frequency), and spatial-domain convolution (GroupCNN)), and a feed-forward network. The conditional input term is input into the conditional feed-forward network to obtain the output result of the conditional feed-forward network; the output result of the conditional feed-forward network is input into layer normalization to obtain the first layer normalization result. The input term is input into layer normalization to obtain the second layer normalization result; the obtained second layer normalization result is input into the attention mechanism to obtain the attention result; the attention result and the input term are concatenated with residuals to obtain the first addition result; the first addition result is input into layer normalization to obtain the third layer normalization result; the first layer normalization result and the third layer normalization result are input into the cross-attention mechanism to obtain the cross-attention result; the cross-attention result and the first addition result are input into layer normalization, and the obtained normalization result is input into the feed-forward network to obtain the output result of the feed-forward network. The output result of the feed-forward network is concatenated with the cross-attention result with residuals to obtain the output result of the conditional hybrid neural operator. It should be noted that the attention mechanism in the conditional hybrid neural operator and the attention mechanism in the hybrid neural operator have the same processing operations.
[0079] Cross-attention is applied to fuse the features of the input tokens and the conditional tokens. The cross-attention mechanism and the output of the HNO are combined using the residual concatenation method to generate the final output of the basic module through subsequent feed-forward processes. In this embodiment, this basic module is called the conditional hybrid neural operator (CHNO).
[0080] The overall architecture diagram of the Xichen model is as Figure 4 shown in Figure (c). It consists of 5 main parts: patch embedding, encoder, feature fusion module, decoder, and fully connected (FC) layer. The input data synthesizes the upper-air and surface variables to form a data cube with dimensions of , where 49, 128, and 256 represent the total number of input variables, latitude ( ), and longitude ( ) grid points, respectively.
[0081] Initially, the high-dimensional input data is reduced in dimension to through patch embedding, where Set to 512, representing the number of channels. The main purpose of patch embedding is to reduce the spatial dimension of the input data, thereby reducing redundancy. Subsequently, the encoder utilizes stacked CHNO layers to condition the conditional tokens on the embedded features, thus fully integrating information from both conditional and input tokens. To optimize the use of shallow and deep features, the outputs of the first three blocks are combined through a feature fusion module consisting of multi-layer perceptrons (MLPs). The fused features are then processed by the decoder to generate predicted features, which are remapped to dimensions through an MLP and a reshaping operation. The predicted features are fed into a fully connected (FC) layer to obtain the weather prediction result.
[0082] Using patch embedding can reduce the spatial dimension of the input and speed up the training process. This method involves splitting the image into patches, and each patch is converted into a feature vector, as used in models such as the Pangu model, Fengwu model, and Fuxi model. Patch embedding applies a two-dimensional convolutional layer with a kernel and stride of (corresponding to ), and generates output channels numbered . After patch embedding, to improve the stability of training, layer normalization (LayerNorm) is performed, resulting in a tensor of dimension .
[0083] The encoder consists of three blocks, each containing four CHNO layers. Each CHNO layer includes an HNO branch and a conditional feed-forward branch. The conditional feed-forward branch consists of an MLP for mapping conditional tokens to hidden features. HNO includes an AFNO branch with 16 heads and a group CON layer (i.e., spatial domain convolution (Group CON), where CON is convolution), characterized by two-dimensional periodic padding, a kernel size of 3*3, and 16 channel groups. The fusion module contains a linear layer and a GELU activation function, which processes the three features generated by the three encoder blocks to produce fused features. The decoder contains 12 HNO layers. Before feeding into the decoder, the features are processed by LayerNorm to improve the stability of training. Finally, the output of the decoder passes through a linear head and is reshaped to dimensions.
[0084] When fine-tuning the Xichen model as an assimilation model, the gradient of the 4DVar cost function is used as the input, while the background field is used as the conditional input. These data are converted into tokens through patch embedding and then fed into CHNO. Figure 5 Shows a schematic diagram of calculating the 4DVar cost function during the assimilation process of conventional observations and satellite observations. Figure 5Figure (a) is a schematic diagram of the calculation process of the 4DVar cost function for assimilating conventional observation data. Figure 5 Figure (b) is a schematic diagram of the calculation process of the 4DVar cost function for assimilating satellite observation data.
[0085] 8. Evaluation Metrics.
[0086] This embodiment aims to evaluate the performance of the model in this embodiment in terms of assimilation and forecasting tasks. Therefore, this embodiment comprehensively evaluates a one-year DA cycle and a 10-day medium-range forecast. The assimilation time interval of the DA cycle is 12 hours, and the DAW is 12 hours. Specifically, the DA cycle runs at 00:00 UTC and 12:00 UTC every day in 2023, which corresponds to the initialization time of the 10-day forecasts conducted by IFS HRES, Pangu-Weather, GraphCast, and FengWu. To evaluate the medium-range forecast, in accordance with the methods of WeatherBench and DABench, this embodiment selects 50 initial fields at intervals of 336 hours for medium-range weather forecasting experiments. The first initial field at 00:00 UTC is set to January 1, 2023, and the first initial field at 12:00 UTC is set to January 8, 2023.
[0087] All metrics are calculated with float32 precision and reported using the original magnitudes of the variables without normalization. It should be noted that due to the uneven regional distribution from the equator to the north and south poles, the latitude weighting factor of the grid points is used in the calculation of all metrics.
[0088] (1) Root Mean Square Error (RMSE).
[0089] This embodiment uses the latitude-weighted root mean square error (RMSE) to evaluate the assimilation and forecasting skills of a given variable The calculation formula is: ; where is the field to be evaluated, is the ERA5 training label, in represents the sample index in the evaluation dataset, represents the latitude coordinate in the grid, represents the longitude coordinate in the grid. The smaller the RMSE, the better the result.
[0090] (2) Bias.
[0091] This embodiment also calculates the bias of a specific variable: ; The closer the deviation is to 0, the better the result.
[0092] (3) Anomaly Correlation Coefficient (ACC).
[0093] To study the proficient available forecast lead time, this embodiment also calculates the Latitude Weighted Anomaly Correlation Coefficient (ACC) according to the following formula: ; where represents the climatological mean of a given variable and the annual day containing the grid point valid time. It is calculated with reference to the GraphCast, FengWu, and DABench models. The climatological mean is calculated using ERA5 data from 2010 to 2021. The higher the ACC, the better the result.
[0094] Compared with the prior art, the method of this embodiment has the following advantages: Existing artificial intelligence-based models still rely on the expensive data assimilation (DA) module of the Numerical Weather Prediction (NWP) system to construct the initial field, which limits the operationalization of a fully end-to-end weather forecasting system. Therefore, this embodiment proposes the Xichen model (i.e., the weather forecasting model), which is a fundamental model integrating the entire NWP workflow, including medium-range weather forecasting, observation operators, and cascaded DA. The Xichen model incorporates four-dimensional variational prior knowledge and effectively absorbs raw satellite observation data through an artificial intelligence-based observation operator. The research results of this embodiment show that the analysis fields and forecasts generated by the Xichen model can be comparable to those generated by an operational NWP system. In addition, the scalability of the Xichen model allows for the integration of more satellite observation data, thereby improving forecast performance. These research results highlight the potential of expertise-guided fundamental models as a promising direction for advancing the next generation of end-to-end weather forecasting systems.
[0095] The above has described the embodiments of the present application in detail with reference to the accompanying drawings, but the present application is not limited to the above embodiments. Various changes can be made without departing from the spirit of the present application within the knowledge scope of those of ordinary skill in the art to which the present application pertains.
Claims
1. An end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge, characterized in that The method includes: Obtaining a target background field, a target forecast lead time, ERA5 reanalysis data, first conventional observation data, first satellite observation data, second conventional observation data, and second satellite observation data, where the conventional observation data is the observation data in the global data assimilation system, and the satellite observation data includes the satellite observation data obtained by the advanced microwave sounding unit and the satellite observation data obtained by the microwave humidity sensor; Pre-training the constructed weather forecasting model into a medium-term weather forecasting model using the ERA5 reanalysis data, and obtaining the prediction result of the medium-term weather forecasting model; Fine-tuning the medium-term weather forecasting model into a first observation operator using the first satellite observation data; Inputting the first observation operator, the prediction result, the first conventional observation data, and the first satellite observation data into a four-dimensional variational cost function to obtain a first gradient of the four-dimensional variational cost function; Taking the prediction result as a first background field, and fine-tuning the weather forecasting model into a data assimilation model using the first gradient and the first background field; Inputting the second conventional observation data and the second satellite observation data into the data assimilation model, fine-tuning the data assimilation model into a second observation operator using the second satellite observation data, and inputting the second observation operator and the second conventional observation data into the four-dimensional variational cost function to obtain a second gradient of the four-dimensional variational cost function; Inputting the second gradient and the target background field into the data assimilation model for fine-tuning to obtain an analysis field output by the fine-tuned data assimilation model; Inputting the analysis field and the target forecast lead time into the fine-tuned data assimilation model to obtain a target weather prediction result.
2. The end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge according to claim 1, wherein The pre-training the constructed weather forecasting model into a medium-term weather forecasting model using the ERA5 reanalysis data includes: Constructing a pre-training model loss function as: ; Among them, represents the loss function of the pre-trained model, represents the expected value, represents the total number of forecast steps, represents the total number of variables, represents the number of latitude coordinates, represents the number of longitude coordinates, represents the variable 's pressure weighting, represents the latitude weighting factor, represents the variable at time and the difference between the predicted state and the training label at the longitude and latitude coordinates , represents the forecast step, represents the F1 norm; Based on the pre-training model loss function, pre-training the constructed weather forecasting model into a medium-term weather forecasting model using the ERA5 reanalysis data.
3. The end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge according to claim 1, characterized in that, The fine-tuning the medium-term weather forecasting model into a first observation operator using the first satellite observation data includes: Constructing a fine-tuning operator loss function as: ; Among them, represents the fine-tuning operator loss function, represents the expected value, represents the channel of satellite observation, represents the observation operator of the channel calculated at time ; the state at represents the auxiliary observation information at time, represents the first satellite observation data of the channel; represents the F1 norm; Based on the fine-tuning operator loss function, fine-tuning the medium-term weather forecasting model into a first observation operator using the first satellite observation data.
4. The end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge according to claim 1, wherein Before inputting the first observation operator, the prediction result, the first conventional observation data, and the first satellite observation data into the four-dimensional variational cost function, the method further includes: Constructing a four-dimensional variational cost function: ; Among them, represents the four-dimensional variational cost function, represents the background field, represents the initial field to be optimized, represents the background error covariance matrix, represents the inverse transformation of the matrix ; represents the total time, represents the time observation value of, represents the prediction result, represents the observation operator, represents the observation error covariance matrix, represents the inverse transformation of the matrix . It should be noted that in the original text, there is a semicolon (;) in line 12 which is missing in the English translation. I have added it in the translation for better semantic integrity. If this is not allowed according to the strict rules, please adjust accordingly.
5. The end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge according to claim 1, characterized in that The fine-tuning the weather forecasting model into a data assimilation model using the first gradient and the first background field includes: Constructing a target loss function: ; Among them, represents the target loss function, represents the expected value, represents the total number of variables, represents the number of latitude coordinates, represents the number of longitude coordinates, represents the variable weighted by pressure, represents the latitude weighting factor, represents the variable at time and the latitude and longitude coordinates in the initial field state, represents the variable at time and the latitude and longitude coordinates in the state, represents the F1 norm; Based on the target loss function, fine-tuning the weather forecasting model into a data assimilation model using the first gradient and the first background field.
6. The end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge according to claim 1, characterized in that The pre-training the constructed weather forecasting model into a medium-term weather forecasting model using the ERA5 reanalysis data includes: Construct a conditional hybrid neural operator including a conditional feed-forward network, a cross-attention mechanism, layer normalization, an attention mechanism, and a feed-forward network; Construct a hybrid neural operator including layer normalization, an attention mechanism, and a feed-forward network; Construct a weather forecasting model according to the conditional hybrid neural operator and the hybrid neural operator; Pre-train the constructed weather forecasting model into a medium-term weather forecasting model using the ERA5 reanalysis data.
7. The end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge according to claim 6, wherein The constructing a weather forecasting model according to the conditional hybrid neural operator and the hybrid neural operator includes: Construct an encoder based on the conditional hybrid neural operator; Construct a decoder based on the hybrid neural operator; Construct a weather forecasting model including the encoder, the decoder, patch embedding, a feature fusion module, and a fully-connected layer, where the feature fusion module includes a multi-layer perceptron.
8. The end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge according to claim 6, wherein The constructing a conditional hybrid neural operator including a conditional feed-forward network, a cross-attention mechanism, layer normalization, an attention mechanism, and a feed-forward network includes: Input a conditional input item into the conditional feed-forward network to obtain an output result of the conditional feed-forward network; Input the output result of the conditional feed-forward network into layer normalization to obtain a first layer normalization result; Input an input item into layer normalization to obtain a second layer normalization result; Input the second layer normalization result into the attention mechanism to obtain an attention result; Perform a residual connection on the attention result and the input item to obtain a first addition result; Input the first addition result into layer normalization to obtain a third layer normalization result; Input the first layer normalization result and the third layer normalization result into the cross-attention mechanism to obtain a cross-attention result; Input the cross-attention result and the first addition result into layer normalization to obtain a fourth layer normalization result; Input the fourth layer normalization result into the feed-forward network to obtain an output result of the feed-forward network; Perform a residual connection on the output result of the feed-forward network and the cross-attention result to obtain an output result of the conditional hybrid neural operator, so as to construct a conditional hybrid neural operator.
9. The end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge according to claim 6, characterized in that The constructing a hybrid neural operator including layer normalization, an attention mechanism, and a feed-forward network includes: Input an input item into layer normalization to obtain a first layer normalization result; Input the first layer normalization result into the attention mechanism to obtain an attention result; Perform a residual connection on the attention result and the input item to obtain a first addition result; Input the first addition result into layer normalization to obtain a second layer normalization result; Input the second layer normalization result into the feed-forward network to obtain an output result of the feed-forward network; Perform a residual connection on the output result of the feed-forward network and the first addition result to obtain an output result of the hybrid neural operator, so as to construct a hybrid neural operator.
10. The end-to-end data-driven weather forecasting method based on four-dimensional variational prior knowledge according to claim 9, characterized in that The attention mechanism includes fast Fourier transform, inverse fast Fourier transform, frequency-domain convolution, and spatial-domain convolution. The inputting the first layer normalization result into the attention mechanism to obtain an attention result includes: Process the first layer normalization result through the fast Fourier transform to obtain a first processed result; Input the first processing result into the frequency-domain convolution to obtain a first convolution result; Process the first convolution result through the inverse fast Fourier transform to obtain a second processing result; Input the first-layer normalization result into the spatial-domain convolution to obtain a second convolution result; Perform residual connection on the second processing result, the second convolution result, and the first-layer normalization result to obtain an attention result.
Citation Information
Patent Citations
Numerical weather forecasting method and device, computer storage medium and electronic equipment
CN113568067A
Data assimilation method, system and equipment based on four-dimensional variational constraint and medium
CN117763065A
Attention neural network data assimilation method based on four-dimensional variational constraint
CN117972625A
Systems and methods for converting live weather data to weather index for offsetting weather risk
EP3891534A1