Global climate mode set average estimation deviation correction method
By identifying the heterogeneity between global climate model models and building a grid-based convolutional neural network, extracting high-dimensional information and assigning weights, the deviation problem of GCMs when simulating precipitation is solved, achieving a more accurate and reliable deviation correction effect.
Patent Information
- Application Number
- CN202510115361.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-06-27
AI Technical Summary
Global climate models (GCMs) have large deviations in simulating precipitation, and existing deviation correction methods fail to fully utilize the model's ability to represent atmospheric circulation and processes.
By identifying the heterogeneity between different models of global climate patterns, a convolutional neural network based on grid points is built to extract high-dimensional information and assign higher weights to models similar to observations in atmospheric dynamics, and bias correction is performed.
More accurate deviation correction results are achieved, and a priori atmospheric physical information is fully utilized, the reliability of physical information is improved, and the convolutional neural network method is better than the spatial correlation-based convolutional neural network method.
Smart Images

Figure CN120217276A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hydrological statistics, and more particularly to a method for correcting the ensemble mean prediction bias of global climate models. Background Art
[0002] Precipitation has important impacts on agriculture, ecosystems, social economy and human life. Climate change has changed the occurrence time, intensity and distribution characteristics of precipitation, and exacerbated extreme hydrological events, leading to more frequent floods, droughts, heatwaves and wildfires. The change of precipitation poses a huge threat to social development, especially to vulnerable developing economies.
[0003] Global climate models (GCMs) are commonly used to simulate climate change and evaluate its impact on water resource systems. Although significant progress has been made in global climate models in recent decades, accurately simulating precipitation remains a key challenge in existing GCMs. This is because there are considerable biases in model predictions due to their coarse resolution, imperfect boundary conditions, parameter uncertainties and unreasonable descriptions of physical processes. It is inaccurate to directly apply the original output of GCM precipitation to drive hydrological models without prior bias correction. Bias correction refers to using statistical algorithms to correct the output of climate models and generate the same statistical characteristics as the reference data. Currently, the bias correction methods developed in the world can be mainly divided into two categories: non-neural network-based and neural network-based. Non-neural network-based bias correction methods mainly include linear scaling, variance scaling, quantile mapping, etc. Neural network-based bias correction methods mainly include random forest, support vector machine, convolutional neural network, long short-term memory network, etc.
[0004] Generally, data-driven neural network-based methods have better performance than non-neural network-based methods. However, currently, neural network-based bias correction methods usually only consider the spatial correlation between different grids, without utilizing the relative performance of the models in representing the atmosphere circulation and processes. Summary of the Invention
[0005] Object of the Invention: To overcome the deficiencies of the background art, the present invention discloses a method for correcting the ensemble mean prediction bias of global climate models. This method corrects the ensemble mean bias of the models by identifying the heterogeneity between different models of global climate models, extracting the high-dimensional information (such as atmospheric dynamics) contained behind these models, and assigning higher weights to the models that are similar to the observations in terms of atmospheric dynamics. The models with poor performance are given very small weights in this process.
[0006] Technical Solution: The method for correcting the ensemble mean prediction bias of global climate models disclosed by the present invention includes the following steps:
[0007] S1. Process the resolution of the climate model data to be consistent with the observed data. At the same time, check for missing values and outliers, fill in the missing values, remove the outliers, and convert the data into a definite format.
[0008] S2. Build a convolutional neural network oriented by model heterogeneity based on grids, including an input grid, three network blocks, two fully connected layers, and an output grid.
[0009] S3. Input the sorted data into the model. In the three network blocks, the channel dimension of the data gradually increases, while the number of models and the length of the time series gradually decrease. Then, map it to the same dimension size as the observed data through the fully connected layer, calculate its loss with the observed data, and calculate the gradient of the loss with respect to each parameter in the model.
[0010] S4. Update the weights of the model by the gradient descent method and repeat S3 until the loss no longer decreases, which means the model training is completed.
[0011] Furthermore, in S1, the specific format of data processing is n*c*x*y, where n is the number of grids, c is the channel dimension representing the characteristics of atmospheric dynamics, and x and y are the number of models and the length of the time series respectively. The processed format represents n images composed of different models and time series, that is, the information of a specific variable at a grid point can be characterized by the information of multiple model results within a time series.
[0012] Furthermore, in S2, each network block includes a convolutional layer, a skip layer, a max pooling layer, and an activation function; it is managed by the built-in model building library Module in pytorch.
[0013] Furthermore, S3 specifically includes:
[0014] S3.1 Divide all the data into a training set and a test set. For the training set, through convolution, fuse the information of all model results within a time series and extract the high-dimensional information of the latent representation of these models:
[0015]
[0016] where represents the output value of the k-th convolutional kernel at the i-th model and time j position, and x is the input feature map; W (k) is the weight of the k-th convolutional kernel; b (k) is the bias of the k-th convolutional kernel; M and N are the height and width of the convolutional kernel respectively.
[0017] S3.2 Input the result of convolution into the skip layer to prevent gradient disappearance and network degradation:
[0018] y = f(x) + x
[0019] where x is the input of the skip layer, f represents the convolution in S3.1, and y represents the output of the skip layer;
[0020] S3.3 inputs the result after the skip layer into the max pooling layer to reduce the spatial size of the feature map to highlight significant features, i.e., the heterogeneity of the model:
[0021] Z i,j = max{x m,n | m ∈ [i, i + p), n ∈ [j, j + p)}
[0022] where Z i,j represents the output of downsampling, x m,n is the input value within the downsampling window, and p is the size of the downsampling window;
[0023] S3.4 inputs the result of the downsampling layer into the activation function to increase the non-linearity of the network:
[0024] f(x) = max(0, x)
[0025] S3.5 runs the above steps in sequence three times and then inputs the result into the fully connected layer, mapping the high-dimensional features reflected by the model heterogeneity extracted by the convolution to the final output, i.e., the result of bias correction:
[0026] z = f(W·x + b)
[0027] where x is the input, W is the weight matrix, f is the activation function, and b is the bias;
[0028] S3.6 compares the output obtained in S3.5 with the observation result, measures the difference through the L2 norm, and calculates the loss:
[0029] loss = ||z - z ob || 2
[0030]
[0031] where loss represents the loss, z ob represents the observation, grad represents the gradient, and θ represents the parameters of the model.
[0032] Furthermore, in S4, the first moment and the second moment of the calculated gradient are computed, and the Adam optimizer is used for gradient descent, and the root mean square error is used to measure the loss:
[0033]
[0034] where θt Denote the current parameter, m t is the bias correction value of the first moment of the gradient, v t is the bias correction value of the second moment of the gradient, α is the learning rate, and ∈ is a small constant to prevent division by zero errors.
[0035] Beneficial effects: Compared with the prior art, the advantages of the present invention are as follows: The present invention has better bias correction results than the convolutional neural network based on spatial correlation, and fully utilizes the prior atmospheric physical information. Specifically, it corrects the average bias of the ensemble by allocating weights based on the relative performance of the model in terms of the ability to represent atmospheric circulation and processes, and has more reliable physical information guarantee. Brief Description of the Drawings
[0036] Figure 1 is the flowchart of the method of the present invention;
[0037] Figure 2 is a schematic diagram comparing the structures of the lattice model heterogeneity-guided convolutional neural network (A) and the space-based convolutional neural network (B);
[0038] Figure 3 is the spatial distribution of the bias ratio before and after ensemble mean bias correction of 30 models in summer (June, July, August) and winter (December, January, February) by four different bias correction methods: linear regression, quantile mapping, convolutional neural network based on spatial correlation, and lattice model heterogeneity-guided convolutional neural network;
[0039] Figure 4 is the mean absolute error and relative bias before and after ensemble mean bias correction of 30 models in summer (June, July, August) and winter (December, January, February) by four bias correction methods in the embodiment;
[0040] Figure 5 is the correlation coefficient and Kling-Gupta efficiency before and after ensemble mean bias correction of 30 models in summer (June, July, August) and winter (December, January, February) by four bias correction methods in the embodiment;
[0041] Figure 6 is the generalized extreme value distribution and extreme value index (EVI) of extreme precipitation (defined as precipitation above the 95th percentile threshold) before and after ensemble mean bias correction of 30 models in summer (June, July, August) and winter (December, January, February) by four bias correction methods in the test set of the embodiment; Detailed Embodiment
[0042] The technical solutions of the present invention will be further described below in conjunction with the drawings and embodiments.
[0043] As Figure 1 shown in the global climate model ensemble mean prediction deviation correction method, including the following steps:
[0044] S1. Process the resolution of the climate model data to be consistent with the observed data. At the same time, check whether there are missing values and outliers, fill in the missing values, remove the outliers, and convert the data into a definite format; the specific format of the data processing is n*c*x*y, where n is the number of grid points, c is the channel dimension representing the characteristics of atmospheric dynamics, x and y are the number of models and the length of the time series respectively. The processed format represents n images composed of different models and time series, that is, the information of a specific variable at a grid point can be characterized by the information of multiple model results within a time series.
[0045] S2. Build a convolutional neural network oriented by grid-based model heterogeneity, including an input grid, three network blocks, two fully connected layers, and an output grid; each network block includes a convolutional layer, a skip layer, a max pooling layer, and an activation function; managed by the model building library Module built in pytorch.
[0046] S3. Input the sorted data into the model. In the process of passing through the three network blocks, the channel dimension of the data gradually increases, and the number of models and the length of the time series gradually decrease. Then, map it to the same dimension size as the observed data through the fully connected layer, calculate its loss with the observed data, and calculate the gradient of the loss with respect to each parameter in the model;
[0047] Specifically including:
[0048] S3.1 Divide all the data into a training set and a test set. For the training set, through convolution, fuse the information of all model results within a time series, and extract the high-dimensional information of the latent representation of these models:
[0049]
[0050] where represents the output value of the kth convolutional kernel at the jth position of the ith model, x is the input feature map; W (k) is the weight of the kth convolutional kernel; b (k) is the bias of the kth convolutional kernel; M and N are the height and width of the convolutional kernel respectively;
[0051] S3.2 Input the result of convolution into the skip layer to prevent gradient disappearance and network degradation:
[0052] y = f(x) + x
[0053] Where x is the input of the skip layer, f represents the convolution in S3.1, and y represents the output of the skip layer;
[0054] S3.3 inputs the result after the skip layer into the max pooling layer to reduce the spatial size of the feature map in order to highlight significant features, that is, the heterogeneity of the model:
[0055] Z i,j = max{x m,n |m ∈ [i, i + p), n ∈ [j, j + p)}
[0056] Where Z i,j represents the output of downsampling, x m,n is the input value within the downsampling window, and p is the size of the downsampling window;
[0057] S3.4 inputs the result of the downsampling layer into the activation function to increase the non-linearity of the network:
[0058] f(x) = max(0, x)
[0059] S3.5 runs the above steps three times in sequence and then inputs the result into the fully connected layer, mapping the high-dimensional features reflected by the model heterogeneity extracted by the convolution to the final output, that is, the result of bias correction:
[0060] z = f(W·x + b)
[0061] Where x is the input, W is the weight matrix, f is the activation function, and b is the bias;
[0062] S3.6 compares the output obtained in S3.5 with the observation result, measures the difference through the L2 norm, and calculates the loss:
[0063] loss = ||z - z ob || 2
[0064]
[0065] Where loss represents the loss, z ob represents the observation, grad represents the gradient, and θ represents the parameters of the model.
[0066] S4. Update the weights of the model through the gradient descent method and repeat S3 until the loss no longer decreases, that is, the model training is completed. Calculate the first moment and second moment of the obtained gradient, and use the Adam optimizer for gradient descent, and use the root mean square error to measure the loss:
[0067]
[0068] Where θ t represents the current parameter, mt is the bias correction value of the first moment of the gradient, v t is the bias correction value of the second moment of the gradient, α is the learning rate, and ∈ is a small constant to prevent division-by-zero errors.
[0069] Taking 30 models of the global climate model as an example, as shown in the following table:
[0070]
[0071]
[0072] S1: In the global region, 30 models of the global climate model are used as the models to be corrected, and CRU TS4.0.4 (Climatic Research Unit gridded Time Series version 4.0.4) data is used as the observation data. Both the model and the observation data are bilinearly interpolated to a resolution of 1°×1°. The data from 1901 - 1996 (a total of 96 years) is taken as the training set, and the data from 1997 - 2013 (a total of 16 years) is taken as the test set. Using the pytorch library in python, first convert the data from the array format to the torch format. Then perform data segmentation and dimension transformation:
[0073] The present invention: [t, m, x, y] → [n, c, m, t0]
[0074] Spatially-based convolutional neural network: [t, m, x, y]
[0075] Where: t represents time (96 in this embodiment, representing 1996), m represents the model (30 in this embodiment, representing 30 models), x represents latitude (180 in this embodiment), y represents longitude (360 in this embodiment), n represents the number of grid points (93228 in this embodiment), c represents the representation of atmospheric physical characteristics (initially set to 1), and t0 represents the segmented time period (16 in this embodiment). The data conversion process is as follows. First, 1996 is divided into 6 time periods, each time period being 16 years, i.e., t0, which is placed in the last dimension. (This can be changed according to different data.) The present invention is based on grid points for correction, and in this embodiment, only the land grid points throughout the process are considered for correction. Therefore, the longitude and latitude x, y are converted into the number of land grid points n0, i.e., 15538, which is placed in the first dimension. At this time, the data shape is [n0, m, t0]. Then, a dimension is expanded at position 2 of the dimension, and the data shape becomes [n0, 1, m, t0]. Here, 1 is the initial value of c. Since we divide it into 6 time periods, there are 6 such arrays. Then, the 6 arrays are stacked together in dimension 1 in chronological order. Therefore, n in this embodiment is 6 × 15538 = 93228. Finally, the shape of the training data is [n, c, m, t0].
[0076] Meanwhile, the CRU observed data is also segmented and its dimensions are transformed:
[0077] The present invention: [t, x, y] → [n, 1]
[0078] Convolutional neural network based on space: [x, y]
[0079] The CRU data is divided into the same 6 time periods, each time period also being 16 years, with a shape of [15538, 16, 6]. The average is taken in the second dimension and it is stacked in the first dimension. Finally, the shape becomes [n, 1], i.e., [93228, 1]. And the torch.utils.data.DataLoader is used to make iterators for the model data and the observed data.
[0080] S2: Use the torch.nn library and torch.nn.Module to build the model, including torch.nn.Conv2, torch.nn.Maxpool2D, torch.nn.ReLu, torch.nn.Linear, and write the skip layer well. The model is as Figure 2 shown.
[0081] S3: Randomly extract a sample batch from the iterator constructed in S1. The shape of the sample is [64, 1, 30, 16], and it is input into the convolutional layer (), that is, the batch is input into the formula Obtain a feature vector that can represent the latent high-dimensional information of the model. During this process, the shape and size of the convolutional kernel are set to [3, 3]. After that, we perform a skip operation, perform a full convolution on the batch (i.e., use the same convolution formula, but the size of the convolutional kernel is set to 1), add the results of the two, and obtain a feature vector with a shape and size of [64, 64, 30, 16]. Output this result to the max pooling layer, i.e., Z i,j = max{x m,n |m ∈ [i, i + p), n ∈ [j, j + p)}, and the shape and size of the downsampling window P are set to [2x2]. That is, by sliding p, retain the maximum value within the window, thereby reducing the spatial dimension of the feature vector to highlight the most significant features. After max pooling, the shape and size of the output result are [64, 64, 15, 8]. Finally, input it to the activation layer to enhance the non-linear fitting ability. This is the process of a sample passing through a network "block". Perform the same operation three times. Finally, after passing through all network "blocks", the size of the feature vector is [64, 512, 4, 2]. Finally, input it to the fully connected layer z = f(W·x + b) to obtain the final result pr correct , with a shape and size of [64, 1], representing the precipitation values of these 64 samples after calibration.
[0082] S4: Calculate the loss between the calibrated precipitation and the observed precipitation through the root mean square error RMSE, i.e., Loss = RMSE(pr ob , pr correct ). Calculate the gradient of the loss with respect to each parameter through the automatic differentiation function of pytorch. Finally, update the parameters through the Adma optimizer. After training, use the trained model to test the test data, and the results are as Figure 3 , Figure 4 , Figure 5 , Figure 6 shown. The results of the ensemble mean bias correction of the present invention are better than those of traditional space-based convolutional neural networks and non-convolutional neural network-based methods in all indicators.
[0083] From the above analysis, it can be seen that the present invention convolves the model with time to identify the heterogeneity between different models of the global climate pattern, extracts the high-dimensional information contained behind these models, assigns higher weights to those models that are similar to the observations in atmospheric dynamics, and those models with poor performance are given very small weights during this process, thereby correcting the ensemble mean bias. Compared with traditional technologies, the present invention makes more full use of prior atmospheric physical information to achieve better bias correction effects.
Claims
1. A method for correcting the bias of a global climate model ensemble mean estimate, characterized in that: The following steps are involved: S1. Process the resolution of climate model data to be consistent with the observed data, check whether there are missing values and outliers, fill in missing values, remove outliers, and convert the data into a certain format; S2. Build a grid-based model heterogeneity-oriented convolutional neural network, including an input grid, three network blocks, two fully connected layers, and an output grid; S3. Input the sorted data into the model. In the three network blocks, the channel dimension of the data gradually increases, the number of models and the length of the time series gradually decrease, and then it is mapped to the same size as the observed data dimension through the fully connected layer, and the loss with the observed data is calculated, and the gradient of the loss for each parameter in the model is calculated; S4: Update the model weights through the gradient descent method and repeat S3 until the loss stops decreasing, indicating that the model training is complete.
2. The method for correcting the global climate model ensemble mean estimate bias according to claim 1, characterized in that: In S1, the specific format of data processing is n*c*x*y, where n is the number of grid points, c is the channel dimension representing the characteristics of atmospheric dynamics, x and y are the number of models and the length of the time series, respectively. The processed format represents n images composed of different models and time series, that is, the information of a specific variable at a grid point can be represented by the information of multiple model results in a time series.
3. The method for correcting the global climate model ensemble mean estimate bias according to claim 2, characterized in that: In S2, each network block includes a convolution layer, a skip layer, a maximum pooling layer, and an activation function; it is managed by the built-in model building library Module of PyTorch.
4. The method for correcting the global climate model ensemble mean estimate bias according to claim 3, characterized in that: S3 specifically includes: S3.1 divides all data into training set and test set. For the training set, the information of all model results in a time series is fused through convolution to extract the high-dimensional information potentially represented by these models: in represents the output value of the kth convolution kernel in the i-th model at time j, and x is the input feature map; W (k) is the weight of the kth convolution kernel; b (k) is the bias of the kth convolution kernel; M and N are the height and width of the convolution kernel respectively; S3.2 inputs the result of the convolution into the jump layer to prevent the gradient from disappearing and the network from degrading: y=f(x)+x Where x is the input of the skip layer, f represents the convolution in S3.1, and y represents the output of the skip layer; S3.3 inputs the result after the skip layer into the maximum pooling layer to reduce the spatial size of the feature map in order to highlight the significant features, that is, the heterogeneity of the model: Z i,j =max{x m,n |m∈[i,i+p),n∈[j,j+p)} Where Z i,j represents the downsampled output, x m,n is the input value within the downsampling window, and p is the size of the downsampling window; S3.4 inputs the result of the downsampling layer into the activation function to increase the nonlinearity of the network: f(x)=max(0,x) S3.5 After running the above steps three times in sequence, the results are input into the fully connected layer, and the high-dimensional features reflected by the model heterogeneity extracted by convolution are mapped to the final output, which is the result of bias correction: z=f(W·x+b) Where x is the input, W is the weight matrix, f is the activation function, and b is the bias; S3.6 compares the output obtained in S3.5 with the observed result, measures the difference by the L2 norm, and calculates the loss: loss=||z-z ob || 2 Where loss represents the loss, z ob represents observation, grad represents gradient, and θ represents the parameters of the model.
5. The method for correcting the global climate model ensemble mean estimate bias according to claim 4, characterized in that: In S4, the first-order and second-order moments of the gradient are calculated, and the Adam optimizer is used for gradient descent, and the root mean square error is used to measure the loss: where θ t Indicates the current parameter, m t is the deviation correction value of the first moment of the gradient, v t is the bias correction for the second moment of the gradient, α is the learning rate, and ∈ is a small constant to prevent division by zero errors.