A Machine Learning-Based Acceleration Method for Random Cloud Generators
By building a multi-layer perceptron model, the anti-correlation thickness is directly calculated, the time-consuming problem of random cloud generators is solved, efficient and accurate cloud overlap hypothesis application is achieved, and the simulation and prediction accuracy of climate modes is improved.
Patent Information
- Application Number
- CN202310285960.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-22
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-03-22
AI Technical Summary
In the prior art, random cloud generators take too long to find suitable anti-correlation thicknesses, affecting the application efficiency and accuracy of cloud overlap assumptions in climate modes.
A multi-layer perceptron model is used to build a machine learning model. By inputting cloud profile and total cloud data, a parameterized relationship between resolution, latitude and anti-correlation thickness is established, and the anti-correlation thickness is directly calculated, replacing the traditional random cloud generator traversal search process.
The calculation speed and accuracy of anti-correlation thickness are significantly improved, the impact of cloud uncertainty on numerical modes is reduced, and the simulation and prediction accuracy of climate modes is improved.
Smart Images

Figure CN116341379B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cloud physical parameterization, and particularly relates to an acceleration method for a random cloud generator based on machine learning. Background Art
[0002] Clouds play a crucial role in the radiation budget balance of the earth - atmosphere system and cover approximately 60% of the earth's surface area. Clouds are the result of the comprehensive action of various dynamic and thermodynamic processes in the climate system. Through their microphysical characteristics and the physical processes involved in their generation and development, they can affect the climate system in various aspects.
[0003] A numerical model in atmospheric science refers to the process of numerically simulating and predicting the atmospheric system using a mathematical model. This model is based on physical, chemical, and mathematical principles and predicts future weather and climate changes by calculating and simulating various variables in the atmospheric system. Numerical models usually include atmospheric dynamics equations, thermodynamics equations, radiation transfer equations, cloud microphysical equations, etc. As an important approach for studying global climate change, the simulation and description of clouds in numerical climate models are relatively weak. The grid scale of global weather and climate models is generally dozens or even hundreds of kilometers or more. Although the vertical profile of clouds can be obtained through diagnostic methods in the model, due to the complexity of the vertical overlap relationship of cloud layers, the vertical overlap characteristics of clouds - that is, how the cloud layers at different heights overlap vertically, cannot be accurately constructed at the grid resolution of global climate models. Currently, only a relatively rough cloud vertical overlap assumption can be given for description.
[0004] Currently, the more popular cloud overlap assumptions include: maximum overlap assumption, random overlap assumption, maximum random overlap assumption, and exponential decay overlap assumption, among which the exponential decay overlap assumption is the most widely used. This assumption represents the cloud amount between cloud layers as a linear combination of maximum overlap and random overlap through a weighting factor called the cloud overlap parameter (α), and further fits α as a function of the anticorrelation thickness L and the inter - layer distance D: This assumption effectively avoids the dependence on the vertical resolution of climate models and fully reflects the diversity of cloud layer overlap characteristics, having obvious advantages compared with other overlap assumptions.
[0005] In order to apply this flexible exponential decay overlap assumption in climate models, a random cloud generator is generally required to generate the vertical structure of clouds. However, this generation process requires us to provide the anticorrelation thickness L between cloud layers, but the random cloud generator requires a long random number generation process in the process of finding the appropriate anticorrelation thickness, which is very time - consuming. Summary of the Invention
[0006] The object of the present invention is to provide an acceleration method for a random cloud generator based on machine learning, which is fast and accurate.
[0007] The acceleration method for a random cloud generator based on machine learning provided by the present invention specifically comprises the following steps:
[0008] (1) Construct a machine learning model for generating a large amount of anti-correlated thickness data; the machine learning model is a Multilayer Perception (MLP) model;
[0009] (2) Input the cloud amount profile data and total cloud amount data obtained from satellite observations into the machine learning model to obtain the corresponding anti-correlated thickness;
[0010] (3) Combine the resolution and latitude information during satellite observation with the anti-correlated thickness data to obtain a triple dataset of resolution, latitude, and anti-correlated thickness; according to this triple dataset, establish a parametric relationship between resolution, latitude, and anti-correlated thickness, specifically a fifth-order polynomial (based on data from May and June 2008):
[0011] F(Res,Lat) = c + a1*Res + a2*Lat + a3*Res 2 + a4*Res*Lat + a5*Lat 2 … + a 20 *Lat 5
[0012] where Res and Lat are the input resolution and latitude, c is the constant term, and a i , i = 1, 2…, 20 are polynomial coefficients;
[0013] (4) Directly apply this parametric relationship to a climate numerical model (the climate numerical model is an existing technology) that uses a mathematical model to numerically simulate and predict the atmospheric system. The numerical model directly obtains the mean value of the anti-correlated thickness at the resolution and latitude of the simulation area according to the resolution and latitude of the simulation area. Using this mean value of the anti-correlated thickness, the cloud overlap situation in this area can be directly simulated, which can greatly reduce the impact of cloud uncertainty on the numerical model and has important significance for the cloud overlap application in the numerical model, and can greatly improve the accuracy of the simulation and prediction of the numerical model.
[0014] The machine learning model constructed by the present invention includes an input layer, a hidden layer and an output layer, with a total of eight levels: the first layer is the input layer, with a total of 105 input neurons, which are used to read the total cloud cover and cloud cover profile information; the hidden layer has a total of six layers of neural networks, which receive feature vectors from the input layer, wherein the first four layers of neural networks are composed of 128 neurons, the fifth layer of neural networks is composed of 64 neurons, and the sixth layer of neural networks is composed of 32 neurons. In the hidden layer, the first four layers are to introduce enough neurons to extract enough features for learning and enhance the fitting ability of the neural network, and the last two layers are to gradually aggregate the previously extracted features and aggregate the features into the output layer. The last layer is the output layer, with a total of 1 neuron, which is used to output the final fitting results of all layers, that is, the anti-correlation thickness.
[0015] In the present invention, the training of the machine learning model uses the mean square error (MSE) as the loss function. By calculating the mean square error between the predicted value and the true value, the difference between the model prediction result and the true value is evaluated, thereby further optimizing the performance of the model. During the training process, the back propagation algorithm is used to calculate the gradient, and the gradient descent method is used to optimize the MSE loss function.
[0016] The present invention uses a training strategy with a batch size of 256 to train the model, that is, 256 samples are calculated simultaneously in one iteration, and their gradient averages are used to update the model parameters. This training strategy can improve training efficiency, reduce memory consumption, and usually achieve better convergence performance in a shorter time.
[0017] After completing the training of the model, you only need to input the observed cloud profile and total cloud amount into the model. After the model's intermediate layer calculation, the model can obtain the anti-correlation thickness corresponding to the cloud profile and the total cloud amount, and output it.
[0018] Beneficial effects of the present invention:
[0019] Compared with the traversal search method of traditional random cloud generators to find suitable anti-correlation thickness, after the machine learning model of the present invention is preliminarily trained for two thousand rounds with four months of data from the 94GHz millimeter-wave cloud profiling radar (Cloud Profiling Radar-CPR) carried by the CloudSat satellite, the coefficient of determination between the output of the machine learning model and the true value on the training set can reach a high value of 0.986, and the coefficient of determination between the output and the true value on the test set can also reach a high value of 0.953. And for about 700,000 groups of data in four months, the machine learning model of the present invention only needs about 6 seconds to obtain all the outputs; while for the traditional random cloud generator, it takes about 2,000 seconds to obtain all the results for only 400 groups of input data. Therefore, it can be seen that the anti-correlation thickness obtained by the present invention has high precision and the operation speed far exceeds that of the random cloud generator. Description of the Drawings
[0020] Figure 1 Schematic diagram of the sub-grid division of the random cloud generator.
[0021] Figure 2 Schematic diagram of the machine learning model adopted by the present invention.
[0022] Figure 3 Comparison of the anti-correlation thickness generated by the present invention on the training set with the results of the random cloud generator.
[0023] Figure 4 Comparison of the anti-correlation thickness generated by the present invention on the test set with the results of the random cloud generator.
[0024] Figure 5 Relationship between the total cloud amount and the true cloud amount obtained when the anti-correlation thickness generated by the present invention on the training set is used for the random cloud generator.
[0025] Figure 6 Relationship between the total cloud amount and the true cloud amount obtained when the anti-correlation thickness generated by the present invention on the test set is used for the random cloud generator.
[0026] Figure 7 Parameterization relationship established between latitude and anti-correlation thickness when using the dataset generated by the present invention at a model grid resolution of 4km.
[0027] Figure 8 Parameterization relationship established between model grid resolution and anti-correlation thickness when using the dataset generated by the present invention at a latitude of 40°N.
[0028] Figure 9 Parameterization relationship established between grid resolution and latitude when using the dataset generated by the present invention. Detailed implementation mode
[0029] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0030] Embodiment
[0031] For the acceleration method of the random cloud generator based on machine learning of the present invention, a machine learning model is constructed to replace the random cloud generator in the search process for the relevant thickness.
[0032] As Figure 1 shown, the search process of the random cloud generator for the relevant thickness: The random cloud generator further divides a grid of the climate pattern into sub-grids, and uses these sub-grids to describe the clouds that cannot be accurately described at the grid scale of the original climate pattern. At this time, the sub-grids are only divided into two situations: cloudy (1) and cloudless (0), and there is no fractional cloud amount anymore. The random cloud generator determines the cloud situation of each sub-grid according to the input cloud amount profile and the anti-correlation thickness L. Finally, based on the cloud situation of each sub-grid, the cloud structure within a grid of the climate pattern can be obtained. However, in order to find the appropriate anti-correlation thickness L, the random cloud generator continuously traverses a range (usually [0,7]) of numbers, and then finds the most suitable one from these numbers as the anti-correlation thickness L. This process is very time-consuming.
[0033] As Figure 2 shown, the output method of the anti-correlation thickness of the machine learning model: In this embodiment, an eight-layer multi-layer perceptron model is established, where the first layer is the input layer, there are six layers of neural networks in the middle layer, and the last layer is the output layer. Through this model, the corresponding anti-correlation thickness L is directly calculated according to the input cloud amount profile and the total cloud amount, avoiding the traversal process of the traditional random cloud generator. Only by inputting a vector formed by 104 cloud amount profiles and 1 corresponding total cloud amount, a total of 105 groups of data, the corresponding anti-correlation thickness L can be obtained.
[0034] The data used in this example is from 2007. It is found that among the 125 radar vertical scan layers, the global average sea level height is located at the 105th layer. Therefore, in this example, only the detection information from the top scan layer to the 104th layer is extracted.
[0035] Figure 3 And Figure 4 respectively show the correlation between the output L of the method of the present invention and the output L of the random cloud generator on the training set and the test set. As Figure 3 can be seen, the determination coefficient between the output L of the method of the present invention and the output L of the random cloud generator on the training set is as high as 0.986, while Figure 4 shows that the determination coefficient between the output L of the method of the present invention and the output L of the random cloud generator on the test set also reaches a high value of 0.953.
[0036] Figure 5 and Figure 6 shows the relationship between the cloud amount generated when the anti-correlation thickness L obtained by the method of the present invention is used in a random cloud generator to generate the vertical structure of clouds and the true cloud amount. It can be seen from Figure 5 that when the output L obtained by the method of the present invention on the training set is applied to the random cloud generator, the coefficient of determination between the generated cloud amount and the true cloud amount is as high as 0.999, and the test set result is also as high as 0.998.
[0037] In addition, the present invention specifically analyzes the L calculated by the multi-layer perceptron, explores the relationship between L and the numerical model grid resolution and geographical location (latitude), and constructs the parameterization relationships of L with the grid resolution ( Figure 7 ), L with latitude ( Figure 8 ), and L with the grid resolution and latitude ( Figure 9 ). This parameterization relationship can be directly applied to numerical models and is of great significance for the simulation and prediction of weather and climate.
[0038] In summary, the method of the present invention not only significantly improves the calculation efficiency of the anti-correlation thickness, but also the generated anti-correlation thickness has very high precision and great application value.
[0039] The technical solution of the present invention is not limited to the above embodiments, and all technical solutions obtained by equivalent replacement fall within the scope of protection required by the present invention.
Claims
1. A method for accelerating a stochastic cloud generator based on machine learning, characterized in that, The specific steps are: (1) constructing a machine learning model for generating a large amount of anti-correlated thickness data; the machine learning model is a multi-layer perceptron model; (2) Input the cloud profile data and total cloud cover data obtained from satellite observations into the machine learning model to the corresponding anti-correlation thickness; (3) The resolution and latitude information during satellite observation are combined with the anti-correlation thickness data to obtain a three-tuple data set of resolution, latitude, and anti-correlation thickness; based on the three-tuple data set, a parameterized relationship is established between resolution, latitude, and anti-correlation thickness, specifically a fifth-order polynomial: F(Res,Lat) = c + a1*Res + a2*Lat + a3*Res 2 + a4*Res*Lat + a5*Lat 2 … + a 20 *Lat 5 Among them, Res and Lat are the input resolution and latitude, c is the constant term, a i , and i = 1, 2, …, 20 are the polynomial coefficients; (4) This parameterized relationship is directly applied to the climate numerical model for numerical simulation and prediction of the atmospheric system. The climate numerical model obtains the mean anti-correlation thickness at the resolution and latitude of the simulation area, and uses the mean anti-correlation thickness to directly simulate the cloud overlap in the area.
2. The method for accelerating a random cloud generator based on machine learning according to claim 1, wherein The constructed machine learning model includes input layer, hidden layer and output layer, with a total of eight layers: the first layer is the input layer, with a total of 105 input neurons, which is used to read the total cloud cover and cloud cover profile information; the hidden layer has a total of six layers of neural networks, which receive feature vectors from the input layer, among which the first four layers of neural networks are composed of 128 neurons, the fifth layer of neural network is composed of 64 neurons, and the sixth layer of neural network is composed of 32 neurons; in the hidden layer, the first four layers are to introduce enough neurons to extract enough features for learning and enhance the fitting ability of the neural network, and the last two layers are to gradually aggregate the previously extracted features and aggregate the features into the output layer; the last layer is the output layer, with a total of 1 neuron, which is used to output the final fitting results of all layers, that is, the anti-correlation thickness.
3. The method for accelerating a random cloud generator based on machine learning according to claim 1, wherein The training of the machine learning model uses the mean square error (MSE) as the loss function. By calculating the mean square error between the predicted value and the true value, the degree of difference between the model prediction result and the true value is evaluated, thereby further optimizing the performance of the model. During the training process, the back propagation algorithm is used to calculate the gradient, and the gradient descent method is used to optimize the MSE loss function.
Citation Information
Patent Citations
Cloud physical parameter inversion method based on convolutional neural network
CN113887118A
KR1018556520000B1