Cloud cover rapid generation method and system based on deep neural network
By generating cloud cover using deep neural networks, the problem of computational complexity and low efficiency in traditional satellite cloud simulators is solved, achieving efficient and accurate cloud cover prediction. It is suitable for high-frequency coupling and real-time forecasting, thus improving the efficiency of numerical weather prediction systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-03-31
AI Technical Summary
Traditional satellite cloud simulators are computationally complex and inefficient in cloud cover prediction, which limits the application of high-resolution prediction data and results in high uncertainty in cloud cover simulation.
A rapid cloud cover generation method based on deep neural networks is adopted. By acquiring high-resolution data, recurrent neural networks and fully connected feedforward neural networks are used for feature extraction and nonlinear mapping. Regularization methods are combined to suppress overfitting and generate cloud cover. The cloud cover is then corrected and adjusted using cloud cover anticorrelation thickness, cloud hydrogel anticorrelation thickness, and cloud hydrogel nonuniformity factor.
It significantly improves the speed and efficiency of cloud cover generation, maintains physical consistency and spatial distribution accuracy, and is suitable for high-frequency coupling and real-time computing needs. It performs particularly well in cloud cover estimation for extreme weather in mid-to-high latitude regions, thus improving the operational efficiency of numerical weather prediction systems.
Smart Images

Figure CN121236291B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of meteorological remote sensing and atmospheric simulation technology, and in particular to a method and system for rapid cloud cover generation based on deep neural networks. Background Technology
[0002] Clouds, as an important component of the atmosphere, play a crucial role in the Earth's radiation energy balance, atmospheric water cycle, and the evolution of weather systems. Cloud cover is a fundamental physical quantity reflecting the cloud density in the atmosphere and is an important parameter in numerical weather prediction, climate simulation, and remote sensing products. Accurately obtaining cloud cover not only helps improve the accuracy of weather forecasts and climate predictions but also provides important data for meteorological disaster early warning and environmental monitoring.
[0003] In handling the vertical structure of clouds, traditional methods typically employ rules such as maximum overlap, random overlap, or maximum-random overlap to simulate the superposition relationships between multiple cloud layers. Furthermore, to more realistically recreate the cloud distribution at the subgrid scale, subgrid columns are constructed and Monte Carlo random sampling is combined to form a sample set representing different possible cloud overlaps and structural states. This process, often referred to as a stochastic cloud generator, is one of the key steps in achieving physical consistency and spatial statistical matching in satellite cloud simulators.
[0004] Meanwhile, traditional satellite simulators also require calling numerous radiative transfer models, performing simplified radiative transfer calculations on each subgrid column (or sample cloud column) to obtain observations such as albedo, brightness temperature, and cloud top brightness, which are then converted into cloud cover. Due to the highly nonlinear process, complex model structure, and reliance on predefined subgrid cloud distribution statistical parameters, the simulated cloud cover is inherently uncertain. Furthermore, the random cloud generator requires generating a large number of sample datasets, each of which necessitates calculations using the radiative module, resulting in extremely high computational costs and limiting the use of satellite cloud simulators for high-resolution prediction data. Therefore, although these methods have a sound physical foundation, they suffer from numerous shortcomings in practical applications. This invention proposes a rapid satellite cloud cover generation technique based on deep neural networks for atmospheric models, which addresses the computational complexity and inefficiency of traditional methods while improving the accuracy of cloud cover estimation. Summary of the Invention
[0005] To address the aforementioned shortcomings in existing technologies, the cloud cover rapid generation method and system based on deep neural networks provided by this invention solve the problem that existing technologies require a large amount of computing resources, resulting in low cloud cover prediction efficiency.
[0006] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:
[0007] Firstly, a method for rapid cloud cover generation based on deep neural networks is provided, which includes the following steps:
[0008] S1. Obtain high-resolution data of the study area during the prediction period and preprocess it. The high-resolution data includes atmospheric variables and surface variables.
[0009] S2. Based on the latitude and longitude of the study area, select the optimal subgrid cloud distribution statistical parameters for each latitude and longitude of the study area from the optimal subgrid cloud distribution statistical parameter set.
[0010] S3. Input atmospheric variables into the recurrent neural network of the deep neural network model for feature extraction, and use regularization to suppress overfitting to obtain atmospheric reshaping features;
[0011] S4. The atmospheric reshaping characteristics, optimal subgrid cloud distribution statistical parameters and surface data corresponding to the same grid point at the same latitude and longitude are spliced together and then input into the fully connected feedforward neural network of the deep neural network model for nonlinear mapping to obtain cloud cover.
[0012] The subgrid cloud distribution statistical parameters include cloud cover and anticorrelation thickness L. cf , cloud hydrophobic material resistance related thickness L cw And the non-uniformity factor v of cloud water condensate.
[0013] Furthermore, methods for constructing deep neural network models include:
[0014] S31. Obtain a global high-resolution dataset A for a preset time period. The input of dataset A is high-resolution data, and the output is the cloud cover of the high-resolution data obtained through the cloud simulator.
[0015] S32. Construct a deep neural network that includes a recurrent neural network and a fully connected feedforward neural network, and train the deep neural network using dataset A.
[0016] S33. Obtain a global high-resolution dataset B for any given time period, randomly generate several sets of subgrid cloud distribution statistical parameters, and generate several sets of corrected cloud amounts based on the subgrid cloud distribution statistical parameters and the cloud amount in dataset B.
[0017] S34. Using high-resolution data from dataset B and several sets of subgrid cloud distribution statistical parameters as input, and several sets of corrected cloud amount as output, a fine-tuning dataset is constructed, and all model parameters of the deep neural network trained in step S32 are frozen.
[0018] S35. Fine-tuning the deep neural network with frozen model parameters using a fine-tuning dataset yields a deep neural network model that responds to changes in the statistical parameters of cloud distribution in the subgrid.
[0019] Furthermore, methods for generating corrected cloud cover include:
[0020] S331, Based on cloud cover and related thickness L cf Correct for cloud overlap using the initial cloud cover estimate:
[0021]
[0022] in, The cloud cover is adjusted for the overlapping structure. This is the initial cloud cover estimate for the i-th cloud layer; The thickness between the i-th and (i-1)-th cloud layers; denoted as cloud cover thickness; e is the natural logarithm; L is the total number of cloud layers.
[0023] S332. Estimate cloud cover based on the cloud condensate heterogeneity factor v and the initial cloud cover estimate:
[0024]
[0025] in, Cloud cover adjusted for non-uniformity factor of hydrocondensate in clouds; θ represents the weight of the i-th layer of cloud water relative to the total cloud amount; θ is the scale parameter. The gamma function normalization factor; The non-uniformity factor of hydrocondensate in clouds; Let be the value of the cloud-water path in the i-th layer;
[0026] S333. Generate corrected cloud cover based on the cloud cover adjusted for overlapping structure and the cloud cover adjusted for cloud hydrophobicity factor:
[0027]
[0028] in, To correct cloud cover; α represents the initial cloud coverage estimate; α is the controllable coefficient.
[0029] Furthermore, in step S32, when training the deep neural network, atmospheric variables are input into the recurrent neural network of the deep neural network for feature extraction, and a regularization method is used to suppress overfitting to obtain atmospheric reshaping features; the atmospheric reshaping features are concatenated with the surface data, and then input into the fully connected feedforward neural network of the deep neural network for nonlinear mapping;
[0030] In step S35, when fine-tuning the trained deep neural network, atmospheric variables are input into the recurrent neural network of the trained deep neural network for feature extraction. Regularization is used to suppress overfitting and obtain atmospheric reshaping features. The atmospheric reshaping features are then concatenated with subgrid cloud distribution statistics and surface data, and then input into the fully connected feedforward neural network of the trained deep neural network for nonlinear mapping.
[0031] Furthermore, the recurrent neural network is a four-layer recurrent neural network, with the number of hidden units in each layer gradually decreasing, and a regularization operation to prevent overfitting is set after each layer of the recurrent neural network.
[0032] Furthermore, the method for constructing the optimal set of statistical parameters for the subgrid cloud distribution includes:
[0033] S21. Extract global low-resolution data and its actual cloud cover for any given time period, interpolate the low-resolution data to the spatial resolution of the high-resolution data, set values that exceed the physical range of cloud cover to NaN, and obtain global high-resolution data.
[0034] S22. Randomly select a high-resolution data point of latitude and longitude grid that was not selected, and then randomly generate several individuals in the parameter space of the subgrid cloud distribution statistics as the initial population. Each individual is a subgrid cloud distribution statistics parameter.
[0035] S23. Randomly select three distinct individuals x from the population. r1 x r2 x r3 A mutated individual is generated by constructing a difference vector:
[0036]
[0037] Where F is the differential scaling factor; For individual x r1 x r2 x r3 The generated new mutated individuals;
[0038] S24. Select an individual x from the population. r1 x r2 x r3 Different individuals Adopting individual For variant individuals Perform crossover operations to generate experimental individuals u k test individual u k Values for each dimension for:
[0039]
[0040] in, A random number between [0,1]; For a randomly selected dimension index; for The j-th dimension; The preset probability; For individuals The j-th dimension;
[0041] S25. Compare the high-resolution data of the selected latitude and longitude grid points with the data of the experimental individual u. k Input a deep neural network model to generate simulated cloud cover, and calculate the root mean square error between the simulated cloud cover and the actual cloud cover corresponding to the grid points at the current latitude and longitude.
[0042] S26. Determine whether the root mean square error is less than that of the experimental individual u. k Corresponding individuals If the root mean square error is , then it is used for the experimental individual u. k Replace individuals in the population Otherwise, do not replace;
[0043] S27. Increment the number of iterations by one and determine whether the number of iterations is greater than the preset number. If yes, proceed to step S28; otherwise, return to step S23.
[0044] S28. Select the minimum value of the root mean square error of all test individuals in step S25, and take the test individual corresponding to the minimum value as the optimal subgrid cloud distribution statistical parameter corresponding to the grid point where the current latitude and longitude are located.
[0045] S29. Repeat steps S22 to S28 to obtain the optimal subgrid cloud distribution statistical parameters for all grid points of all latitudes and longitudes, and use all the optimal subgrid cloud distribution statistical parameters to form the optimal subgrid cloud distribution statistical parameter set.
[0046] Furthermore, the atmospheric variables are multi-layered atmospheric variables with 16 pressure layers, including temperature, specific humidity, cloud liquid water, ice water content, and cloud cover; the surface variables include surface temperature, air pressure, wind speed, and altitude; and the high-resolution data is ERA5 data.
[0047] Furthermore, preprocessing is also included for the high-resolution data in the global high-resolution datasets A and B:
[0048] Align all high-resolution data to a uniform latitude and longitude grid and time sequence;
[0049] Perform a mean-removal and standard deviation-based standardization operation on all inputs and outputs in the high-resolution data;
[0050] Temporal linear interpolation and spatial co-interpolation methods are used to fill in missing values in high-resolution data.
[0051] Secondly, a cloud cover rapid generation system based on deep neural networks is provided, which includes:
[0052] The data acquisition module is used to acquire high-resolution data of the study area during the prediction period and to preprocess the data, including atmospheric variables and surface variables.
[0053] The parameter selection module is used to select the optimal subgrid cloud distribution statistical parameters for each grid point at each latitude and longitude of the study area from the optimal subgrid cloud distribution statistical parameter set based on the latitude and longitude of the study area.
[0054] The atmospheric reshaping feature generation module is used to input atmospheric variables into the recurrent neural network of the deep neural network model for feature extraction, and to use regularization methods to suppress overfitting and obtain atmospheric reshaping features.
[0055] The cloud cover generation module is used to stitch together atmospheric reshaping features, optimal subgrid cloud distribution statistical parameters and surface data corresponding to the same grid point at the same latitude and longitude. Then, it is input into the fully connected feedforward neural network of the deep neural network model for nonlinear mapping to obtain cloud cover.
[0056] The subgrid cloud distribution statistical parameters include cloud cover and anticorrelation thickness L. cf , cloud hydrophobic material resistance related thickness L cw And the non-uniformity factor v of cloud water condensate.
[0057] The beneficial effects of this invention are as follows: This solution, based on a deep neural network model, can be trained and recognized end-to-end. It can directly learn the implicit mapping relationship of cloud cover from multi-layer atmospheric variables and surface information, avoiding the cumbersome physical parameterization process and greatly improving the speed and efficiency of cloud cover generation. The deep learning-based rapid cloud cover generation technology can not only maintain high physical consistency and spatial distribution accuracy, but also effectively support the high-frequency coupling and real-time calculation requirements in high-resolution atmospheric models, meeting the requirements of modern numerical weather prediction systems for high spatiotemporal resolution and high computational efficiency of cloud cover data.
[0058] This approach offers advantages in cloud cover generation, including fewer input variables, a simpler model structure, and higher efficiency. While maintaining physical consistency and simulation accuracy, it significantly reduces the computational time cost of cloud cover. Compared to traditional satellite simulators, this method avoids the cumbersome process of cloud microphysical parameterization and multi-layered cloud structure reconstruction, significantly improving computational speed and stability. It is more suitable for high-frequency coupling and real-time rapid forecasting applications, particularly excelling in cloud cover estimation for extreme weather events such as severe summer convection, heavy rain, and typhoons in mid-to-high latitude regions. This effectively enhances the operational efficiency and intelligence of numerical weather prediction systems, providing a highly reliable, efficient, and scalable new technological path for climate trend analysis, extreme weather research, radiative transfer simulation, and numerical model evaluation. Attached Figure Description
[0059] Figure 1 This is a flowchart of a method for rapidly generating cloud cover based on deep neural networks.
[0060] Figure 2 This is a block diagram illustrating the principle of a cloud data rapid generation system based on deep neural networks.
[0061] Figure 3 A comparison of the spatial distribution of cloud cover based on satellite observations (left) and cloud cover predicted by a pre-trained deep neural network (right).
[0062] Figure 4 A comparison of the spatial distribution of cloud cover based on satellite observations (left) and cloud cover predicted by a fine-tuned deep neural network model (right). Detailed Implementation
[0063] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.
[0064] refer to Figure 1 , Figure 1 A flowchart of a method for rapid cloud cover generation based on deep neural networks is shown; for example... Figure 1 As shown, the method S includes steps S1 to S4.
[0065] In step S1, high-resolution data of the study area during the prediction period is acquired and preprocessed. The high-resolution data includes atmospheric variables and surface variables. The atmospheric variables are multi-layer atmospheric variables of 16 pressure layers, including temperature, specific humidity, cloud liquid water, ice water content and cloud cover. The surface variables include surface temperature, air pressure, wind speed and terrain height. The high-resolution data is ERA5 data.
[0066] One approach to preprocessing high-resolution data is to first align all data to a uniform latitude and longitude grid (0.25°) and time period (UTC hours), and then perform standardization operations on the high-resolution data by removing the mean and dividing by the standard deviation to eliminate dimensional differences and seasonal fluctuations.
[0067] In step S2, based on the latitude and longitude of the study area, the optimal subgrid cloud distribution statistical parameters for each latitude and longitude of the study area are selected from the optimal subgrid cloud distribution statistical parameter set. In this scheme, the optimal subgrid cloud distribution statistical parameter set stores the optimal subgrid cloud distribution statistical parameters for each grid point under each latitude and longitude globally.
[0068] In step S3, atmospheric variables are input into the recurrent neural network of the deep neural network model for feature extraction, and a regularization method is used to suppress overfitting to obtain atmospheric reshaping features. The recurrent neural network is a four-layer recurrent neural network, with the number of hidden units in each layer gradually decreasing, and a regularization operation to prevent overfitting is set after each layer of the recurrent neural network.
[0069] In step S4, atmospheric reshaping features, optimal subgrid cloud distribution statistics, and surface data corresponding to the same grid point at the same latitude and longitude are spliced together, and then input into the fully connected feedforward neural network of the deep neural network model for nonlinear mapping to obtain cloud cover.
[0070] The subgrid cloud distribution statistical parameters include cloud cover and anticorrelation thickness L. cf , cloud hydrophobic material resistance related thickness L cw And the non-uniformity factor v of cloud water condensate.
[0071] In this scheme, the main input of the deep neural network model is a time-series meteorological variable sequence with shape (16, 5), and the other is an auxiliary static or semi-static variable with shape (1, 4). The main input is first passed through a four-layer recurrent neural network, with each layer having a gradually decreasing number of hidden units (256→128→64→32). Regularization is applied after each layer to prevent overfitting. This structure can progressively extract and compress long-term dependency information in the time series.
[0072] After the last recurrent neural network layer, the output is reshaped into a two-dimensional tensor, flattening the time and feature dimensions into a single vector. This vector is then concatenated with the auxiliary input and optimal parameters, integrating the main sequence features and additional auxiliary information. The concatenated vector is expanded using Flatten and then enters a fully connected layer, where a non-linear transformation is performed using the tanh activation function (Dense(516, tanh activation)). Finally, the linear prediction result is output.
[0073] The output dimension is the cloud cover distribution corresponding to n pressure layers × m optical thickness (tau) categories. This output method can characterize the percentage distribution of cloud cover at different pressure altitudes and different optical thickness ranges in detail. The final total cloud cover can be obtained by summing these n×m category cloud cover values.
[0074] In one embodiment of the present invention, the method for constructing a deep neural network model includes:
[0075] S31. Obtain a global high-resolution dataset A for a preset time period. The input of dataset A is high-resolution data, and the output is the cloud cover of the high-resolution data obtained through the cloud simulator.
[0076] S32. Construct a deep neural network that includes a recurrent neural network and a fully connected feedforward neural network, and train the deep neural network using dataset A.
[0077] S33. Obtain a global high-resolution dataset B for any given time period, randomly generate several sets of subgrid cloud distribution statistical parameters, and generate several sets of corrected cloud amounts based on the subgrid cloud distribution statistical parameters and the cloud amount in dataset B.
[0078] Among them, cloud cover resistance thickness L cf The value range is 1000–10000, and the thickness L of the cloud hydrophobic material is related to the thickness. cw The values are set to 500–5000, and the non-uniformity factor v of cloud hydrocondensates is set to 0.5–5. The random sampling process of the three parameters follows a logarithmic distribution, meaning that it tends to explore different orders of magnitude of parameter combinations in the parameter space, thereby covering more possibilities of structural changes.
[0079] S34. Using high-resolution data from dataset B and several sets of subgrid cloud distribution statistical parameters as input, and several sets of corrected cloud amount as output, a fine-tuning dataset is constructed, and all model parameters of the deep neural network trained in step S32 are frozen.
[0080] S35. Fine-tuning the deep neural network with frozen model parameters using a fine-tuning dataset yields a deep neural network model that responds to changes in the statistical parameters of cloud distribution in the subgrid.
[0081] When fine-tuning a deep neural network, in addition to splicing auxiliary inputs, a new structural parameter input layer is added, namely the three-dimensional structural adjustment parameter (L... cf L cw The three parameters (v, y, v) are used as the third input to the model. After being standardized, these three parameters are input into the network and merged with the original concatenated vector to form an expanded vector of length 517 (originally 516).
[0082] The concatenated vector is then fed into a fully connected layer (Dense(517, tanh activation)) for nonlinear transformation, thereby achieving deep fusion of the original time-series meteorological features, surface information, and structural perturbation signals. The final output layer remains linear, used to generate cloud cover distribution outputs corresponding to an n×m dimension, where n represents the number of vertical pressure layers and m represents different optical thickness (tau) categories. The final total cloud cover can be obtained by summing these n×m categorized cloud cover values.
[0083] The high-resolution data in both datasets A and B are derived from ERA5, covering multiple layers of atmospheric state variables (such as specific humidity, temperature, cloud water content, ice water content, cloud cover, etc.) and surface variables (such as surface air pressure, surface temperature, topographic height, land-sea mask, etc.), with high spatial (0.25°) and temporal (1 hour) resolution.
[0084] In steps S31 and S35, the high-resolution data in the global high-resolution datasets A and B also include preprocessing: aligning all high-resolution data to a uniform latitude and longitude grid and time sequence; performing mean removal and standard deviation standardization operations on the inputs and outputs of all high-resolution data; and filling in missing values in the high-resolution data using temporal linear interpolation and spatial co-interpolation methods.
[0085] Temporal linear interpolation and spatial co-interpolation methods are used to fill in missing values, ensuring the integrity and density of the input tensor, thereby providing a stable data foundation for subsequent model training and reducing systematic errors caused by missing values.
[0086] In this scheme, each group (L) cf , L cw The parameter combinations (v) will be substituted into the pre-trained cloud simulator structure, and combined with the fixed input variable x (atmospheric and surface state variables from ERA5 reanalysis data), a corresponding set of simulated cloud amount outputs y will be generated. In the pre-trained cloud simulator, the methods for generating corrected cloud amount include:
[0087] S331, Based on cloud cover and related thickness L cf Correct for cloud overlap using the initial cloud cover estimate:
[0088]
[0089] in, The cloud cover is adjusted for the overlapping structure. This is the initial cloud cover estimate for the i-th cloud layer; The thickness between the i-th and (i-1)-th cloud layers; denoted as cloud cover thickness; e is the natural logarithm; L is the total number of cloud layers.
[0090] S332. Estimate cloud cover based on the cloud condensate heterogeneity factor v and the initial cloud cover estimate:
[0091]
[0092] in, Cloud cover adjusted for non-uniformity factor of hydrocondensate in clouds; θ represents the weight of the i-th layer of cloud water relative to the total cloud amount; θ is the scale parameter. The gamma function normalization factor; The non-uniformity factor of hydrocondensate in clouds; Let be the value of the cloud-water path in the i-th layer;
[0093] S333. Generate corrected cloud cover based on the cloud cover adjusted for overlapping structure and the cloud cover adjusted for cloud hydrophobicity factor:
[0094]
[0095] in, To correct cloud cover; α represents the initial cloud coverage estimate; α is the controllable coefficient.
[0096] The generated y, obtained through the above method, is no longer a unique value due to the different cloud distribution statistical parameters for each sub-grid. Instead, it represents different cloud cover outputs obtained from different sub-grid cloud distribution statistical parameters. Therefore, the cloud simulator acts as a "structure response generator" at this stage. y is calculated by the model using different cloud cover values calculated through sampling paths with different sub-grid cloud distribution statistical parameters, while keeping x constant. Ultimately, a structure response-type sample library, i.e., a fine-tuning dataset, can be constructed. Here, x is the constant atmospheric input, and y is the cloud cover output by the model under different structural parameter perturbations, forming a mapping relationship of "sub-grid cloud distribution statistical parameters → cloud cover response". Subsequent training tasks are based on this sample library. The trained deep neural network, with a fixed input x, learns how a set of sub-grid cloud distribution statistical parameters affects the cloud cover prediction results, thereby realizing the model's perception and response characteristics to structural perturbations, resulting in a deep neural network model. Through this fine-tuning strategy, the model not only maintains its original time-series learning ability but also possesses structural adjustability.
[0097] In implementation, in the preferred step S32 of this scheme, when training the deep neural network, atmospheric variables are input into the recurrent neural network of the deep neural network for feature extraction. Regularization is used to suppress overfitting and obtain atmospheric reshaping features. The atmospheric reshaping features are then concatenated with the surface data and input into the fully connected feedforward neural network of the deep neural network for nonlinear mapping.
[0098] In step S35, when fine-tuning the trained deep neural network, atmospheric variables are input into the recurrent neural network of the trained deep neural network for feature extraction. Regularization is used to suppress overfitting and obtain atmospheric reshaping features. The atmospheric reshaping features are then concatenated with subgrid cloud distribution statistics and surface data, and then input into the fully connected feedforward neural network of the trained deep neural network for nonlinear mapping.
[0099] In one embodiment of the present invention, the method for constructing the optimal set of statistical parameters for subgrid cloud distribution includes:
[0100] S21. Extract global low-resolution data and its actual cloud cover for any given time period, interpolate the low-resolution data to the spatial resolution of the high-resolution data, set values that exceed the physical range of cloud cover to NaN, and obtain global high-resolution data.
[0101] In step S21, a low-resolution version consistent with the high-resolution data structure is first extracted directly from the ERA5 reanalysis dataset. ERA5 itself provides multiple resolution data download options, such as 0.5°×0.5° or 1°×1° low-resolution products corresponding to the original 0.25°×0.25° high-resolution data. By using the low-resolution version of ERA5, the physical consistency and temporal alignment of the input data can be ensured, avoiding additional errors introduced by the downsampling method. This step aims to construct a test path of "downsampling input → high-resolution estimation" to test the model's scale generalization ability, i.e., whether the model can still output accurate cloud cover estimates under coarser input conditions. At the same time, the output corresponding to the low-resolution input data is also directly extracted from the ERA5 cloud cover product and used as a benchmark reference in the model's low-precision mode.
[0102] The interpolation method in step S21 is a linear extrapolation method. This method not only performs accurate linear interpolation within the known observation range but also achieves stable extrapolation in edge regions. That is, in areas where the ERA5 grid exceeds the observation range, the algorithm can still calculate approximate values based on boundary trends, enhancing the spatial integrity of the data field. After interpolation, to ensure the physical rationality of the data, all outliers exceeding the physical value range (i.e., cloud cover less than 0% or greater than 100%) are uniformly set to NaN. This processing not only avoids numerical contamination in subsequent training and validation phases but also improves the model's robustness in simulating real cloud field structures.
[0103] S22. Randomly select a high-resolution data point of latitude and longitude grid that was not selected, and then randomly generate several individuals in the parameter space of the subgrid cloud distribution statistics as the initial population. Each individual is a subgrid cloud distribution statistics parameter.
[0104] S23. Randomly select three distinct individuals x from the population. r1 x r2 x r3 A mutated individual is generated by constructing a difference vector:
[0105]
[0106] Where F is the differential scaling factor; For individual x r1 x r2 x r3 The generated new mutated individuals;
[0107] S24. Select an individual x from the population. r1 x r2 x r3 Different individuals Adopting individual For variant individuals Perform crossover operations to generate experimental individuals u k test individual u k Values for each dimension for:
[0108]
[0109] in, A random number between [0,1]; For a randomly selected dimension index; for The j-th dimension; The preset probability; For individuals The j-th dimension;
[0110] To facilitate understanding of step S24, a small example will be used below for illustration:
[0111] Suppose an individual's parameters have 3 dimensions: the selected x... k =[1000,2000,3.0], generated through differential mutation. =[1500,2200,4.0], further CR = 0.7, randomly generated rand=[0.3,0.9,0.5], j rand =1, then we have:
[0112] j=1: rand1=0.3≤CR → Take v k1 =1500;
[0113] j=2: rand2=0.9>CR, but j≠rand j →Take x k2 =2000;
[0114] j=3: rand3=0.5≤CR → Take v k3 =4.0;
[0115] The generated test individual u k =[1500,2000,4.0].
[0116] S25. Compare the high-resolution data of the selected latitude and longitude grid points with the data of the experimental individual u. k Input a deep neural network model to generate simulated cloud cover, and calculate the root mean square error between the simulated cloud cover and the actual cloud cover corresponding to the grid points at the current latitude and longitude.
[0117]
[0118] Where T is the length of any time period; Let t be the actual cloud cover corresponding to the grid points of latitude and longitude at time t; This is the simulated cloud cover at grid points with latitude and longitude at time t, predicted by a deep neural network model.
[0119] S26. Determine whether the root mean square error is less than that of the experimental individual u. k Corresponding individuals If the root mean square error is , then it is used for the experimental individual u. k Replace individuals in the population Otherwise, do not replace;
[0120] S27. Increment the number of iterations by one and determine whether the number of iterations is greater than the preset number. If yes, proceed to step S28; otherwise, return to step S23.
[0121] S28. Select the minimum value of the root mean square error corresponding to all test individuals in step S25, and take the test individual corresponding to the minimum value as the optimal subgrid cloud distribution statistical parameter corresponding to the grid point of the current latitude and longitude.
[0122] S29. Repeat steps S22 to S28 to obtain the optimal subgrid cloud distribution statistical parameters for all grid points of all latitudes and longitudes, and use all the optimal subgrid cloud distribution statistical parameters to form the optimal subgrid cloud distribution statistical parameter set.
[0123] This scheme uses the above method to find the optimal subgrid cloud distribution statistics parameters for each grid point under each latitude and longitude, thereby achieving the global optimal determination of key subgrid cloud distribution statistics parameters, significantly improving the prediction accuracy and spatial structure consistency of the model at each latitude and longitude grid point, and enhancing the accuracy and reliability of cloud cover estimation.
[0124] This scheme obtains the optimal subgrid cloud distribution statistics at each latitude and longitude grid point, with each grid point being processed independently. This approach enables the model to adaptively adjust to spatially heterogeneous structures, making cloud cover estimation results closer to actual observations on a global or regional scale, thus improving simulation accuracy and geographical consistency. Furthermore, this method possesses good scalability and parallelizability, making it suitable for batch correction and optimization of large-scale, high-resolution cloud cover data fields.
[0125] like Figure 2 As shown, this solution also provides a cloud cover rapid generation system based on deep neural networks, which includes:
[0126] The data acquisition module is used to acquire high-resolution data of the study area during the prediction period and to preprocess the data, including atmospheric variables and surface variables.
[0127] The parameter selection module is used to select the optimal subgrid cloud distribution statistical parameters for each grid point at each latitude and longitude of the study area from the optimal subgrid cloud distribution statistical parameter set based on the latitude and longitude of the study area.
[0128] The atmospheric reshaping feature generation module is used to input atmospheric variables into the recurrent neural network of the deep neural network model for feature extraction, and to use regularization methods to suppress overfitting and obtain atmospheric reshaping features.
[0129] The cloud cover generation module is used to stitch together atmospheric reshaping features, optimal subgrid cloud distribution statistical parameters and surface data corresponding to the same grid point at the same latitude and longitude. Then, it is input into the fully connected feedforward neural network of the deep neural network model for nonlinear mapping to obtain cloud cover.
[0130] The subgrid cloud distribution statistical parameters include cloud cover and anticorrelation thickness L. cf , cloud hydrophobic material resistance related thickness L cw And the non-uniformity factor v of cloud water condensate.
[0131] To evaluate the effectiveness of the proposed subgrid cloud distribution statistical parameter fine-tuning mechanism in global cloud cover estimation, this scheme compares and analyzes the differences between the cloud cover distribution generated by the model before and after fine-tuning and satellite observation data. Figure 3 and Figure 4 In the image, the left side shows the actual observed cloud cover (OBS), and the right side shows the model output before and after fine-tuning.
[0132] from Figure 3 It can be seen that a deep neural network without fine-tuning ( Figure 3 The model (marked "Model" on the right) exhibits significant biases in several key regions. For example, cloud cover is systematically underestimated in the equatorial Pacific, South America, and the waters surrounding Antarctica, while mid-latitude landmasses in the Northern Hemisphere show excessively smoothed cloud cover. The root mean square error (RMSE) between this version of the model and the observed data is 8.46, indicating that there are still significant discrepancies in terms of spatial structure and intensity.
[0133] In this scheme, the differential evolution algorithm is introduced to L cf L cw After grid-by-grid optimization and adjustment of the three structural parameters (v), the cloud cover output by the fine-tuned model is ( Figure 4 On the right, the model labeled "Fine-tuned" shows greater consistency with the overall structure of the observed results, particularly in the tropical convection zone, mid-latitude ocean regions, and polar margins. The fine-tuned model's RMSE is significantly reduced to 4.07, almost halved, demonstrating the clear advantage of the proposed structural response mechanism in improving model accuracy and spatial consistency.
[0134] The above analysis shows that structural parameter disturbance control (L... cf L cw The joint optimization strategy of combining the algorithm with the differential evolution algorithm can not only capture the local cloud structure characteristics, but also significantly improve the model's performance on a global scale, further verifying the adaptability and reliability of this scheme in the simulation of complex cloud systems.
Claims
1. A method for generating cloud cover based on deep neural network, comprising the steps of: S1, obtaining high-resolution data of a study area in a prediction period and preprocessing the data, wherein the high-resolution data comprises atmospheric variables and surface variables; S2, selecting optimal secondary grid cloud distribution statistical parameters of each latitude and longitude of the study area from a set of optimal secondary grid cloud distribution statistical parameters according to the latitude and longitude of the study area; S3, inputting the atmospheric variables into a recurrent neural network of a deep neural network model to extract features, and using a regularization method to suppress overfitting to obtain atmospheric remodeling features; S4, splicing the atmospheric remodeling features, the optimal secondary grid cloud distribution statistical parameters and the surface data corresponding to the same grid point of the same latitude and longitude, and then inputting them into a fully connected feedforward neural network of the deep neural network model to perform nonlinear mapping to obtain the cloud cover; said subgrid cloud distribution statistical parameters include cloud amount anti-correlation thickness L cf , cloud water anti-correlation thickness L cw and cloud water inhomogeneity factor v; The method for constructing the deep neural network model comprises: S31, obtaining a global high-resolution data set A in a preset period, wherein the input of the data set A is high-resolution data, and the output is the cloud cover of the high-resolution data obtained by a cloud simulator; S32, constructing a deep neural network comprising a recurrent neural network and a fully connected feedforward neural network, and training the deep neural network using the entire data set A; S33, obtaining a global high-resolution data set B in any period, randomly generating a plurality of sets of secondary grid cloud distribution statistical parameters, and generating a plurality of sets of corrected cloud cover according to the secondary grid cloud distribution statistical parameters and the cloud cover in the data set B; S34, using the high-resolution data of the data set B and the plurality of sets of secondary grid cloud distribution statistical parameters as inputs, and the plurality of sets of corrected cloud cover as outputs to form a fine-tuning data set, and freezing all model parameters of the deep neural network trained in step S32; S35, fine-tuning the deep neural network with frozen model parameters using the fine-tuning data set to obtain a deep neural network model in which the predicted cloud cover responds to changes in the secondary grid cloud distribution statistical parameters; The method for generating the corrected cloud cover comprises: S331. Correcting the cloud cover overlap from the cloud amount anti-correlation thickness L cf and the initial cloud cover estimate, the cloud cover overlap is corrected: wherein, is the adjusted cloud cover for the overlapping structure; is the initial cloud cover estimate for the i-th cloud layer; is the thickness between the i-th and i-1-th cloud layers; is the cloud cover anti-correlation thickness; e is the natural logarithm; L is the total number of cloud layers; S332, estimating the cloud cover according to a cloud hydrometeor non-uniformity factor v and an initial cloud coverage estimate value; wherein, is the cloud amount adjusted for cloud water inhomogeneity factor; is the weight of the i-th layer of cloud water to the total cloud amount; θ is the scale parameter; is the gamma function normalization factor; is the cloud water inhomogeneity factor; is the value of the i-th layer of cloud water path; S333, generating the corrected cloud cover according to the cloud cover adjusted according to the overlapping structure and the cloud hydrometeor non-uniformity factor adjusted cloud cover: wherein, is the correction of cloud cover; is the initial cloud cover estimate; and a is a controllable coefficient. 2.The method for generating cloud cover based on deep neural network according to claim 1, wherein when training the deep neural network in step S32, the atmospheric variables are input into the recurrent neural network of the deep neural network to extract features, a regularization method is used to suppress overfitting to obtain atmospheric remodeling features, and the atmospheric remodeling features and the surface data are spliced and then input into the fully connected feedforward neural network of the deep neural network to perform nonlinear mapping; When fine-tuning the trained deep neural network in step S35, the atmospheric variables are input into the recurrent neural network of the trained deep neural network to extract features, a regularization method is used to suppress overfitting to obtain atmospheric remodeling features, and the atmospheric remodeling features, the secondary grid cloud distribution statistical parameters and the surface data are spliced and then input into the fully connected feedforward neural network of the trained deep neural network to perform nonlinear mapping.
3. The cloud cover fast generation method based on deep neural network according to claim 1, wherein the recurrent neural network is a four-layer recurrent neural network, the number of hidden units of each layer of the recurrent neural network gradually decreases, and a regularization operation for preventing overfitting is arranged after each layer of the recurrent neural network.
4. The cloud cover fast generation method based on deep neural network according to any one of claims 1-3, wherein the method for constructing the set of optimal secondary grid cloud distribution statistical parameters comprises: S21, extracting global low-resolution data and its true cloud cover in any time period, and interpolating the low-resolution data to the spatial resolution of high-resolution data to obtain global high-resolution data; S22, randomly selecting a high-resolution data of an unselected latitude and longitude grid point, and then randomly generating a plurality of individuals in the parameter space of the secondary grid cloud distribution statistical parameter as an initial population, each individual being a secondary grid cloud distribution statistical parameter; S27, adding one to the iteration number, and determining whether the iteration number is greater than a preset number, if yes, proceeding to step S28, otherwise returning to step S23; S28, selecting the minimum value in the root mean square errors corresponding to all test individuals in step S25, and taking the test individual corresponding to the minimum value as the optimal secondary grid cloud distribution statistical parameter corresponding to the grid point at the current latitude and longitude; S29, repeating steps S22-S28 to obtain the optimal secondary grid cloud distribution statistical parameters of all grid points at all latitudes and longitudes, and using all the optimal secondary grid cloud distribution statistical parameters to form the set of optimal secondary grid cloud distribution statistical parameters.
5. The cloud cover fast generation method based on deep neural network according to any one of claims 1-3, wherein the atmospheric variables are 16 pressure layer multi-layer atmospheric variables, including temperature, specific humidity, cloud liquid water, ice water content and cloud cover, and the surface variables include surface temperature, air pressure, wind speed and terrain height; and the high-resolution data is ERA5 data.
6. The cloud cover fast generation method based on deep neural network according to any one of claims 1-3, wherein the high-resolution data in the global high-resolution data sets A and B further comprises preprocessing: aligning all high-resolution data to a unified latitude and longitude grid and time; performing a standardization operation of removing mean value and dividing by standard deviation on the input and output of all high-resolution data; and filling in missing values in the high-resolution data by using time series linear interpolation and spatial collaborative interpolation methods. S23, randomly select three different individuals x from the population r1 , x r2 , x r3 , generate a mutated individual by constructing a difference vector: where F is a difference scaling factor; is the individual x r1 , x r2 , x r3 generated new mutated individual; S24. Select an individual x from the population. r1 x r2 x r3 Different individuals Adopting individual For variant individuals Perform crossover operations to generate experimental individuals u k test individual u k Values for each dimension for: wherein, is a random number between 0 and 1 ; is a randomly chosen dimension index; is the jth dimension of is a predetermined probability; is the jth dimension of the individual is the jth dimension of the individual S25, high-resolution data of the selected latitude and longitude grid point is combined with the test individual u k The input deep neural network model generates a simulated cloud amount, and calculates a root mean square error between the simulated cloud amount and the real cloud amount corresponding to the grid point of the current latitude and longitude. S26, determine whether the root mean square error is less than the test individual u k corresponding individual of the root mean square error, if yes, replace the individual in the test individual u k population , otherwise not replace; 7. The cloud cover fast generation method based on deep neural network according to any one of claims 1-3, comprising: a data acquisition module for acquiring high-resolution data of a research area in a prediction time period and preprocessing the high-resolution data, the high-resolution data including atmospheric variables and surface variables; a parameter selection module for selecting, according to the latitude and longitude of the research area, the optimal secondary grid cloud distribution statistical parameter of each grid point at each latitude and longitude of the research area from the set of optimal secondary grid cloud distribution statistical parameters; and an atmospheric remodeling feature generation module for inputting the atmospheric variables into the recurrent neural network of the deep neural network model for feature extraction, and using a regularization method to suppress overfitting to obtain atmospheric remodeling features. 7. A generating system applied to the cloud cover fast generation method based on a deep neural network according to any one of claims 1-6, characterized in that, The cloud cover generation module is used for splicing the atmospheric remodeling features, the optimal secondary grid cloud distribution statistical parameters and the ground data corresponding to the same grid point of the same latitude and longitude, and then inputting a full connection feedforward neural network of a deep neural network model for nonlinear mapping to obtain the cloud cover. The subgrid cloud distribution statistical parameters include cloud amount anti-correlation thickness L cf , cloud water anti-correlation thickness L cw , and cloud water inhomogeneity factor v.