A method for predicting supercooled water content in clouds based on deep learning and spatiotemporal data

Through deep learning and spatiotemporal data-driven methods, combined with multi-headed attention mechanism and Kuyu algorithm to optimize network parameters, ConvLSTM-Transformer hybrid network is built, which solves the uncertainty of forecasting supercooled water content in the cloud, achieves high-precision prediction, and improves the prediction accuracy and timeliness of supercooled water content in the cloud.

CN120255027BActive Publication Date: 2025-09-02辽宁省人工影响天气办公室
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510733813.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-02
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

The prior art has uncertainties and deviations in the forecast of supercooled water content in the cloud, making it difficult to achieve accurate predictions, especially in strong convective weather, and aircraft on-board detection equipment and foundation observations cannot achieve extrapolated forecasts of horizontal ranges.

Method used

Using a method based on deep learning and spatiotemporal data-driven, a convolutional neural network is used to extract features of multidimensional live meteorological data, combine multi-headed attention mechanism for feature fusion, and optimize network parameters through the Kuyu algorithm to construct a ConvLSTM-Transformer hybrid network for prediction of supercooled water distribution in the cloud in the next 3 hours.

Benefits of technology

It improves the accuracy and timeliness of the prediction of supercooled water content, provides scientific and intelligent decision-making support, alleviates water resource shortage, improves the ecological environment and ensures agricultural, electricity and aviation safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255027B_ABST
    Figure CN120255027B_ABST
Patent Text Reader

Abstract

This paper discloses a method for predicting supercooled water content in clouds based on deep learning and spatiotemporal data. The method involves acquiring multi-channel satellite radiation transmission data and performing channel screening and feature extraction, acquiring radar reflectivity data and performing data interpolation and feature extraction, acquiring conventional meteorological data from ground-based automatic weather stations and performing feature extraction, fusing these features using a multi-head attention mechanism to generate fused features, and establishing a deep learning neural network model to predict supercooled water content in clouds for future time periods. This method enables more comprehensive and accurate prediction of supercooled water content and distribution in clouds, and has important scientific and practical significance for alleviating water resource shortages, improving the ecological environment, and ensuring economic security at low altitudes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of supercooled water prediction, and in particular to a method for predicting supercooled water content in clouds based on deep learning and spatiotemporal data driving. Background Art

[0002] Supercooled water refers to liquid water with a temperature below 0°C but not frozen. Accurately grasping and accurately predicting the supercooled water content in clouds is of great significance to weather forecasting, agricultural production, ecological environmental restoration, power safety and aviation safety. With the continuous advancement of meteorological observation technology, the amount of data generated by various observation methods has exploded. These data contain rich temporal and spatial information, providing new opportunities for predicting the supercooled water content in clouds.

[0003] Currently, the main means of forecasting supercooled water is numerical models. Due to the settings of the initial field and parameters, the forecast results of the models have some uncertainty. In particular, in some severe convective weather, the models are difficult to accurately grasp the development and evolution of clouds. At the same time, although aircraft-borne detection equipment and ground-based observations can provide real-time detection of supercooled water content along the route or at a single point, they cannot achieve extrapolated forecasts of the horizontal range, resulting in biased evaluation results and the inability to accurately predict the supercooled water content in the cloud. Therefore, the present invention proposes a method for predicting supercooled water content in clouds based on deep learning and spatiotemporal data. A convolutional neural network is used to extract features from multi-dimensional real-time meteorological data, and a multi-head attention mechanism is used for feature fusion. A deep learning neural network is established, and the bitter fish algorithm is used to optimize the parameters during the network training process to achieve a prediction of supercooled water distribution in clouds for the next 3 hours. This method effectively integrates satellite, radar and conventional meteorological data, extracts and fuses the characteristics of these data, efficiently captures the supercooled water information in clouds contained in multi-source observation data, and effectively improves the accuracy of supercooled water content prediction. It has important scientific and practical significance for alleviating water resource shortages, improving the ecological environment, and ensuring agricultural production safety, power safety and aviation safety. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for predicting supercooled water content in clouds based on deep learning and spatiotemporal data driving.

[0005] To achieve the above object, the present invention is implemented according to the following technical solutions:

[0006] The present invention comprises the following steps:

[0007] Obtain satellite multi-channel radiation data for the past 6 hours, perform calibration, projection conversion, and channel screening on the multi-channel radiation data, and use a first feature extraction model to extract features to obtain satellite features; the multi-channel radiation data dimensions include the number of samples, the number of channels, the number of longitudinal grid points, the number of latitudinal grid points, and the time step;

[0008] Acquire radar data from the past six hours and use the second feature extraction model to extract radar features. The radar data dimensions include the number of features, the number of longitudinal grid points, the number of latitudinal grid points, and the time step.

[0009] Obtain conventional meteorological data observed by ground automatic stations over the past six hours, and use the third feature extraction model to extract features to obtain conventional meteorological features; the conventional meteorological data includes temperature, air pressure, wind direction, wind speed, precipitation, and relative humidity; the dimensions of the conventional meteorological data include the number of features, the number of stations, and the time step;

[0010] Using a multi-head attention mechanism, the satellite features, the radar features, and the conventional meteorological features are subjected to feature fusion to obtain fused features;

[0011] Constructing a deep learning network based on the fusion features and the supercooled water content at the corresponding future time, determining an optimization objective function, and optimizing the deep learning network based on the optimization objective function; the deep learning network is a ConvLSTM-Transformer hybrid network;

[0012] The fused features to be predicted are input into the optimized deep learning network to obtain the predicted value of supercooled water content in the cloud in the next 3 hours;

[0013] The method for performing feature fusion to obtain fusion features comprises the following steps:

[0014] The satellite features, the radar features, and the conventional meteorological features are fused using a multi-head attention mechanism to obtain fused features. The attention mechanism module uses a multi-head attention mechanism to perform feature fusion, and the expression is:

[0015]

[0016]

[0017]

[0018] in is the attention mechanism, is the query vector, is the key vector, is a value vector, is the key vector Dimensions for scaling, For multiple attentions, For the The hidden state of the head, , is the weight matrix used to combine the outputs of each head, 、 、 To input 、 、 Learnable weight matrices mapped to different subspaces;

[0019] The method for determining the optimization objective function is based on the input feature category and the number of features Determine the optimization objective function for the deep learning network:

[0020]

[0021] in To optimize the objective function, is the threshold stability weight, is the eigenvalue bias weight, is the time constraint weight, is the spatial constraint weight, is the threshold type weight, for Class dynamic identification threshold prediction value, for The historical mean of the class threshold is calculated using a sliding window. Grid points middle Weather-like characteristic values, Time penalty coefficient, is the spatial penalty coefficient, For the input feature categories, is the input feature category index.

[0022] Furthermore, the method for obtaining satellite characteristics comprises the following steps:

[0023] Calibration, projection conversion, and channel screening are performed on the multi-channel radiation data to obtain first radiation data; the multi-channel radiation data has data dimensions including the number of samples, the number of channels, the number of longitudinal grid points, the number of latitudinal grid points, and the time step;

[0024] Inputting the first radiation data into a first feature extraction model for feature extraction to obtain satellite features; the first feature extraction model includes an input layer, a convolution layer, a pooling layer, an activation layer, a fully connected layer, and an output layer;

[0025] The input layer is used for receiving the first radiation data;

[0026] The convolution layer uses a 4D convolutional neural network to perform a convolution operation on the input data to capture the local features of the data in space and time to obtain the first spatiotemporal radiation feature; the convolution kernel of the 4D convolutional neural network is 2×3×3×3 or 2×5×5×3, with a step size of 2×2×2×2;

[0027] The pooling layer downsamples the first spatiotemporal radiation feature and simplifies the dimension of the first spatiotemporal radiation feature to obtain a second spatiotemporal radiation feature; the pooling layer performs pooling operations in spatial and temporal dimensions;

[0028] The activation layer uses an activation function to perform a nonlinear transformation on the output of the convolution layer or the pooling layer; the activation layer includes a first activation layer and a second activation layer; the first activation layer receives the output of the convolution layer and outputs it to the pooling layer; the second activation layer receives the output of the pooling layer and outputs it to the fully connected layer;

[0029] The fully connected layer maps the second spatiotemporal radiation features after convolution, pooling and activation operations to the output space. The fully connected layer includes multiple neurons, performs linear transformation on the second spatiotemporal radiation features through weight matrix and bias term, and performs nonlinear processing through activation function to obtain satellite features;

[0030] The output layer connects to the fully connected layer to output satellite features.

[0031] Furthermore, the method for obtaining radar characteristics comprises the following steps:

[0032] Inputting the radar data into a second feature extraction model to extract features to obtain radar features; the radar data dimensions include the number of samples, the number of features, the number of longitudinal grid points, the number of latitudinal grid points, and the time step; the second feature extraction model includes an input layer, a convolution layer, a pooling layer, an activation layer, a fully connected layer, and an output layer;

[0033] The input layer is used to receive radar data;

[0034] The convolution layer uses a 4D convolutional neural network to perform a convolution operation on the input data to capture the local features of the data in space and time to obtain the first spatiotemporal radar features. The convolution kernel of the 4D convolutional neural network is 1×3×3×3 or 1×5×5×3, with a step size of 2×2×2×2.

[0035] The pooling layer downsamples the first spatiotemporal radar feature and simplifies the dimension of the first spatiotemporal radar feature to obtain a second spatiotemporal radar feature; the pooling layer performs pooling operations in spatial and temporal dimensions;

[0036] The activation layer uses an activation function to perform a nonlinear transformation on the output of the convolution layer or the pooling layer; the activation layer includes a first activation layer and a second activation layer; the first activation layer receives the output of the convolution layer and outputs it to the pooling layer; the second activation layer receives the output of the pooling layer and outputs it to the fully connected layer;

[0037] The fully connected layer maps the second spatiotemporal radar features after convolution, pooling, and activation operations to the output space. The fully connected layer includes multiple neurons, performs linear transformation on the second spatiotemporal radar features through weight matrices and bias terms, and performs nonlinear processing through activation functions to obtain satellite features.

[0038] The output layer is connected to the fully connected layer to output radar features.

[0039] Furthermore, the method for obtaining conventional meteorological characteristics comprises the following steps:

[0040] Inputting conventional meteorological data into a third feature extraction model for feature extraction to obtain conventional meteorological features; the conventional meteorological data includes temperature, air pressure, wind direction, wind speed, precipitation, and relative humidity; the dimensions of the conventional meteorological data include the number of features, the number of stations, and the time step; the third feature extraction model includes an input layer, a convolutional layer, a pooling layer, an activation layer, a fully connected layer, and an output layer;

[0041] The input layer is used to receive regular meteorological data;

[0042] The convolution layer uses a 3D convolutional neural network to perform a convolution operation on the input data to capture the local features of the data in space and time to obtain the first spatiotemporal conventional meteorological features; the convolution kernel of the 3D convolutional neural network uses a 2×3×3 or 2×5×3 convolution kernel with a step size of 2×2×2;

[0043] The pooling layer downsamples the first spatiotemporal conventional meteorological data features and simplifies the dimensions of the first spatiotemporal conventional meteorological features to obtain the second spatiotemporal conventional meteorological features; the pooling layer performs pooling operations in the spatial and temporal dimensions;

[0044] The activation layer uses an activation function to perform a nonlinear transformation on the output of the convolution layer or the pooling layer; the activation layer includes a first activation layer and a second activation layer; the first activation layer receives the output of the convolution layer and outputs it to the pooling layer; the second activation layer receives the output of the pooling layer and outputs it to the fully connected layer;

[0045] The fully connected layer maps the second spatiotemporal conventional meteorological features after convolution, pooling and activation operations to the output space. The fully connected layer includes multiple neurons, performs linear transformation on the second spatiotemporal conventional meteorological features through weight matrix and bias term, and obtains conventional meteorological features through nonlinear processing through activation function;

[0046] The output layer is connected to the fully connected layer to output radar features.

[0047] Furthermore, the method for optimizing the deep learning network according to the optimization objective function includes:

[0048] The optimal network model parameters are determined using the bitter fish optimization algorithm, and the definition For the The network model parameters correspond to the spawning points of bitter fish. Input feature dimensions within the past 3 hours, and the bitter fish population is composed of multiple bitter fish individuals , is the population size, and the population is initialized with the standard network model parameters to obtain , where the boundary is , is the random perturbation number;

[0049] Search for suitable oysters to determine the spawning point of bittern, update the position of bittern (update the dynamic network model parameters), and when the oyster is successfully caught, the updated position of bittern is expressed as:

[0050]

[0051]

[0052]

[0053] in For the Fish in the The updated position in the iteration, For the Fish in the The current position in the iteration, is the dynamic inertia weight, For the The number of steps the bitter fish takes to escape the oyster in the iteration, 、 for A random number, For the best oysters, the best spawning spots to attract bitterlings, For the most worthwhile oyster in a randomly selected population, 、 are the maximum inertia weight and the minimum inertia weight, For the The population variance at the iteration is is the initial variance of the population, is the maximum number of iterations, is a random function;

[0054] When the oyster escapes, it will re-explore and the bitter fish will update its position expression as follows:

[0055]

[0056] in is the escape reset probability, take ;

[0057] According to the position of the successfully captured oysters, the female fish lays eggs at that position to produce new individuals. The updated position expression of the new individual bitterling is:

[0058]

[0059] in For the The location of the new bitter fish generated during iteration, is the distribution radius of the newly generated bitter fish inside the oyster shell, The mean is 0 and the variance is Gaussian noise;

[0060] Calculating population fitness to identify the best oysters and the most worthy oysters of the stock , the expression is:

[0061]

[0062]

[0063] in is the population fitness, is the fitness dynamic weight, For the Iteration No. Dynamic network model parameter locations, For the standard network model parameter positions, For the The corresponding position of the maximum value of the dynamic network model parameter, No. The position corresponding to the minimum value of the dynamic network model parameter, is the weight of coordination entropy, For the The coordination entropy of the dynamic recognition threshold, is the time-varying penalty weight, is the time-varying penalty decay rate, For the Iteration No. Dynamic network model parameter positions and target positions relevance;

[0064] The strategy of using oysters to kill newborn bitterlings is used to eliminate individuals in the population. The probability of eliminating newborn bitterlings is:

[0065]

[0066]

[0067] in The probability of elimination of new bitter fish, Selection pressure for the population, For the new bitter fish population Iterative fitness, The new bitterfish population includes new bitterfish and the original bitterfish. Iterative fitness, is the diversity penalty weight, For the Iterated population diversity, the historical average position of the population;

[0068] Repeat the above steps until the objective function is optimized The iteration is stopped when the minimum or maximum number of iterations is reached, the optimal network model parameters are output, and the optimized deep learning network is obtained.

[0069] Furthermore, the satellite characteristics include satellite multi-channel radiation data; the radar characteristics include combined reflectivity and VIL; and the conventional meteorological characteristics include temperature, relative humidity, wind direction, wind speed, and precipitation.

[0070] The beneficial effects of the present invention are:

[0071] The present invention is a method for predicting supercooled water content in clouds based on deep learning and spatiotemporal data. Compared with the existing technology, the present invention has the following technical effects:

[0072] Through feature extraction, feature fusion, model construction, and parameter optimization, this invention enhances data preprocessing capabilities in the dynamic evaluation of cold cloud artificial rainfall potential areas. It can effectively extract and integrate information on supercooled water content in clouds contained in spatiotemporal data, improve data utilization, reduce computational complexity, and enhance the accuracy and timeliness of supercooled water content predictions in clouds. This method provides scientific and intelligent decision-making support for analyzing operational conditions and dynamically identifying potential areas during cold cloud artificial rainfall operations, easing water resource conservation, improving the ecological environment, and ensuring agricultural production safety, power safety, and aviation safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 This is a flowchart of the steps of the method for predicting supercooled water content in clouds based on deep learning and spatiotemporal data driving. DETAILED DESCRIPTION

[0074] The present invention will be further described below through specific examples. The illustrative examples and descriptions of the present invention are used to explain the present invention but are not intended to limit the present invention.

[0075] The method for predicting supercooled water content in clouds based on deep learning and spatiotemporal data driving includes the following steps:

[0076] like Figure 1 As shown, in this embodiment, the following steps are included:

[0077] Obtain satellite multi-channel radiation data for the past 6 hours, perform calibration, projection conversion, and channel screening on the multi-channel radiation data, and use a first feature extraction model to extract features to obtain satellite features; the multi-channel radiation data dimensions include the number of samples, the number of channels, the number of longitudinal grid points, the number of latitudinal grid points, and the time step;

[0078] Acquire radar data from the past six hours and use the second feature extraction model to extract radar features. The radar data dimensions include the number of features, the number of longitudinal grid points, the number of latitudinal grid points, and the time step.

[0079] Obtain conventional meteorological data observed by ground automatic stations over the past six hours, and use the third feature extraction model to extract features to obtain conventional meteorological features; the conventional meteorological data includes temperature, air pressure, wind direction, wind speed, precipitation, and relative humidity; the dimensions of the conventional meteorological data include the number of features, the number of stations, and the time step;

[0080] Using a multi-head attention mechanism, the satellite features, the radar features, and the conventional meteorological features are subjected to feature fusion to obtain fused features;

[0081] Constructing a deep learning network based on the fusion features and the supercooled water content at the corresponding future time, determining an optimization objective function, and optimizing the deep learning network based on the optimization objective function; the deep learning network is a ConvLSTM-Transformer hybrid network;

[0082] The fused features to be predicted are input into the optimized deep learning network to obtain the predicted value of supercooled water content in the cloud in the next 3 hours;

[0083] The method for performing feature fusion to obtain fusion features comprises the following steps:

[0084] The satellite features, the radar features, and the conventional meteorological features are fused using a multi-head attention mechanism to obtain fused features. The attention mechanism module uses a multi-head attention mechanism to perform feature fusion, and the expression is:

[0085]

[0086]

[0087]

[0088] in is the attention mechanism, is the query vector, is the key vector, is a value vector, is the key vector Dimensions for scaling, For multiple attentions, For the The hidden state of the head, , is the weight matrix used to combine the outputs of each head, 、 、 To input 、 、 Learnable weight matrices mapped to different subspaces;

[0089] The method for determining the optimization objective function is based on the input feature category and the number of features Determine the optimization objective function for the deep learning network:

[0090]

[0091] in To optimize the objective function, is the threshold stability weight, is the eigenvalue bias weight, is the time constraint weight, is the spatial constraint weight, is the threshold type weight, for Class dynamic identification threshold prediction value, for The historical mean of the class threshold is calculated using a sliding window. Grid points middle Weather-like characteristic values, Time penalty coefficient, is the spatial penalty coefficient, For the input feature categories, is the input feature category index.

[0092] In this embodiment, the method for obtaining satellite characteristics includes:

[0093] The multi-channel radiation data is calibrated, projected, and channel-filtered to obtain first radiation data; the multi-channel radiation data has dimensions including the number of samples, the number of channels, the number of longitudinal grid points, the number of latitudinal grid points, and the time step.

[0094] In this embodiment, the method for obtaining radar characteristics includes:

[0095] The radar data is input into a second feature extraction model for feature extraction to obtain radar features; the radar data dimensions include the number of samples, the number of features, the number of longitudinal grid points, the number of latitudinal grid points and the time step; the second feature extraction model includes an input layer, a convolution layer, a pooling layer, an activation layer, a fully connected layer and an output layer.

[0096] In this embodiment, the method for obtaining conventional meteorological characteristics includes:

[0097] The conventional meteorological data is input into the third feature extraction model for feature extraction to obtain conventional meteorological features; the conventional meteorological data includes temperature, air pressure, wind direction, wind speed, precipitation and relative humidity; the dimensions of the conventional meteorological data include the number of features, the number of stations and the time step; the third feature extraction model includes an input layer, a convolution layer, a pooling layer, an activation layer, a fully connected layer and an output layer.

[0098] In this embodiment, the method for constructing a deep learning network includes:

[0099] The fused features and the corresponding supercooled water content in the future time are taken as a comprehensive set, which is divided into a training set and a test set according to a ratio of 6:4. The training set is used to train the deep learning network, and the test set is used to evaluate the performance of the deep learning network.

[0100] In this embodiment, the method for optimizing the deep learning network according to the optimization objective function includes:

[0101] The optimal network model parameters are determined using the bitter fish optimization algorithm, and the definition For the The network model parameters correspond to the spawning points of bitter fish. Input feature dimensions within the past 3 hours, and the bitter fish population is composed of multiple bitter fish individuals , is the population size, and the population is initialized with the standard network model parameters to obtain , where the boundary is , is the random perturbation number;

[0102] Search for suitable oysters to determine the spawning point of bittern, update the position of bittern (update the dynamic network model parameters), and when the oyster is successfully caught, the updated position of bittern is expressed as:

[0103]

[0104]

[0105]

[0106] in For the Fish in the The updated position in the iteration, For the Fish in the The current position in the iteration, is the dynamic inertia weight, For the The number of steps the bitter fish takes to escape the oyster in the iteration, 、 for A random number, For the best oysters, the best spawning spots to attract bitterlings, For the most worthwhile oyster in a randomly selected population, 、 are the maximum inertia weight and the minimum inertia weight, For the The population variance at the iteration is is the initial variance of the population, is the maximum number of iterations, is a random function;

[0107] When the oyster escapes, it will re-explore and the bitter fish will update its position expression as follows:

[0108]

[0109] in is the escape reset probability, take ;

[0110] According to the position of the successfully captured oysters, the female fish lays eggs at that position to produce new individuals. The updated position expression of the new individual bitterling is:

[0111]

[0112] in For the The location of the new bitter fish generated during iteration, is the distribution radius of the newly generated bitter fish inside the oyster shell, The mean is 0 and the variance is Gaussian noise;

[0113] Calculating population fitness to identify the best oysters and the most worthy oysters of the stock , the expression is:

[0114]

[0115]

[0116] in is the population fitness, is the fitness dynamic weight, For the Iteration No. Dynamic network model parameter locations, For the standard network model parameter positions, For the The corresponding position of the maximum value of the dynamic network model parameter, No. The position corresponding to the minimum value of the dynamic network model parameter, is the weight of coordination entropy, For the The coordination entropy of the dynamic recognition threshold, is the time-varying penalty weight, is the time-varying penalty decay rate, For the Iteration No. Dynamic network model parameter positions and target positions relevance;

[0117] The strategy of using oysters to kill newborn bitterlings is used to eliminate individuals in the population. The probability of eliminating newborn bitterlings is:

[0118]

[0119]

[0120] in The probability of elimination of new bitter fish, Selection pressure for the population, For the new bitter fish population Iterative fitness, The new bitterfish population includes new bitterfish and the original bitterfish. Iterative fitness, is the diversity penalty weight, For the Iterated population diversity, the historical average position of the population;

[0121] Repeat the above steps until the objective function is optimized Stop iteration when the minimum or maximum number of iterations is reached, output the optimal network model parameters, and obtain the optimized deep learning network;

[0122] In this embodiment, the satellite characteristics include satellite multi-channel radiation data; the radar characteristics include combined reflectivity and VIL; and the conventional meteorological characteristics include temperature, relative humidity, wind direction, wind speed, and precipitation.

[0123] In actual evaluations, to predict the distribution of supercooled water in a specific area at a specific time, the deep learning network model must first be trained using years of historical data. Satellite radiation data (15 channels), weather radar network mosaic data (combined reflectivity and VIL), and conventional meteorological data (temperature, air pressure, relative humidity, wind direction, wind speed, and precipitation) were obtained from 20:00 to 23:00 on July 21, 2024, with a cumulative duration of 3 hours.

[0124] Using the first feature extraction model to extract satellite multi-channel radiation data to obtain satellite features, using the second feature extraction model to extract radar data to obtain radar features, and using the third feature extraction model to extract conventional meteorological data to obtain conventional meteorological features;

[0125] Use multi-head attention mechanism to fuse satellite features, radar features and conventional meteorological features;

[0126] The feature image size (height * width) obtained by the multi-branch CNN network is 4*4, with 4 channels. The channel attention mechanism divides the feature map into two groups (height + temperature, wind speed + humidity). The lightweight multi-layer perceptron can learn the importance weights of the groups to be 0.6 and 0.4 respectively. The weights of color, texture, and shape features after channel attention are 0.40, 0.30, and 0.30 respectively.

[0127] The convolution kernel weights of the multi-head attention mechanism are 3*3 / 0.4, 5*5 / 0.3, and 7*7 / 0.7, respectively. The polar coordinate convolution (radial distance Take 1, angle Take 45°) weight is 0.1, position Variance on Take 0.2, variance adjustment coefficient Taking 0.5, the weights of satellite, radar and conventional meteorological features after attention weighting are 0.30, 0.35 and 0.35 respectively;

[0128] The fully connected module is used to concatenate the weighted features and output the field features corresponding to the three moments;

[0129] Optimize the objective function Mid-threshold stability weight Take 0.3, eigenvalue bias weight Take 0.5, the time constraint weight Take 0.2, the spatial constraint weight Take 0.1, the threshold type weight is (satellite feature 0.30, radar feature 0.35, conventional meteorological feature 0.35), time penalty coefficient Take 0.1, the spatial penalty coefficient Take 0.05;

[0130] Initial population size Take 10 to initialize the network weight parameters, , the maximum inertia weight and the minimum inertia weight are respectively and , maximum number of iterations Take 100 times;

[0131] Taking the optimization of the prediction model parameters in the first prediction period (21:00-23:00 on the 21st) as an example, the prediction step is 3 hours, the lag time step is 6 hours (15:00-20:00 on the 21st), and the target function is optimized when the bitter fish search algorithm is iterated to 56, 57, 58, and 59 times. They are 0.11, 0.09, 0.09, and 0.09 respectively. The optimal position of the population in the 57th iteration is taken as the optimized total population position, and the optimized network weight parameters are output; the fusion features of the past 6 hours are input into the constructed prediction network to obtain the prediction results of the supercooled water content in the next 3 hours.

[0132] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for predicting supercooled water content in clouds based on deep learning and spatiotemporal data, characterized by: The following steps are involved: S1. Obtain satellite multi-channel radiation data for the past 6 hours, perform calibration, projection conversion, and channel screening on the multi-channel radiation data, and use a first feature extraction model to perform feature extraction to obtain satellite features; The multi-channel radiation data dimensions include the number of samples, the number of channels, the number of longitudinal grid points, the number of latitudinal grid points and the time step; S2. Acquire radar data for the past 6 hours and use a second feature extraction model to extract radar features; the radar data dimensions include the number of features, the number of longitudinal grid points, the number of latitudinal grid points, and the time step; S3, obtaining conventional meteorological data observed by the ground automatic station in the past 6 hours, and using the third feature extraction model to perform feature extraction to obtain conventional meteorological features; The conventional meteorological data includes temperature, air pressure, wind direction, wind speed, precipitation and relative humidity; the conventional meteorological data dimensions include the number of features, the number of stations and the time step; S4. Using a multi-head attention mechanism, the satellite features, the radar features, and the conventional meteorological features are subjected to feature fusion to obtain fused features; S5. Construct a deep learning network based on the fusion features and the supercooled water content at the corresponding future time, determine an optimization objective function, and optimize the deep learning network based on the optimization objective function; the deep learning network is a ConvLSTM-Transformer hybrid network; S6. Input the fused features to be predicted into the optimized deep learning network to obtain the predicted value of supercooled water content in the cloud in the next 3 hours; The method for performing feature fusion to obtain fusion features comprises the following steps: The satellite features, the radar features, and the conventional meteorological features are fused using a multi-head attention mechanism to obtain fused features. The attention mechanism module uses a multi-head attention mechanism to perform feature fusion, and the expression is: in is the attention mechanism, is the query vector, is the key vector, is a value vector, is the key vector Dimensions for scaling, For multiple attentions, For the The hidden state of the head, , is the weight matrix used to combine the outputs of each head, 、 、 To input 、 、 Learnable weight matrices mapped to different subspaces; The method for determining the optimization objective function is: according to the input feature category and the number of features Determine the optimization objective function for the deep learning network: in To optimize the objective function, is the threshold stability weight, is the eigenvalue bias weight, is the time constraint weight, is the spatial constraint weight, is the threshold type weight, for Class dynamic identification threshold prediction value, for The historical mean of the class threshold is calculated using a sliding window. Grid points middle Weather-like characteristic values, Time penalty coefficient, is the spatial penalty coefficient, For the input feature categories, is the input feature category index; The method for optimizing the deep learning network according to the optimization objective function comprises the following steps: The optimal network model parameters are determined using the bitter fish optimization algorithm, and the definition For the The network model parameters correspond to the spawning points of bitter fish. Input feature dimensions within the past 3 hours, and the bitter fish population is composed of multiple bitter fish individuals , is the population size, and the population is initialized with the standard network model parameters to obtain , where the boundary is , is the random perturbation number; Search for suitable oysters to determine the spawning point of bitter fish and update the position of bitter fish. When oysters are successfully caught, the updated position of bitter fish is expressed as: in For the Fish in the The updated position in the iteration, For the Fish in the The current position in the iteration, is the dynamic inertia weight, For the The number of steps the bitter fish takes to escape the oyster in the iteration, 、 for A random number, For the best oysters, the best spawning spots to attract bitterlings, For the most worthwhile oyster in a randomly selected population, 、 are the maximum inertia weight and the minimum inertia weight, For the The population variance at the iteration is is the initial variance of the population, is the maximum number of iterations, is a random function; When the oyster escapes, it will re-explore and the bitter fish will update its position expression as follows: in is the escape reset probability, take ; According to the position of the successfully captured oysters, the female fish lays eggs at that position to produce new individuals. The updated position expression of the new individual bitterling is: in For the The location of the new bitter fish generated during iteration, is the distribution radius of the newly generated bitter fish inside the oyster shell, The mean is 0 and the variance is Gaussian noise; Calculating population fitness to identify the best oysters and the most worthy oysters of the stock , the expression is: ; in is the population fitness, is the fitness dynamic weight, For the Iteration No. Dynamic network model parameter locations, For the standard network model parameter positions, For the The corresponding position of the maximum value of the dynamic network model parameter, No. The position corresponding to the minimum value of the dynamic network model parameter, is the weight of coordination entropy, For the The coordination entropy of the dynamic recognition threshold, is the time-varying penalty weight, is the time-varying penalty decay rate, For the Iteration No. Dynamic network model parameter positions and target positions relevance; The strategy of using oysters to kill newborn bitterlings is used to eliminate individuals in the population. The probability of eliminating newborn bitterlings is: in is the elimination probability of new bitter fish, Selection pressure for the population, For the new bitter fish population Iterative fitness, The new bitterfish population includes new bitterfish and the original bitterfish. Iterative fitness, is the diversity penalty weight, For the Iterated population diversity, historical average position of the population; Repeat the above steps until the objective function is optimized The iteration is stopped when the minimum or maximum number of iterations is reached, the optimal network model parameters are output, and the optimized deep learning network is obtained.

2. The method for predicting supercooled water content in clouds based on deep learning and spatiotemporal data driving according to claim 1 is characterized in that: The method for obtaining satellite characteristics comprises: Calibration, projection conversion, and channel screening are performed on the multi-channel radiation data to obtain first radiation data; the multi-channel radiation data has data dimensions including the number of samples, the number of channels, the number of longitudinal grid points, the number of latitudinal grid points, and the time step; Inputting the first radiation data into a first feature extraction model for feature extraction to obtain satellite features; the first feature extraction model includes an input layer, a convolution layer, a pooling layer, an activation layer, a fully connected layer, and an output layer; The input layer is used for receiving the first radiation data; The convolution layer uses a 4D convolutional neural network to perform a convolution operation on the input data to capture the local features of the data in space and time to obtain the first spatiotemporal radiation feature; the convolution kernel of the 4D convolutional neural network is 2×3×3×3 or 2×5×5×3, with a step size of 2×2×2×2; The pooling layer downsamples the first spatiotemporal radiation feature and simplifies the dimension of the first spatiotemporal radiation feature to obtain a second spatiotemporal radiation feature; the pooling layer performs pooling operations in spatial and temporal dimensions; The activation layer uses an activation function to perform a nonlinear transformation on the output of the convolution layer or the pooling layer; the activation layer includes a first activation layer and a second activation layer; the first activation layer receives the output of the convolution layer and outputs it to the pooling layer; the second activation layer receives the output of the pooling layer and outputs it to the fully connected layer; The fully connected layer maps the second spatiotemporal radiation features after convolution, pooling and activation operations to the output space. The fully connected layer includes multiple neurons, performs linear transformation on the second spatiotemporal radiation features through weight matrix and bias term, and performs nonlinear processing through activation function to obtain satellite features; The output layer connects to the fully connected layer to output satellite features.

3. The method for predicting supercooled water content in clouds based on deep learning and spatiotemporal data driving according to claim 1 is characterized in that: The method for obtaining radar characteristics comprises: Inputting the radar data into a second feature extraction model to extract features to obtain radar features; the radar data dimensions include the number of samples, the number of features, the number of longitudinal grid points, the number of latitudinal grid points, and the time step; the second feature extraction model includes an input layer, a convolution layer, a pooling layer, an activation layer, a fully connected layer, and an output layer; The input layer is used to receive radar data; The convolution layer uses a 4D convolutional neural network to perform a convolution operation on the input data to capture the local features of the data in space and time to obtain the first spatiotemporal radar features. The convolution kernel of the 4D convolutional neural network is 1×3×3×3 or 1×5×5×3, with a step size of 2×2×2×2. The pooling layer downsamples the first spatiotemporal radar feature and simplifies the dimension of the first spatiotemporal radar feature to obtain a second spatiotemporal radar feature; the pooling layer performs pooling operations in spatial and temporal dimensions; The activation layer uses an activation function to perform a nonlinear transformation on the output of the convolution layer or the pooling layer; the activation layer includes a first activation layer and a second activation layer; the first activation layer receives the output of the convolution layer and outputs it to the pooling layer; the second activation layer receives the output of the pooling layer and outputs it to the fully connected layer; The fully connected layer maps the second spatiotemporal radar features after convolution, pooling, and activation operations to the output space. The fully connected layer includes multiple neurons, performs linear transformation on the second spatiotemporal radar features through weight matrices and bias terms, and performs nonlinear processing through activation functions to obtain satellite features. The output layer is connected to the fully connected layer to output radar features.

4. The method for predicting supercooled water content in clouds based on deep learning and spatiotemporal data driving according to claim 1 is characterized in that: The method for obtaining conventional meteorological characteristics comprises: Inputting conventional meteorological data into a third feature extraction model for feature extraction to obtain conventional meteorological features; the conventional meteorological data includes temperature, air pressure, wind direction, wind speed, precipitation, and relative humidity; the dimensions of the conventional meteorological data include the number of features, the number of stations, and the time step; the third feature extraction model includes an input layer, a convolutional layer, a pooling layer, an activation layer, a fully connected layer, and an output layer; The input layer is used to receive regular meteorological data; The convolution layer uses a 3D convolutional neural network to perform a convolution operation on the input data to capture the local features of the data in space and time to obtain the first spatiotemporal conventional meteorological features; the convolution kernel of the 3D convolutional neural network uses a 2×3×3 or 2×5×3 convolution kernel with a step size of 2×2×2; The pooling layer downsamples the first spatiotemporal conventional meteorological data features and simplifies the dimensions of the first spatiotemporal conventional meteorological features to obtain the second spatiotemporal conventional meteorological features; the pooling layer performs pooling operations in the spatial and temporal dimensions; The activation layer uses an activation function to perform a nonlinear transformation on the output of the convolution layer or the pooling layer; the activation layer includes a first activation layer and a second activation layer; the first activation layer receives the output of the convolution layer and outputs it to the pooling layer; the second activation layer receives the output of the pooling layer and outputs it to the fully connected layer; The fully connected layer maps the second spatiotemporal conventional meteorological features after convolution, pooling and activation operations to the output space. The fully connected layer includes multiple neurons, performs linear transformation on the second spatiotemporal conventional meteorological features through weight matrix and bias term, and obtains conventional meteorological features through nonlinear processing through activation function; The output layer is connected to the fully connected layer to output radar features.

5. The method for predicting supercooled water content in clouds based on deep learning and spatiotemporal data-driven according to claim 1 is characterized by: The satellite characteristics include satellite multi-channel radiation data; the radar characteristics include combined reflectivity and VIL; and the conventional meteorological characteristics include temperature, relative humidity, wind direction, wind speed, and precipitation.

Citation Information

Patent Citations

  • Cloud physical parameter prediction method, system, equipment and medium

    CN119202903A

  • Shield tunneling carbon emission prediction method and system based on LSTM and bitter fish optimization

    CN119783877A

  • Cloud water resource forecasting method based on multi-source data fusion

    CN120046511A