Rainfall prediction method based on multi-source meteorological data fusion
By constructing a precipitation prediction model through the fusion of multi-source meteorological data and long short-term memory network algorithm, and combining real-time verification and dynamic optimization, the problem of single data source and fixed parameters in existing technologies is solved, and high-precision precipitation prediction under complex terrain and extreme weather conditions is achieved.
Patent Information
- Application Number
- CN202511402366.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2026-02-10
AI Technical Summary
Existing precipitation forecasting technologies suffer from limitations of single data sources, inability to capture nonlinear relationships and multi-scale interactions, and lack of multi-source data fusion and adaptive optimization mechanisms, resulting in insufficient forecast accuracy and stability, especially under complex terrain and extreme weather conditions.
By fusing multi-source meteorological data and using a long short-term memory network algorithm to construct a precipitation prediction model, combined with real-time verification and evaluation and dynamic parameter optimization, precipitation prediction and error analysis at multiple temporal and spatial scales are achieved, and model parameters are dynamically adjusted to improve prediction accuracy and stability.
It significantly improves the accuracy and stability of precipitation forecasting, maintains high forecasting performance under complex terrain and extreme weather conditions, and realizes the synergistic advantages of multi-source data and adaptive optimization of the model.
Smart Images

Figure CN121503746A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a precipitation prediction method based on multi-source meteorological data fusion. BACKGROUND
[0002] Traditional precipitation prediction methods mainly rely on numerical weather prediction models and statistical prediction techniques. Numerical weather prediction solves the equations of atmospheric dynamics and thermodynamics, and uses finite difference or finite element methods to numerically simulate atmospheric motion, to predict future weather conditions. Statistical prediction methods are based on historical meteorological observation data, and use statistical methods such as multiple regression analysis and time series analysis to establish an empirical relationship between precipitation and previous meteorological factors. Existing precipitation prediction systems usually use a single data source, mainly relying on ground meteorological station observations or satellite remote sensing data, and use linear regression models or simple artificial neural networks to estimate precipitation. In terms of data processing, traditional methods use simple quality control processes to remove or replace abnormal data, and lack complex data fusion techniques. The verification of the prediction results mainly relies on basic statistical indicators such as correlation coefficient and root mean square error, and once the prediction model is established, the parameters are relatively fixed, and there is a lack of dynamic adjustment mechanism.
[0003] However, the existing precipitation prediction technology has significant limitations and shortcomings. First, the limitations of a single data source result in incomplete prediction information, which cannot fully reflect the complex atmospheric processes, especially in complex terrain areas and extreme weather conditions, the prediction accuracy is significantly reduced. Second, traditional linear statistical models are difficult to capture the non-linear relationships and multi-scale interactions in the formation of precipitation, and cannot effectively handle the chaotic characteristics and complex physical mechanisms of the atmospheric system. Third, the existing methods lack effective multi-source data fusion techniques, and different sources of meteorological data are often used independently, without fully utilizing the synergistic advantages of various observation data. Fourth, the parameters of traditional prediction models are fixed and cannot be adaptively adjusted and continuously optimized according to the prediction results, resulting in a decline in model performance over time. Finally, the existing error evaluation and model optimization mechanisms are relatively simple, lacking in-depth error cause analysis and targeted improvement strategies, making it difficult to continuously improve prediction accuracy and dynamically optimize system performance. SUMMARY
[0004] The present application provides a precipitation prediction method based on multi-source meteorological data fusion, which is used to effectively integrate multi-source heterogeneous meteorological data, deeply mine non-linear features, and adaptively optimize prediction models, to significantly improve the accuracy and stability of precipitation prediction.
[0005] The application provides a precipitation prediction method based on multi-source meteorological data fusion, which comprises the following steps: collecting and preprocessing multi-source meteorological data through a meteorological observation network to obtain a standardized meteorological data set; extracting and analyzing precipitation-related characteristic parameters according to the standardized meteorological data set to obtain a precipitation prediction factor matrix; constructing and training a model for the precipitation prediction factor matrix by using a long short-term memory network algorithm to obtain a precipitation prediction model; predicting and calculating multi-time and space scale precipitation based on the precipitation prediction model to obtain a gridded precipitation prediction field; performing error analysis on the gridded precipitation prediction field through a real-time verification and evaluation system to obtain a prediction accuracy evaluation result; and dynamically adjusting and optimizing model parameters according to the prediction accuracy evaluation result to obtain an optimized precipitation prediction system.
[0006] In the technical scheme provided by the application, the beneficial effects are BRIEF DESCRIPTION OF DRAWINGS
[0007] In order to more clearly illustrate the technical scheme of the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0008] Figure 1 An embodiment of the precipitation prediction method based on multi-source meteorological data fusion in the application is shown in the figure. DETAILED DESCRIPTION
[0009] The application provides a precipitation prediction method based on multi-source meteorological data fusion. The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the term "includes" or "has" and any variation thereof is intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0010] For the sake of understanding, the specific process of the embodiments of the application will be described below. Please refer to Figure 1 The embodiment of the precipitation prediction method based on multi-source meteorological data fusion in the application comprises the following steps:
[0011] In step S101, multi-source meteorological data are collected and preprocessed through a meteorological observation network to obtain a standardized meteorological data set.
[0012] In step S102, precipitation-related characteristic parameters are extracted and analyzed according to the standardized meteorological data set to obtain a precipitation prediction factor matrix.
[0013] In step S103, a long short-term memory network algorithm is used to construct and train a model for the precipitation prediction factor matrix to obtain a precipitation prediction model.
[0014] In step S104, multi-temporal and multi-spatial scale precipitation is predicted and calculated based on the precipitation prediction model to obtain a gridded precipitation prediction field.
[0015] In step S105, an error analysis is performed on the gridded precipitation prediction field through a real-time verification and evaluation system to obtain a prediction accuracy evaluation result.
[0016] In step S106, model parameters are dynamically adjusted and optimized according to the prediction accuracy evaluation result to obtain an optimized precipitation prediction system.
[0017] It can be understood that the execution subject of the present application can be a precipitation prediction system based on multi-source meteorological data fusion.
[0018] Specifically, the meteorological observation network synchronously collects parameters such as temperature, humidity, air pressure, wind speed and wind direction through ground automatic weather stations, simultaneously obtains vertical profile data of atmospheric temperature, humidity, wind direction and wind speed at different height layers through upper air sounding equipment, and provides spatial data such as cloud pixel values and water vapor content distribution through a satellite remote sensing system. The quality control module performs outlier detection on the original data, marks and removes an observation value when it exceeds the normal range, and uses a time series interpolation method to complete missing data. Temporal and spatial registration processing unifies data from different sources to the same time interval and spatial grid, and coordinate system conversion ensures that all data use a unified geographic coordinate reference. Standardization processing uses a Z-score normalization method to convert each meteorological element into a standardized value with a mean of 0 and a standard deviation of 1, eliminating the dimensional differences between different physical quantities.
[0019] The standardized meteorological dataset input feature extraction module needs to obtain the temperature decrement rate of different height layers when calculating the atmospheric stability index. The atmospheric stratification stability is judged by the ratio of the vertical temperature gradient to the dry adiabatic decrement rate. The convective available potential energy calculation involves the temperature change of the air mass during the ascending process. The energy parameter is obtained by integrating the difference between the air mass temperature and the ambient temperature. The wind shear strength is obtained by calculating the difference between the wind speed vectors of adjacent height layers. The water vapor flux divergence needs to be calculated by vector operation combined with the wind field and humidity field. The correlation analysis algorithm calculates the Pearson correlation coefficient of each feature parameter and the historical precipitation observation data, and selects the key factors with an absolute correlation coefficient greater than a threshold value. The principal component analysis reduces the data dimension by eigenvalue decomposition, retains the principal components with a cumulative variance contribution rate reaching a set threshold, and extracts the periodic change mode and trend characteristics of each factor through time series analysis, to finally form a precipitation prediction factor matrix.
[0020] The precipitation prediction factor matrix is divided into a training set and a validation set according to the time sequence, and the training set accounts for 80% of the total data for model parameter learning. The LSTM network architecture design includes an input layer receiving a multi-dimensional feature vector, multiple LSTM hidden layers processing time series information through a gating mechanism, and an output layer generating precipitation prediction values. During the forward propagation process, the input data is subjected to nonlinear transformation through the forget gate, input gate, and output gate, and the cell state transmits long-term information between time steps. The backpropagation algorithm updates the network weight parameters through gradient descent, and the loss function uses mean square error to measure the difference between the predicted value and the true value. The attention mechanism module assigns weight coefficients to different features, highlighting meteorological factors that contribute more to precipitation prediction, and normalizes the attention weights through a softmax function. Cross-validation uses time series segmentation methods to avoid data leakage, and hyperparameter tuning determines the optimal learning rate, batch size, and other parameter configurations through grid search.
[0021] Real-time meteorological data is input into the trained precipitation prediction model, which outputs precipitation prediction values for future time steps. The rolling prediction strategy updates the input data every hour to generate new prediction results. Numerical weather prediction products provide medium-term meteorological element prediction fields, which are time-scaled with short-term prediction results to extend the prediction time. The study area is divided into regular grids according to latitude and longitude intervals, and each grid point corresponds to a geographic coordinate position. Spatial interpolation uses the inverse distance weighting method to calculate the prediction value of a grid point based on the distance and numerical value of surrounding observation points. The closer the observation points, the greater the weight. In the interpolation formula, the weight coefficient is inversely proportional to the square of the distance, ensuring spatial continuity. Visualization rendering converts grid point values into contour maps and color-filled maps, with different precipitation intensities corresponding to different color levels, forming an intuitive grid-based precipitation prediction field.
[0022] A gridded precipitation forecast field is compared point-by-point with concurrent observed precipitation data, and spatiotemporal matching ensures that predicted and observed values correspond to the same time and location. Prediction error calculation involves the difference between predicted and observed values at each grid point, forming a spatially distributed error matrix. The root mean square error (RMSE) quantifies prediction accuracy using the arithmetic mean of the sum of squared errors, while the mean absolute error (MAE) assesses systematic bias using the arithmetic mean of absolute error values. Correlation coefficient analysis calculates the linear correlation between predicted and observed value sequences, reflecting the consistency between predicted trends and actual conditions. A comprehensive evaluation function weights and combines multiple error indicators, with weighting coefficients determined based on indicator importance. A skill score reflects the degree of improvement in prediction skills by comparing with a reference prediction method. Spatiotemporal distribution statistical analysis analyzes the variation patterns of errors in different regions and time periods, identifying spatiotemporal ranges with better and worse prediction performance.
[0023] The prediction accuracy evaluation results are input into the error diagnosis module, which identifies the main sources and distribution characteristics of errors through statistical analysis. The classification results include systematic bias, random error, and prediction bias due to extreme events. The parameter adjustment strategy determines the corresponding optimization direction based on the error type; systematic bias requires adjusting the model output bias parameters, while excessive random error necessitates enhancing model stability. The adaptive learning algorithm employs an online gradient descent method, updating the model weight parameters based on new prediction error information, with the learning rate dynamically adjusted according to the error change trend. A quality control algorithm monitors the stability of the parameter update process, triggering a rollback mechanism to restore the system to a stable state when parameter changes exceed a threshold. The system reconstruction process integrates the optimized model parameters into the prediction process, verifying the optimization effect through independent test data to ensure the effectiveness and stability of the improved prediction accuracy.
[0024] In one specific embodiment, the process of performing step S101 may specifically include the following steps:
[0025] Temperature, humidity, air pressure, and wind speed and direction data are collected synchronously from ground meteorological stations to obtain a set of ground meteorological observation data.
[0026] The atmospheric stratification parameter matrix is obtained by scanning the vertical profile data of the atmosphere layer by layer using high-altitude sounding equipment.
[0027] The ground meteorological observation data set and atmospheric stratification parameter matrix are input into the quality control module for outlier detection and removal to obtain a purified meteorological dataset.
[0028] The purified meteorological dataset is subjected to spatiotemporal registration and coordinate system unification processing to obtain a synchronized meteorological data sequence;
[0029] The synchronous meteorological data sequence is numerically normalized using a standardization processing algorithm to obtain the standardized meteorological dataset.
[0030] Specifically, ground-based meteorological stations synchronously acquire real-time observation data using equipment such as temperature sensors, humidity sensors, air pressure sensors, and wind direction and speed meters. All sensors use a unified timestamp to ensure the synchronization of data acquisition, forming a ground-based meteorological observation data set containing multiple meteorological elements. Upper-air sounding equipment, such as radiosondes, is released from the ground and rises layer by layer, measuring parameters such as temperature, humidity, air pressure, and wind speed at regular altitude intervals. This acquires vertical distribution data from the ground to different altitude layers, arranging them according to altitude to form an atmospheric stratification parameter matrix.
[0031] The quality control module sets the normal value range for each meteorological element. Observations exceeding the physically reasonable range or differing significantly from data from neighboring observation points are marked as outliers and removed. Missing data is filled using linear interpolation of nearby time points. After outlier detection and data completion, a purified meteorological dataset is obtained. Spatiotemporal registration unifies data from different sources and time intervals onto the same time grid. The coordinate system is unified by converting geographic coordinates to a uniform projected coordinate system, ensuring data consistency in both time and space, forming a synchronized meteorological data sequence.
[0032] The standardization process employs zero-mean normalization to calculate the mean and standard deviation of each meteorological element. The original values are subtracted from the mean and then divided by the standard deviation to convert meteorological elements of different dimensions into dimensionless standardized values. This eliminates the impact of numerical differences on subsequent analysis and ultimately forms a standardized meteorological dataset.
[0033] In one specific embodiment, the process of performing step S102 may specifically include the following steps:
[0034] Based on the standardized meteorological dataset, the atmospheric stability index and convective available potential energy are calculated and extracted to obtain a set of atmospheric stability characteristic parameters.
[0035] Based on the vertical wind field data, the wind shear intensity and water vapor flux divergence are quantitatively analyzed to obtain a set of dynamic characteristic parameters.
[0036] The atmospheric stability characteristic parameter set and the dynamic characteristic parameter set are input into a correlation analysis algorithm to calculate the correlation degree, and the characteristic importance weight vector is obtained.
[0037] Principal component analysis is performed on the feature importance weight vector to reduce its dimensionality, resulting in a set of key predictive factors.
[0038] The precipitation prediction factor matrix is obtained by extracting periodic features from the set of key prediction factors based on time series analysis.
[0039] Specifically, the atmospheric stability index calculation requires temperature data from different altitude layers. The stability of the atmospheric stratification is determined by calculating the ratio of the actual temperature lapse rate to the dry adiabatic lapse rate. A ratio less than 1 indicates an unstable stratification conducive to convection. The calculation of convective available potential energy involves the adiabatic ascent of an air parcel. The air parcel rises along the dry adiabatic line to the condensation height and then continues along the wet adiabatic line. The integral of the temperature difference between the air parcel and the ambient temperature yields the energy value usable for convection development. A larger value indicates a stronger convection potential. These two parameters together constitute the characteristic parameter set of atmospheric stability.
[0040] Wind shear intensity is obtained by calculating the difference in wind speed vectors between adjacent altitude layers. Specifically, it calculates the rate of change of wind speed in the horizontal and vertical directions at each altitude layer, reflecting the degree of non-uniformity of atmospheric motion. Water vapor flux divergence needs to be calculated in conjunction with the wind field and specific humidity field. First, the water vapor flux in each direction is calculated, and then the divergence of the water vapor flux is calculated. Positive values indicate that water vapor divergence is unfavorable for precipitation, while negative values indicate that water vapor convergence is favorable for precipitation formation. These parameters constitute a set of dynamic characteristic parameters.
[0041] Correlation analysis algorithms calculate the Pearson correlation coefficients between each feature parameter and historical precipitation observation data. The correlation coefficients, ranging from -1 to 1, are obtained by dividing the covariance by the product of the standard deviations of the two variables; a larger absolute value indicates a stronger correlation. The correlation coefficients of all feature parameters are then used to construct a feature importance weight vector. Principal component analysis projects the original multidimensional feature space to a lower-dimensional space through eigenvalue decomposition, retaining principal components with larger eigenvalues. These principal components contain the main information from the original data and are orthogonal to each other. Principal components with a cumulative variance contribution rate of over 85% are selected as the set of key predictive factors.
[0042] Time series analysis uses Fourier transform to identify the periodic variation patterns of each predictor, extracts variation patterns at different time scales such as daily and seasonal cycles, and calculates trend terms to reflect long-term variation characteristics. These time features are combined with the original predictor factors to form a precipitation predictor matrix containing time dimension information. The rows of the matrix represent time series, and the columns represent different predictor factors and their time features.
[0043] In one specific embodiment, the process of executing step S103 may specifically include the following steps:
[0044] The precipitation prediction factor matrix is divided into a training sample set and a validation sample set to construct a dataset, thus obtaining model training data pairs.
[0045] Based on the Long Short-Term Memory network architecture, the network structure of the input layer, hidden layer and output layer is designed to obtain the LSTM network topology.
[0046] Based on the model training data, the LSTM network topology is trained by forward and backward propagation to obtain the network weight parameter matrix.
[0047] The attention mechanism module is embedded into the network weight parameter matrix to adaptively adjust the feature importance, thereby obtaining an optimized network parameter set;
[0048] The optimized network parameter set is subjected to cross-validation and hyperparameter tuning to obtain the precipitation prediction model.
[0049] Specifically, the precipitation predictor matrix is divided in chronological order. The first 80% of the data is used as a training sample set for model parameter learning, and the last 20% is used as a validation sample set for model performance evaluation. Each sample contains multiple predictors within a continuous time window as input features, and the corresponding precipitation observations are used as target outputs, forming input-output paired model training data pairs.
[0050] In the LSTM network topology design, the input layer receives a multi-dimensional predictor vector, where the dimension equals the number of predictors. The hidden layer adopts a multi-layer LSTM unit structure, with each LSTM unit containing three gate structures: a forget gate, an input gate, and an output gate. The forget gate determines which historical information to discard, the input gate controls the storage of new information, and the output gate determines the output content at the current moment. Cell states transmit long-term memory information between time steps. The output layer uses a fully connected layer to map the LSTM output to the precipitation prediction value.
[0051] During forward propagation, the input data passes through each layer of the network structure sequentially. The LSTM unit performs a nonlinear transformation on the input sequence through a gating mechanism. The cell state update formula combines the current input and the state at the previous time step to calculate the new state value. Backpropagation training uses the temporal backpropagation algorithm, calculating the gradient of the loss function with respect to the parameters of each layer starting from the output layer. The gradient propagates layer by layer to the input layer using the chain rule. The gradient descent algorithm is used to update the weight matrix and bias vector in the network. After multiple rounds of iterative training, a converged network weight parameter matrix is obtained.
[0052] The attention mechanism module calculates the importance of each predictor to the current prediction task. It maps the predictors to query vectors, key vectors, and value vectors through a learnable weight matrix. The inner product of the query vector and the key vector is calculated to obtain the attention score. After softmax normalization, the score is used as the weight coefficient of each feature. The weight coefficient is multiplied by the value vector to obtain the weighted feature representation. This mechanism enables the model to automatically focus on meteorological factors that contribute significantly to precipitation prediction. The attention weights are combined with the original network parameters to form an optimized network parameter set.
[0053] Cross-validation employs a time-series partitioning method, dividing the training data into multiple folds in chronological order. Each time, the model is trained using earlier folds, and its performance is validated using later folds, preventing future information from being leaked into historical predictions. Hyperparameter tuning systematically tests different combinations of parameters such as learning rate, batch size, number of hidden layer neurons, and number of LSTM layers using a grid search method. The optimal parameter configuration is selected based on the performance on the validation set, ultimately resulting in a fully trained and optimized precipitation prediction model.
[0054] In one specific embodiment, the process of executing step S104 may specifically include the following steps:
[0055] Real-time meteorological data is input into the precipitation prediction model for rolling prediction calculation to obtain hourly precipitation prediction sequences.
[0056] The hourly precipitation prediction sequence is extended on a time scale based on numerical weather prediction products to obtain multi-time-dependent precipitation prediction data.
[0057] The predicted area is divided into regular grids based on the geographical coordinates of the study area to obtain a spatial grid node matrix.
[0058] The multi-time-dependent precipitation forecast data are allocated to the spatial grid node matrix for spatial interpolation calculation to obtain the precipitation forecast value of the grid point.
[0059] The predicted precipitation values at the grid points are spatially continuous and visualized to obtain the gridded precipitation prediction field.
[0060] Specifically, real-time meteorological data, including the latest observations of temperature, humidity, and wind speed, is preprocessed according to the data format used during training and then input into the precipitation prediction model. The model calculates the probability and intensity of precipitation at each future moment based on current meteorological conditions and historical evolution patterns. The rolling forecast employs a time window sliding strategy, using a fixed-length historical data window for each forecast. As new observation data arrives, the window slides forward one time step, generating new forecast results and forming a continuous hourly precipitation forecast sequence.
[0061] Numerical weather forecast products provide forecast fields of meteorological elements for the next 72 hours or longer, including three-dimensional grid data such as temperature, humidity, wind field, and air pressure. Short-term hourly forecast sequences are stitched together with medium- and long-term numerical weather forecast products along the time dimension. Short-term forecasts cover the next 24 hours, while medium-term forecasts are extended to the next 72 hours using numerical weather forecast products. Time interpolation methods are used to smoothly connect forecasts with different timeframes, forming multi-timeframe precipitation forecast data from short-term to medium-term.
[0062] The study area is divided into regular grids according to latitude and longitude intervals. The grid spacing is determined based on the required prediction accuracy, typically using 0.1 degrees or 0.05 degrees. Each grid point corresponds to a unique latitude and longitude coordinate. Grid division begins at the southwest corner of the area and is numbered sequentially by row and column, forming a two-dimensional spatial grid node matrix. The number of rows in the matrix corresponds to the number of grids in the latitudinal direction, and the number of columns corresponds to the number of grids in the longitudinal direction.
[0063] Spatial interpolation calculations employ an inverse distance weighting method. For any grid point, observation stations or prediction points within a certain surrounding range are selected, and weight coefficients are assigned based on their distance; the closer the distance, the greater the weight, and the weight is inversely proportional to the square of the distance. The precipitation prediction value for each observation point is multiplied by its corresponding weight and then summed to obtain the precipitation prediction value for that grid point. After calculations are performed for all grid points, a grid point precipitation prediction value covering the entire area is formed.
[0064] Spatial continuity processing fills the gaps between grid points using bilinear interpolation, ensuring a smooth spatial transition of the precipitation field. Visualization rendering converts grid point values into contour maps and color-filled maps, assigning different color levels based on precipitation intensity: no precipitation areas are displayed as white or light blue, light rain as green, moderate rain as yellow, heavy rain as orange, and torrential rain as red, forming an intuitive gridded precipitation forecast field for users to view and analyze.
[0065] In one specific embodiment, the process of executing step S105 may specifically include the following steps:
[0066] The prediction error distribution matrix is obtained by comparing the gridded precipitation prediction field with the actual precipitation observation data in a spatiotemporal manner.
[0067] Based on statistical methods, the root mean square error and mean absolute error of the prediction error distribution matrix are calculated to obtain a set of quantitative error indicators.
[0068] The linear correlation between the predicted and observed values is tested using the correlation coefficient analysis algorithm to obtain the prediction correlation evaluation parameters.
[0069] The quantitative error index set and the prediction correlation evaluation parameters are input into the comprehensive evaluation function for weighted calculation to obtain the comprehensive prediction skill score.
[0070] The spatiotemporal distribution statistics and trend analysis of the comprehensive prediction skill score are processed to obtain the prediction accuracy evaluation result.
[0071] Specifically, each grid point in the gridded precipitation prediction field has corresponding geographic coordinates and a prediction time. The actual precipitation observation data comes from the real-time observation records of the rain gauge network. The spatiotemporal matching process first performs spatial matching, projecting the rain gauge coordinates onto the prediction grid and finding the nearest grid point as the correspondence. Then, it performs temporal matching to ensure that the predicted values and observed values correspond to the same time period. By comparing each point, the difference between the predicted and observed values is calculated, and the error values of all grid points are arranged according to their spatial location to form a prediction error distribution matrix.
[0072] The root mean square error (RMSE) is calculated by squaring the prediction error at each grid point, taking the arithmetic mean of all squared errors, and then taking the square root. This RMSE reflects the overall degree of deviation between the predicted and observed values. The mean absolute error (MAO) is calculated by averaging the absolute values of the prediction errors at each grid point. This MAO measures the average deviation between the predicted and observed values. These two indicators constitute a quantitative error index set; the smaller the value, the higher the prediction accuracy.
[0073] Correlation analysis tests the linear correlation between the predicted and observed value sequences of all grid points, and calculates the Pearson correlation coefficient, which is the product of the covariance of the two sequences and their respective standard deviations. The correlation coefficient ranges from -1 to 1. A value close to 1 indicates a strong positive correlation, a value close to -1 indicates a strong negative correlation, and a value close to 0 indicates a weak linear correlation. The absolute value of the correlation coefficient reflects the degree of consistency between the trend of the predicted value and the trend of the actual observed value, forming a parameter for evaluating the predictive correlation.
[0074] The comprehensive evaluation function employs a weighted average method, linearly combining the root mean square error (RMSE), mean absolute error (MAE), and correlation coefficient with different weights. Error-related indicators have negative weights, while the correlation coefficient has a positive weight, ensuring that the skill score increases with prediction accuracy. The weighting coefficients are determined based on the importance and sensitivity of each indicator; typically, the RMSE has a weight of 0.4, the MAE has a weight of 0.3, and the correlation coefficient has a weight of 0.3. The weighted calculation yields the comprehensive prediction skill score; a higher score indicates better prediction skill.
[0075] The spatiotemporal distribution statistical analysis examines the variation patterns of skill scores across different geographical regions and time periods, calculating the average skill score and standard deviation for each region to identify spatial locations with better and worse predictive performance. Trend analysis employs linear regression to examine the changing trends of skill scores over time; a positive trend indicates gradually improving predictive ability, while a negative trend indicates a decline. Combined with seasonality analysis, the differences in predictive performance across different seasons are identified, ultimately forming a predictive accuracy evaluation result that incorporates spatial distribution characteristics, temporal trends, and seasonal differences.
[0076] In one specific embodiment, the process of performing step S106 may specifically include the following steps:
[0077] Based on the prediction accuracy evaluation results, the prediction error is analyzed for causes and pattern recognition is performed to obtain the error source classification results.
[0078] Based on the error source classification results, the model weight parameters and network structure are adjusted and calculated in a targeted manner to obtain a parameter optimization scheme;
[0079] The parameter optimization scheme is input into the adaptive learning algorithm for online parameter updating, resulting in a dynamic optimization model parameter set;
[0080] Based on the quality control algorithm, the stability test and anomaly handling of the dynamic optimization model parameter set are performed to obtain a stable model parameter configuration.
[0081] The stable model parameters are integrated into the prediction process for system reconstruction and performance testing, resulting in the optimized precipitation prediction system.
[0082] Specifically, the prediction accuracy evaluation results include the spatiotemporal distribution characteristics and statistical properties of the errors. Causal analysis uses clustering algorithms to group similar error patterns into one category, identifying different types such as systematic bias, random errors, and missed reporting of extreme events. Pattern recognition employs decision tree methods to analyze the correlation between errors and meteorological conditions. When the temperature gradient is large, the model tends to underestimate precipitation; when humidity changes drastically, time-shift errors are easily generated; and when the terrain is complex, spatial interpolation accuracy decreases. Based on these patterns, the errors are categorized into source classifications such as terrain influence, thermal processes, and dynamic processes.
[0083] The parameter optimization scheme formulates corresponding strategies based on the type of error source. For systematic biases, the bias parameters of the output layer are adjusted, and fixed values are added or reduced to correct the prediction results. For random errors, the number of neurons in the LSTM hidden layer is increased to enhance model complexity. For errors under specific meteorological conditions, the weight coefficients of the corresponding features in the attention mechanism are adjusted. Network structure adjustments include adding or removing the number of hidden layers, modifying the activation function type, and adjusting regularization parameters. Each adjustment scheme has a clear target error type and expected improvement effect.
[0084] The adaptive learning algorithm employs an online gradient descent method, calculating the gradient direction based on the latest prediction error information. The learning rate uses an adaptive adjustment strategy: increasing the learning rate to accelerate convergence as the error continues to decrease, and decreasing it to ensure stability when the error oscillates. Parameter updates utilize a momentum mechanism, combining historical gradient information and the current gradient to calculate parameter adjustments, avoiding getting trapped in local optima. After each update, the new parameter configuration is saved, forming a dynamically optimized model parameter set that evolves over time.
[0085] The quality control algorithm monitors the numerical stability during parameter updates. When the parameter change exceeds a preset threshold, an anomaly handling mechanism is triggered to check for gradient explosion or vanishing phenomena. Stability testing includes multiple aspects such as parameter range checks, gradient norm checks, and loss function convergence checks. Anomaly handling employs a parameter rollback strategy, restoring unstable parameters to their previous stable state. Through multiple iterations, a model parameter configuration that improves prediction accuracy while ensuring numerical stability is obtained.
[0086] The system refactoring involves replacing the original parameters with optimized ones, updating model components in the prediction process, and keeping the data preprocessing and post-processing modules unchanged. Performance testing uses an independent test dataset to verify the optimization effect, comparing prediction accuracy metrics before and after optimization to ensure the effectiveness and generalization ability of the improvement. Testing includes prediction accuracy under different weather conditions, prediction skill scores at different lead times, and prediction error distribution in different regions. After a comprehensive evaluation of system performance, it is deployed as a formal optimized precipitation prediction system.
[0087] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A precipitation prediction method based on multi-source meteorological data fusion, characterized in that, The method includes: A standardized meteorological dataset is obtained by collecting and preprocessing multi-source meteorological data through a meteorological observation network. Based on the standardized meteorological dataset, precipitation-related characteristic parameters are extracted and analyzed to obtain a precipitation prediction factor matrix; The precipitation prediction factor matrix is trained using a long short-term memory network algorithm to obtain a precipitation prediction model. Based on the precipitation prediction model, precipitation at multiple temporal and spatial scales is predicted and calculated to obtain a gridded precipitation prediction field. Error analysis of the gridded precipitation prediction field is performed using a real-time verification and evaluation system to obtain prediction accuracy evaluation results; Based on the prediction accuracy evaluation results, the model parameters are dynamically adjusted and optimized to obtain an optimized precipitation prediction system.
2. The precipitation prediction method based on multi-source meteorological data fusion according to claim 1, characterized in that, The process of collecting and preprocessing multi-source meteorological data through a meteorological observation network to obtain a standardized meteorological dataset includes: Temperature, humidity, air pressure, and wind speed and direction data are collected synchronously from ground meteorological stations to obtain a set of ground meteorological observation data. The atmospheric stratification parameter matrix is obtained by scanning the vertical profile data of the atmosphere layer by layer using high-altitude sounding equipment. The ground meteorological observation data set and atmospheric stratification parameter matrix are input into the quality control module for outlier detection and removal to obtain a purified meteorological dataset. The purified meteorological dataset is subjected to spatiotemporal registration and coordinate system unification processing to obtain a synchronized meteorological data sequence; The synchronous meteorological data sequence is numerically normalized using a standardization processing algorithm to obtain the standardized meteorological dataset.
3. The precipitation prediction method based on multi-source meteorological data fusion according to claim 1, characterized in that, The step involves extracting and analyzing precipitation-related characteristic parameters based on the standardized meteorological dataset to obtain a precipitation prediction factor matrix, including: Based on the standardized meteorological dataset, the atmospheric stability index and convective available potential energy are calculated and extracted to obtain a set of atmospheric stability characteristic parameters. Based on the vertical wind field data, the wind shear intensity and water vapor flux divergence are quantitatively analyzed to obtain a set of dynamic characteristic parameters. The atmospheric stability characteristic parameter set and the dynamic characteristic parameter set are input into a correlation analysis algorithm to calculate the correlation degree, and the characteristic importance weight vector is obtained. Principal component analysis is performed on the feature importance weight vector to reduce its dimensionality, resulting in a set of key predictive factors. The precipitation prediction factor matrix is obtained by extracting periodic features from the set of key prediction factors based on time series analysis.
4. The precipitation prediction method based on multi-source meteorological data fusion according to claim 1, characterized in that, The precipitation prediction model is obtained by training the precipitation prediction factor matrix using a long short-term memory network algorithm, including: The precipitation prediction factor matrix is divided into a training sample set and a validation sample set to construct a dataset, thus obtaining model training data pairs. Based on the Long Short-Term Memory network architecture, the network structure of the input layer, hidden layer and output layer is designed to obtain the LSTM network topology. Based on the model training data, the LSTM network topology is trained by forward and backward propagation to obtain the network weight parameter matrix. The attention mechanism module is embedded into the network weight parameter matrix to adaptively adjust the feature importance, thereby obtaining an optimized network parameter set; The optimized network parameter set is subjected to cross-validation and hyperparameter tuning to obtain the precipitation prediction model.
5. The precipitation prediction method based on multi-source meteorological data fusion according to claim 1, characterized in that, The prediction calculation based on the precipitation prediction model at multiple spatiotemporal scales yields a gridded precipitation prediction field, including: Real-time meteorological data is input into the precipitation prediction model for rolling prediction calculation to obtain hourly precipitation prediction sequences. The hourly precipitation prediction sequence is extended on a time scale based on numerical weather prediction products to obtain multi-time-dependent precipitation prediction data. The predicted area is divided into regular grids based on the geographical coordinates of the study area to obtain a spatial grid node matrix. The multi-time-dependent precipitation forecast data are allocated to the spatial grid node matrix for spatial interpolation calculation to obtain the precipitation forecast value of the grid point. The predicted precipitation values at the grid points are spatially continuous and visualized to obtain the gridded precipitation prediction field.
6. The precipitation prediction method based on multi-source meteorological data fusion according to claim 1, characterized in that, The error analysis of the gridded precipitation prediction field through a real-time verification and evaluation system to obtain prediction accuracy evaluation results includes: The prediction error distribution matrix is obtained by comparing the gridded precipitation prediction field with the actual precipitation observation data in a spatiotemporal manner. Based on statistical methods, the root mean square error and mean absolute error of the prediction error distribution matrix are calculated to obtain a set of quantitative error indicators. The linear correlation between the predicted and observed values is tested using the correlation coefficient analysis algorithm to obtain the prediction correlation evaluation parameters. The quantitative error index set and the prediction correlation evaluation parameters are input into the comprehensive evaluation function for weighted calculation to obtain the comprehensive prediction skill score. The spatiotemporal distribution statistics and trend analysis of the comprehensive prediction skill score are processed to obtain the prediction accuracy evaluation result.
7. The precipitation prediction method based on multi-source meteorological data fusion according to claim 1, characterized in that, The step of dynamically adjusting and optimizing the model parameters based on the prediction accuracy evaluation results to obtain an optimized precipitation prediction system includes: Based on the prediction accuracy evaluation results, the prediction error is analyzed for causes and pattern recognition is performed to obtain the error source classification results. Based on the error source classification results, the model weight parameters and network structure are adjusted and calculated in a targeted manner to obtain a parameter optimization scheme; The parameter optimization scheme is input into the adaptive learning algorithm for online parameter updating, resulting in a dynamic optimization model parameter set; Based on the quality control algorithm, the stability of the dynamic optimization model parameter set is checked and anomalies are handled to obtain a stable model parameter configuration. The stable model parameters are integrated into the prediction process for system reconstruction and performance testing, resulting in the optimized precipitation prediction system.