Method for simulating extreme rainfall meteorological influence in Asian monsoon area based on artificial intelligence
By preprocessing meteorological data and establishing mathematical models, the most influential variables are selected and the simulation results are explained using interpretability methods, the problem of lack of theoretical basis for variable selection and insufficient model interpretability in the existing technology is solved, and efficient and accurate simulation of extreme precipitation in the East Asian monsoon region is achieved.
Patent Information
- Application Number
- CN202411985867.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-23
Smart Images

Figure CN120030880A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of meteorological simulation, and in particular to a method for simulating the meteorological impact of extreme precipitation in the Asian monsoon region based on artificial intelligence. Background Art
[0002] Artificial intelligence technology has shown great potential and broad application prospects in all aspects, but there are relatively few studies on the application of artificial intelligence to simulate extreme precipitation. This means that in the study of extreme precipitation, an important meteorological phenomenon, the role of artificial intelligence has not been fully explored and exerted, and there is still a lot of room for exploration waiting for researchers to explore. In existing studies, the selection of simulation variables often has a certain degree of randomness, but this random selectivity lacks corresponding theoretical basis as support. This makes the variable selection process appear to be more arbitrary, which may affect the accuracy and reliability of the simulation results. The variable selection method without theoretical basis is also difficult to promote and apply in different research scenarios, which limits the further development of related research. Deep learning models have achieved remarkable results in many fields, but in the process of simulating extreme precipitation, the interpretability of deep learning models is seriously insufficient. This means that although the model can give simulation results, it is difficult for researchers to understand how the model obtains these results and cannot deeply analyze the internal working mechanism of the model. The lack of interpretability not only brings difficulties to the optimization and improvement of the model, but also affects people's trust in the model to a certain extent. Summary of the invention
[0003] In order to solve the above technical problems, the present invention proposes a method for simulating the meteorological impact of extreme precipitation in the Asian monsoon region based on artificial intelligence, comprising the following steps:
[0004] S1, preprocessing meteorological data to obtain a variable set;
[0005] S2. Establish a mathematical model to simulate extreme precipitation, input the variable set obtained after preprocessing into the mathematical model in its entirety, compare the simulation performance before and after variable selection, and screen out the most influential variables;
[0006] S3. Use the most influential variables screened out in step S2 to re-simulate extreme precipitation and evaluate the effectiveness of the most influential variables in improving the simulation effect.
[0007] In a preferred embodiment, in step S1, the correlation coefficient between each pair of variables and the variance inflation factor VIF of all input variables are calculated. i , the expression is:
[0008]
[0009]
[0010] Among them, Cov(X,Y) is the covariance of the two variables, Var(X), Var(Y) is the variance of variables X and Y, r is the correlation coefficient, r i is the correlation coefficient between the ith pair of variables, and p is the total number of variable pairs.
[0011] In a preferred embodiment, in step S2, the expression for evaluating the performance of the mathematical model is:
[0012]
[0013] Among them, P is precision, R is recall and F1Score is score, TP stands for true positive, FP stands for false positive, and TN stands for true negative.
[0014] In a preferred embodiment, permutation importance is used as a feature importance calculation method: a feature is selected, all values of the feature in the data set are randomly permuted, and new simulation results are calculated. If the difference between the new and old results is small, it indicates that the importance of the feature is low; if the difference is significant, it indicates that the feature has a great impact on the model.
[0015] In a preferred embodiment, in step S2, Shapley regression values are used to explain the simulation results. The Shapley regression value φ i The expression is:
[0016]
[0017] Where F represents the set of all input meteorological variables, S represents a subset of meteorological variables in F, and f S∪{i}(x S∪{i}) represents the SHAP value of i contained in the set S, f s (x s ) represents the SHAP value after i is excluded.
[0018] In a preferred embodiment, in step S3, the six factors that have a significant impact on precipitation meteorology are arranged in descending order: total water vapor TMQ, precipitation on the previous day pr_b, total water vapor transport IVT, sea level pressure MSLP, 300hPa vertical velocity ω300 and 850hPa meridional wind V850.
[0019] In a preferred embodiment, in step S1, the preprocessing process includes: cleaning, resampling and normalizing all potential meteorological variables, while removing highly correlated input variables for dimensionality reduction to obtain a variable set.
[0020] Compared with the prior art, the present invention has the following beneficial technical effects:
[0021] The present invention can effectively simulate extreme precipitation in the East Asian monsoon region. Through an interpretable deep learning model, six meteorological variables that have a significant impact on extreme precipitation are determined, and extreme precipitation in the East Asian region can be automatically identified and simulated. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 The present invention is a flow chart of a method for simulating the meteorological impact of extreme precipitation in the Asian monsoon region based on artificial intelligence. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0024] In the drawings of the specific embodiments of the present invention, in order to better and more clearly describe the working principles of the various components in the system, the connection relationship of the various parts in the device is shown, which only clearly distinguishes the relative position relationship between the various components, and cannot constitute a limitation on the signal transmission direction, connection sequence and size, dimensions and shape of the components or structures.
[0025] like Figure 1 As shown, it is a flow chart of the method for simulating the meteorological impact of extreme precipitation in the Asian monsoon region based on artificial intelligence of the present invention, which includes the following steps:
[0026] S1. Preprocess the meteorological data to obtain a variable set.
[0027] The preprocessing process includes cleaning, resampling and normalizing all potential meteorological variables, while removing highly correlated input variables for dimensionality reduction to obtain a variable set.
[0028] The present invention uses the Multi-Source Weighted Ensemble Precipitation Dataset (MSWEP) as observational data with a temporal resolution of 3 hours and a spatial resolution of 0.1°. In order to provide the best quality precipitation estimates for each location, MSWEP integrates data from multiple sources, including satellite remote sensing, radar observations, ground meteorological stations, and numerical weather prediction models. ERA5 reanalysis data are selected as the source of meteorological factors that potentially affect extreme precipitation (EP), which are then used to simulate extreme precipitation. The dataset includes six upper-air variables: temperature (T), geopotential height (Z), specific humidity (q), vertical velocity (ω), and horizontal wind speed (U and V) at 300, 500, and 850 hPa. In addition, three single-layer variables are included: sea level pressure (SLP), convective effective potential energy (CAPE), and precipitation (pr_b) of the previous day, as well as two commonly used variables for diagnosing water vapor: water vapor flux (IVT) and whole-layer water vapor (IWV).
[0029] The present invention focuses on EP in the warm season, but it is crucial to simulate the annual daily precipitation first. This is because there is a difference in the occurrence time of EP in the definition of the observed data and the reference data. If only EP is simulated and other types of precipitation are ignored, even if the input variables are irrelevant to EP or the overall performance of the model is poor, the model may still achieve ideal results on EP due to the simulated precipitation being high enough. However, when simulating annual daily precipitation, EP is unbalanced data relative to total precipitation, which causes problems in regression analysis, so preprocessing is necessary, and preprocessing of the original data is essential. The present invention adopts SMOGN (named after SmoteR and the introduction of Gaussian noise) as a preprocessing method. SMOGN generates synthetic samples by combining these strategies. By introducing Gaussian noise, a more conservative strategy, SMOGN mitigates the potential risks associated with SmoteR while increasing the diversity of synthetic samples.
[0030] The central assumption of variable selection is that a high-quality set of variables should include those variables that are highly correlated with the categories but remain uncorrelated with each other. The present invention calculates the correlation coefficient (r) between each pair of variables and the variance inflation factor (VIF) of all input variables, which is expressed as:
[0031]
[0032] Among them, Cov(X,Y) is the covariance of two variables, Var(X), Var(Y) is the variance of variables X and Y. r is the correlation coefficient, r i is the correlation coefficient between the ith pair of variables, and p is the total number of variable pairs.
[0033] VIF is a key indicator for quantifying the degree of multicollinearity. The higher the VIF value, the more serious the multicollinearity. In statistics and related fields, VIF values greater than 5 or 10 are usually used as screening criteria. In order to more strictly consider the impact of multicollinearity, the present invention sets thresholds of r>0.7 and VIF>5. First, the present invention calculates r between each pair of variables. When r>0.7, the correlation between the two and extreme precipitation events (EP) is evaluated, and variables with weaker correlation are deleted. Subsequently, the VIF of all remaining variables is calculated, and variables with VIF lower than 5 are deleted.
[0034] Specifically, r is used to remove pairwise correlated variables. Multicollinearity can lead to an increase in the variance of the regression coefficient estimator, thereby reducing the stability and reliability of the model. VIF is an effective tool to measure the severity of this multicollinearity. VIF is calculated based on the linear relationship between each independent variable and other independent variables in the regression model. Each independent variable has a corresponding VIF value that is greater than 1. The closer the VIF value is to 1, the less severe the multicollinearity, and vice versa.
[0035] S2. Establish a mathematical model to simulate extreme precipitation, input the variable set obtained after preprocessing into the mathematical model in its entirety, compare the simulation performance before and after variable selection, and screen out the most influential variables.
[0036] In order to demonstrate the interpretability of the selected input variables for EP (extreme precipitation), the present invention is first verified by simulating EP. Therefore, the output of the model simulation is named EP. The random forest algorithm is applied to the training and simulation in the XAI (Explainable Artificial Intelligence) model. As an integrated decision tree algorithm, random forest (RF) improves the accuracy of simulation and prevents overfitting by combining multiple decision trees, and its superiority has been demonstrated in many studies.
[0037] The present invention models the data in a grid-based manner, using the potential influencing variables introduced above as input. For each grid point, the present invention uses variable data at two time points: the previous moment and the current moment. The current grid point and its eight neighboring points (up, down, left, right, upper left, lower left, upper right, and lower right) are selected spatially. The training data set covers 1979 to 2012, the validation data set covers 2013 to 2017, and the test data set covers 2018 to 2022. In order to verify the simulation performance, it is first necessary to prove that it can accurately identify EP from the complete precipitation data set. In this context, precision (P), recall (R), and F1 scores - indicators that are commonly used to evaluate the degree of match with a specified threshold - are used to evaluate the performance of the mathematical model, expressed as:
[0038]
[0039] Among them, TP (True Positive) represents true positive examples, FP (False Positive) represents false positive examples, and TN (True Negative) represents true negative examples.
[0040] From the perspective of interpretability, the random forest (RF) model provides a certain degree of interpretability compared to the traditional decision tree model. In regression tasks, random forest (RF) uses feature importance as a significance measure.
[0041] Feature importance is the average of feature importance in all individual trees, and the principle of feature importance calculation method for individual trees is: the sum of the reduction in squared loss after partitioning according to the feature. In addition, permutation importance is a model-independent feature importance calculation method. Its basic method is as follows: select a feature (variable), randomly permute all values of the feature in the data set, and calculate new simulation results. If the difference between the new and old results is small, it means that the feature is less important; if the difference is significant, it means that the feature has a greater impact on the model.
[0042] In order to further enhance the robustness of the variable selection results, Shapley regression values are used to explain the simulation results. Shapley regression values are based on game theory and local interpretation, revealing the complex relationship between input and output. The SHAP interpreter uses the Shapley regression value theory to explain the simulation results of the model by calculating the contribution of each feature to the model output. It not only reflects the impact of each feature in a single sample, but also reveals the positive and negative contributions of the feature. SHAP explanation can be applied to explain any machine learning model, including neural networks and integrated models, to provide more comprehensive and accurate explanation results. Shapley regression value φ i The expression is:
[0043]
[0044] The Shapley regression value is an important feature of linear models in the presence of multicollinearity. In this framework, F represents the set of all input meteorological variables (all meteorological variables), and S represents a subset of meteorological variables in F (a subset of all meteorological variables). This expression reflects the impact of variable i on the result when it is included in the set S (a certain meteorological variable) compared to the impact when it is excluded. S∪{i} (x S∪{i} ) represents the SHAP value of i contained in the set S, f s (x s ) represents the SHAP value after i is excluded.
[0045] Based on this interpretation, the SHAP method can be used to calculate the Shapley regression value for each individual variable and then rank them according to these values. As shown in the figure, in the ranking plot, the horizontal position of the variables indicates their importance in the simulation, while the color indicates the size of their effect: positive (red) or negative (blue) relative to the observed value.
[0046] S3. Use the most influential variables screened out in step S2 to re-simulate extreme precipitation and evaluate the effectiveness of the most influential variables in improving the simulation effect.
[0047] In the simulation, the present invention utilized all potential meteorological variables related to EP to train the machine learning model, aiming to use interpretability to explain which variables have a stronger relationship with the occurrence of EP.
[0048] In order to reduce redundancy and ensure efficient use of information, the best combination of variables for simulating EP was selected based on the r value and VIF value. As shown in the third part of the figure, by using the correlation coefficient r, the following variables were retained in the final model: U300, U850, V300, V850, ω300, ω500, ω850, MSLP, CAPE, IVT, TMQ, and pr_b. These variables were removed because their correlation with the retained variables exceeded 0.7 and their correlation with EP was weak. Subsequently, all the retained variables were retained because their VIF values were all below 5.
[0049] To explore the relationship between EP and its potential influencing factors, the interpretability metrics described in the third part of the figure are used to evaluate the importance of the screened variables. Among the three metrics, TMQ ranks the highest, followed by pr_b. When the values of TMQ and pr_b are high, their effects on EP are positive. This means that when the values of TMQ or pr_b are high, their presence leads to larger EP values when simulating EP with other variables. In addition, IVT, MSLP, and ω300 consistently show high importance scores in all three metrics. When the values of MSLP and ω300 are high, their effects on EP are negative. Variables with lower importance rankings and larger uncertainties include U, V, and CAPE.
[0050] To quantify the importance of each variable, a structured ranking approach was used.
[0051] First, the variables are ranked according to each individual interpretability metric. Next, the rankings for the three metrics are averaged to produce a final ranking. If two variables have the same average ranking, their correlation with EP is used as the tiebreaker, giving priority to the variable with a stronger correlation with EP. The final ranking of variable importance is: TMQ, pr_b, IVT, MSLP, ω300, V850, ω500, V300, U850, ω850, CAPE, and U300.
[0052] After variable selection and importance ranking, in order to verify the influence of variables with different importance levels on EP simulation, these variables were used in sequence according to the ranking order and added one by one for EP simulation, thereby reducing the inclusion of unnecessary variables. In addition, the relationship between each variable and EP was further analyzed.
[0053] In a preferred embodiment, the six factors that significantly affect EP are arranged in descending order: total water vapor quantity (TMQ), precipitation on the previous day (pr_b), total water vapor transport (IVT), sea level pressure (MSLP), 300hPa vertical velocity (ω300) and 850hPa meridional wind (V850).
[0054] The accuracy of EP simulated using these six variables is similar to that obtained by using all variables that may affect EP and is significantly better than ERA5.
[0055] When the variables that affect EP are not extreme, they have less impact on the strength of EP. The more extreme the values of the variables, the greater the strength of EP.
[0056] Not only is TMQ the factor that has the greatest impact on the occurrence of EP, but the emergence of extreme TMQ is usually accompanied by a stronger EP intensity.
[0057] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or a data center that includes one or more available media integrated. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid state disk (SSD)), etc.
[0058] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A method for simulating the meteorological impact of extreme precipitation in the Asian monsoon region based on artificial intelligence, characterized in that: The steps include: S1, preprocessing meteorological data to obtain a variable set; S2. Establish a mathematical model to simulate extreme precipitation, input the variable set obtained after preprocessing into the mathematical model in its entirety, compare the simulation performance before and after variable selection, and screen out the most influential variables; S3. Use the most influential variables screened out in step S2 to re-simulate extreme precipitation and evaluate the effectiveness of the most influential variables in improving the simulation effect.
2. The method for simulating the meteorological impact of extreme precipitation in the Asian monsoon region based on artificial intelligence according to claim 1, characterized in that: In step S1, the correlation coefficient between each pair of variables and the variance inflation factor VIF of all input variables are calculated. i , the expression is: Among them, Cov(X,Y) is the covariance of the two variables, Var(X), Var(Y) is the variance of variables X and Y, r is the correlation coefficient, r i is the correlation coefficient between the ith pair of variables, and p is the total number of variable pairs.
3. The method for simulating the meteorological impact of extreme precipitation in the Asian monsoon region based on artificial intelligence according to claim 1, characterized in that: In step S2, the expression for evaluating the performance of the mathematical model is: Among them, P is precision, R is recall and F1Score is score, TP stands for true positive, FP stands for false positive, and TN stands for true negative.
4. The method for simulating the meteorological impact of extreme precipitation in the Asian monsoon region based on artificial intelligence according to claim 3, characterized in that: Permutation importance is used as a feature importance calculation method: select a feature, randomly permute all values of the feature in the data set, and calculate new simulation results. If the difference between the new and old results is small, it means that the importance of the feature is low; if the difference is significant, it indicates that the feature has a great impact on the model.
5. The method for simulating the meteorological impact of extreme precipitation in the Asian monsoon region based on artificial intelligence according to claim 3, characterized in that: In step S2, the Shapley regression value is used to explain the simulation results. The Shapley regression value φ i The expression is: Where F represents the set of all input meteorological variables, S represents a subset of meteorological variables in F, and f S∪{i} (x S∪{i} ) represents the SHAP value of i contained in the set S, f s (x s ) represents the SHAP value after i is excluded.
6. The method for simulating the meteorological impact of extreme precipitation in the Asian monsoon region based on artificial intelligence according to claim 1, characterized in that: In step S3, the six factors that have a significant impact on precipitation meteorology are arranged in descending order: total water vapor TMQ, precipitation on the previous day pr_b, total water vapor transport IVT, sea level pressure MSLP, 300hPa vertical velocity ω300 and 850hPa meridional wind V850.
7. The method for simulating the meteorological impact of extreme precipitation in the Asian monsoon region based on artificial intelligence according to claim 1, characterized in that: In step S1, the preprocessing process includes: cleaning, resampling and normalizing all potential meteorological variables, while removing highly correlated input variables for dimensionality reduction to obtain a variable set.