Quantitative identification method of hail based on gaussian distribution of FY-2G satellite data
By constructing a spatiotemporal matching dataset and Gaussian distribution modeling, and combining single-indicator triggering and multi-indicator collaborative rules, the problem of insufficient regional adaptability and quantification of hail identification methods is solved, and accurate quantitative identification and early warning of hail occurrence probability are achieved.
Patent Information
- Application Number
- CN202510972514.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-07-15
AI Technical Summary
Existing hail identification methods have poor adaptability to different regions and insufficient quantification, making it difficult to meet the needs of refined early warning. Furthermore, relying on a single indicator or simple threshold judgment is easily affected by environmental factors, leading to misjudgment and missed judgment.
The method for quantitative hail identification based on FY-2G satellite data based on Gaussian distribution constructs a spatiotemporal matching dataset, uses quantile-quantile plots to verify data distribution characteristics, establishes a two-sample probability density function and confidence interval, and combines single-indicator triggering and multi-indicator collaborative rules to determine the probability of hail.
It improved the accuracy of hail identification by 20% to 30%, enhanced the method's adaptability to different data scenarios and its quantitative expression ability, provided more accurate weather warning support, and reduced losses.
Smart Images

Figure CN120470271B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of meteorological disaster early warning, and relates to a FY-2G satellite data hail quantitative identification method based on Gaussian distribution. BACKGROUND
[0002] Hail is a severe convective weather that frequently occurs in low-latitude plateau regions (such as Guizhou) in China, and it has the characteristics of strong locality, short life history, and great difficulty in early identification. This disaster often causes devastating damage to specialty agriculture such as tobacco and tea, and facility agriculture, and a single disaster can result in economic losses of hundreds of millions of yuan. Satellite observation has the advantages of wide coverage and real-time, and has become a core technical means for hail short-impending warning. Among them, FY-2G satellite is the main satellite for weather modification business in China, and the 7 items of inversion product data such as cloud top height (ztop), cloud top temperature (ttop), and supercooled layer thickness (hsc) provided by it can effectively reflect the macro and micro physical characteristics of hail cloud. However, the existing hail identification methods generally have the problems of insufficient regional adaptability and lack of quantitative indicators, resulting in significant fluctuations in identification accuracy under different meteorological backgrounds and terrain conditions, and cannot meet the needs of fine early warning.
[0003] Based on the stable index such as CAPE, TT index and ground radar data, the hail identification model constructed has a high dependence on the geographical and meteorological conditions of a specific region, and its performance is significantly reduced when applied across regions. For example, the stable index developed by European scholars for local meteorological conditions has poor effect in the prediction of convective weather in tropical and North America; the Logistic regression model based on radar and sounding data in northern Spain has an identification accuracy of 0.87 for hail, but the false positive rate is as high as 18%. In complex terrain regions in China, due to the existence of terrain shielding and radar coverage blind area, the false negative rate of traditional models often exceeds 20%.
[0004] Most of the existing methods use statistical models such as decision tree and Bayesian to make comprehensive judgments on multiple factors (such as literature), but lack single-index quantitative criteria based on data distribution characteristics, and the threshold setting relies on subjective experience, making it difficult to grade the probability of different intensity hail. Traditional methods mainly use a single data source such as radar or ground observation, and fail to fully exploit the multi-dimensional features of FY-2G satellite inversion product data. Although some studies attempt to introduce machine learning algorithms such as support vector machine (SVM), the prior analysis of data distribution is not enough, the model has poor interpretability, and the identification accuracy is limited in the scenario of satellite single-source data.
[0005] As patent number CN117055051A patent name is based on multi-source data hail identification method, system, device and storage medium, the scheme proposes to fuse three-dimensional radar, FY4B satellite, ERA5 environment field and other data, realizes identification through FEMU-Net network, and automatically extracts deep features of multi-source data by using neural network;However, this scheme depends on radar and other auxiliary data and large-scale labeled samples, and has high demand for computing resources;No quantitative index is designed for satellite single-source data, and it is difficult to adapt to the rapid early warning demand of small-scale disasters.
[0006] The present application proposes a single-index quantitative modeling and multi-index collaborative identification framework based on Gaussian distribution. The technical core is to verify the Gaussian distribution adaptability of FY-2G satellite multi-item inversion product data through quantile-quantile plot (Q-Q plot), to establish a quantitative identification index of "low-middle-high" three probability gradients by using double-sample probability density function and 95% confidence interval, and to realize comprehensive determination through single-index triggering and multi-index collaborative rules. This framework has the advantages of quantitative and interpretability, regional adaptability optimization, and lightweight and real-time, which can effectively solve the problems in production practice. SUMMARY
[0007] The present application provides a FY-2G satellite data hail quantitative identification method based on Gaussian distribution, which can solve the problems of poor regional adaptability and insufficient quantitative degree of existing hail identification methods, and realize single-index and multi-index joint hail probability quantitative determination by constructing spatiotemporal matching dataset, Gaussian distribution modeling and confidence interval division.
[0008] To solve the above problems, the technical scheme adopted by the application is:
[0009] The FY-2G satellite data hail quantitative identification method based on Gaussian distribution comprises the following steps: S1. Constructing spatiotemporal matching dataset: obtaining the inversion product data of cloud top height, cloud top temperature, supercooled layer thickness, optical thickness, effective particle radius, liquid water path and blackbody brightness temperature of FY-2G satellite, and extracting the satellite inversion data within a period of time before and after the hail falling point as hail sample; selecting the same number of satellite inversion data in non-hail period and non-hail area in space as non-hail sample to form a dataset containing spatiotemporal matching features; S2. Gaussian distribution adaptability test: using quantile-quantile plot to test the distribution characteristics of the inversion product data, taking the coincidence degree of quantile points and standard normal distribution straight line as the judgment basis to determine that the cloud top height, cloud top temperature and supercooled layer thickness conform to Gaussian distribution, and the optical thickness, liquid water path and blackbody brightness temperature approximately conform to Gaussian distribution; S3. Double-sample probability density modeling: calculating the mathematical expectation of hail sample and non-hail sample respectively for the inversion product data conforming or approximately conforming to Gaussian distribution and standard deviation , construct the bivariate Gaussian distribution probability density function:
[0010] ; S4. Confidence interval dynamic division: establish unilateral confidence interval for hail sample and non-hail sample respectively:
[0011] Calculate the lower bound b of hail sample, get the high hail probability interval (b, +∞);
[0012] Calculate the upper bound a of non-hail sample, get the low hail probability interval (-∞, a);
[0013] According to the interval overlap part, define the hail probability interval , form the quantitative identification index containing three-level probability gradient; S5. Multi-dimensional joint identification: match the corresponding three-level probability interval of step S4 to the inversion product data of the sample to be identified, and determine the hail probability by the following rules:
[0014] Single index trigger rule: if any index falls into the high hail probability interval, it is directly determined as high probability hail;
[0015] Multi-index collaborative rule: if at least 3 indexes fall into the medium hail probability interval and no index falls into the low probability interval, it is determined as medium probability hail.
[0016] The principle and advantages of the scheme are:
[0017] By constructing the spatiotemporal matching dataset, collecting multiple inversion product data of FY-2G satellite, including cloud top height, cloud top temperature, etc., taking the ground observation hail point as the core, selecting hail sample and non-hail sample, ensuring that the data has spatiotemporal matching characteristics, and laying a foundation for subsequent analysis. Secondly, the quantile-quantile diagram is used to test the Gaussian distribution adaptability of the inversion product data, and by comparing the coincidence degree of the quantile points and the standard normal distribution straight line, the data indexes conforming or approximately conforming to the Gaussian distribution are selected, providing reliable basis for probability density modeling. Then, for the selected indexes, the mathematical expectation and standard deviation of the hail sample and non-hail sample are calculated, and the bivariate Gaussian distribution probability density function is constructed to quantify the data distribution characteristics of the two types of samples. Then, unilateral confidence interval is established for hail sample and non-hail sample respectively, and high, medium and low three-level hail probability intervals are divided to form quantitative identification index. Finally, through single index trigger rule and multi-index collaborative rule, multi-dimensional joint identification of hail probability is realized, when any index falls into the high hail probability interval or at least 3 indexes fall into the medium hail probability interval and no index falls into the low probability interval, the hail probability is determined.
[0018] Compared with the prior art, in the identification accuracy, the existing method depends on a single index or a simple threshold judgment, which is easily disturbed by environmental factors, leading to misjudgment and missed judgment. For example, only according to the cloud top temperature to judge the hail, the cloud layer with low temperature but without hail may be misjudged as hail. However, the present scheme can more comprehensively capture the characteristics of hail formation by multi-dimensional data fusion and Gaussian distribution modeling, and can greatly improve the identification accuracy by using single index triggering and multi-index cooperative rule. According to the actual data verification, under complex weather conditions, the hail identification accuracy of the present scheme is about 20% to 30% higher than that of the traditional method. In terms of reliability, the prior art lacks in-depth analysis of the data distribution characteristics, and it is difficult to cope with data fluctuations. The present scheme can enhance the adaptability of the identification method to different data scenarios and reduce the risk of false judgment caused by data fluctuations by scientific screening of data indicators and establishment of probability density model through Gaussian distribution adaptability test. In terms of quantification, the prior art is mostly qualitative or semi-quantitative judgment, and it is difficult to provide accurate hail occurrence probability. The present scheme realizes the quantitative expression of hail occurrence probability by constructing a three-level probability interval, providing more accurate and intuitive data support for meteorological warning and decision-making, so that the meteorological department can more reasonably arrange disaster prevention and reduction work, and effectively reduce the loss caused by hail disasters.
[0019] Further, the method for constructing the spatio-temporal matching data set in step S1 is: taking 184 hail points on 30 hail days as the center, extracting the corresponding satellite inversion data at the moment according to the latitude and longitude coordinates, and selecting non-hail point data by spatial random sampling method to ensure that the two types of samples are aligned in time resolution and spatial resolution. The data is extracted from the center of the hail point, which can accurately capture the core characteristics of the hail cloud system, such as high cloud top height, low cloud top temperature and other key parameters, so that the hail sample directly reflects the true environment when the hail occurs; spatial random sampling selects non-hail point data, which can avoid human selection bias and ensure that the non-hail sample covers a wide range of non-hail cloud types, such as stratiform cloud, ordinary cumulus cloud, etc., truly reflecting the data distribution of non-hail scene. In terms of spatio-temporal consistency, strictly aligning the time resolution and spatial resolution can eliminate false correlations caused by spatio-temporal misalignment, such as avoiding forcibly matching non-hail data at different times or in different regions with hail points, ensuring that the two types of samples are under the same spatio-temporal reference, making the subsequent statistical test and probability modeling based on Gaussian distribution more scientific. From the model training effect, the data set constructed by this method has balanced spatio-temporal characteristics, which can effectively improve the ability of the model to distinguish between hail and non-hail scenes. For example, when comparing the probability density distribution of cloud top temperature and other indicators, it can clearly show the significant difference between hail and non-hail samples, avoiding model misjudgment caused by spatio-temporal confusion, and laying a reliable data foundation for the establishment of subsequent quantitative identification indicators.
[0020] Further, the determination criterion of the Gaussian distribution adaptability test in step S2 is: when the Pearson correlation coefficient of the quantile point of the inversion product data and the standard normal distribution straight line is ≥0.9, it is determined to be subject to Gaussian distribution; when the correlation coefficient is between 0.8 and 0.9, it is determined to be approximately subject to Gaussian distribution. The Pearson correlation coefficient quantifies the linear correlation degree of the quantile point and the standard normal distribution straight line through a mathematical formula, avoiding subjective judgment errors, such as avoiding ambiguous conclusions caused by visual coincidence of Q-Q plots, so that the test result is more credible. In terms of operability, this standard provides a clear numerical threshold, which facilitates researchers to quickly judge the data distribution type. For example, when dealing with inversion product data such as cloud top height and cloud top temperature, the correlation coefficient can be directly classified by calculation, improving the test efficiency. The correlation coefficient ≥0.9 indicates that the linear fitting degree of the data points and the standard normal distribution is very high, such as cloud top height, cloud top temperature, and supercooled layer thickness, and the distribution characteristics can be directly modeled using the probability density function and confidence interval of Gaussian distribution. For the correlation coefficient interval of 0.8-0.9, such as optical thickness and liquid water path, although there is a certain deviation, the overall distribution trend is close to Gaussian distribution, and the statistical characteristics of Gaussian distribution can still be used to establish quantitative identification indicators through approximate processing, which widens the data applicability while ensuring the model accuracy, avoids information loss caused by strictly excluding approximate distribution data, and provides a reliable basis for subsequent two-sample probability density modeling and confidence interval division.
[0021] Further, the step S3 of modeling the probability density of the two samples further includes long-tail distribution correction of optical thickness, liquid water path, and blackbody brightness temperature: by introducing a truncation threshold, the mathematical expectation and standard deviation are recalculated after removing outliers, and the fitting accuracy of the model for non-ideal distribution data is improved. The outliers in the long-tail distribution, such as sudden increase in optical thickness or extremely low liquid water path, may be caused by satellite observation errors, cloud boundary interference, and other factors, and are not characteristic data of real hail or non-hail scenes. By setting a truncation threshold, such as removing data points that are more than ±3 times the standard deviation from the mean based on the 3σ principle, such noise can be effectively filtered out, avoiding distortion of the overall distribution parameters such as the mathematical expectation and standard deviation. For example, if there are individual extremely high non-hail samples in the liquid water path data, the mathematical expectation of the non-hail data set may be artificially high, causing the subsequent confidence interval division to deviate from the true distribution. After removal, the probability density curve of the non-hail sample can be closer to the actual liquid water content characteristics of the cloud system. From the model fitting effect, the corrected mathematical expectation and standard deviation can more accurately reflect the concentration trend and dispersion degree of the data main body. Taking the blackbody brightness temperature as an example, after removing the abnormally high temperature values in the hail sample caused by instrument failure, the mathematical expectation of the hail data set is closer to the true low temperature characteristics, and the peak position and distribution range of the probability density curve can better reflect the temperature characteristics of the strong convective cloud top, thereby enhancing the recognition of the distribution difference between hail and non-hail samples in the two-sample modeling, and improving the rationality of the subsequent confidence interval division and the reliability of the identification index. This correction mechanism optimizes the applicability of the Gaussian distribution model while preserving the distribution characteristics of the data main body, effectively balancing data integrity and model accuracy, especially for parameters that approximately follow a Gaussian distribution.
[0022] Further, the step S4 of dynamically dividing the confidence interval by numerically integrating the probability density function, satisfies: the one-sided confidence interval [b, +∞) of the hail sample, satisfies ; the one-sided confidence interval (-∞, a] of the non-hail sample, satisfies , where and are the Gaussian probability density functions of the hail and non-hail samples, respectively. The area under the probability density curve is directly calculated by mathematical integration to ensure that the confidence interval [b, +∞) of the hail sample strictly covers 95% of the hail data, for example, the cloud top height of the hail sample is determined by integration to be b = 7.20, so that the interval contains 95% of the cloud top height values of the hail points, and the confidence interval (-∞, ]Precise contains 95% non-hail data, such as cloud top temperature non-hail sample calculation a = -42.72, avoid the subjectivity of setting threshold. Dynamic adaptability, for different inversion product data probability density characteristics, such as supercooled layer thickness, liquid water path, numerical integral method can be flexible to adapt to the difference of data distribution, for example, supercooled layer thickness non-hail sample is limited by the maximum value of 8 kilometers, through the integral combined with data truncation adjustment , make the confidence interval more close to the real distribution. Statistical reliability, based on 95% confidence interval division conforms to the statistical standard, high hail probability interval , + ∞) only contains 5% non-hail data, can be used as a single index to trigger high probability hail, such as black body brightness temperature ≤-37.70℃ directly determine high probability hail; Medium hail probability interval , b] covers the overlapping area of data, through the multi-index coordination rule, at least 3 indicators fall into the interval and no low probability indicator comprehensive judgment, effectively reduce the risk of misjudgment caused by single index data overlap.
[0023] Further, the identification index of supercooled layer thickness in step S4 is optimized as follows: combined with the constraint condition that the maximum value of supercooled layer thickness of non-hail points in actual observation data is 8 kilometers, the upper bound a of non-hail sample is forcibly limited to 8 kilometers, forming the modified medium hail probability interval [4.20, 8] and high hail probability interval (8, + ∞). The maximum value of supercooled layer thickness of non-hail points in actual observation is 8 kilometers, which indicates that when the index exceeds 8 kilometers, the cloud system has broken through the physical boundary of non-hail scene. By forcibly limiting = 8 kilometers, the misjudgment caused by the fact that the pure mathematical integral result deviates from the actual observation can be avoided, for example, the false data scene of “non-hail point supercooled layer thickness exceeding 8 kilometers” is excluded, so that the non-hail sample confidence interval (-∞, 8] strictly fits the observation fact, and the credibility of the index is enhanced; The modified high hail probability interval (8, + ∞) directly corresponds to the physical characteristics of “supercooled layer thickness exceeding the possible range of non-hail cloud system”, for example, when the supercooled layer thickness increases significantly when the strong convective cloud develops vigorously, it can be determined as high probability hail signal when it exceeds 8 kilometers, reducing the index window problem in the original mathematical solution (8, 14.18] interval because there is no actual non-hail data support. The medium hail probability interval [4.20, 8] covers the overlapping area of supercooled layer thickness, which contains part of the hail sample, accounting for 95% of the hail data that is not covered by the high probability interval, and also contains the extreme value of non-hail sample. Through the multi-index coordination rule, false alarm can be further filtered, for example, when other indicators such as cloud top temperature and liquid water path fall into the medium and high probability interval, it is determined as medium probability hail, which improves the rigor of the identification logic.
[0024] Further, the multi-dimensional joint recognition in step S5 further includes a weight distribution mechanism: the indicators of cloud top temperature and blackbody brightness temperature are assigned a weight, and the weight coefficient is ≥0.2; the auxiliary indicators of optical thickness and liquid water path are assigned a weight, and the weight coefficient is ≤0.1; and the comprehensive hail reduction probability value is calculated by weighted summation.
[0025] Further, the method further includes a model verification step: the effect of the quantitative identification indicators is tested by using leave-one-out cross-validation (LOOCV) to calculate the hail sample recognition accuracy, the non-hail sample recognition accuracy, and the overall recognition accuracy; the data set is separated into a training set and a test set one by one by LOOCV, and each time, one sample is reserved as the test set, and the remaining samples are used as the training set, so that each sample participates in model training and verification, and evaluation deviation caused by random sampling is avoided. For example, in 368 samples, each hail point and non-hail point is individually reserved for testing, so that the hail sample recognition accuracy and the non-hail sample recognition accuracy, such as 91.85%, can truly reflect the distinguishing ability of the model for the two types of data, and the overall recognition accuracy, such as 87.50%, comprehensively reflects the overall effectiveness of the indicators. From the perspective of reliability verification, the method can quantitatively calculate the stability of the model by multiple iterations. For example, when verifying the SVM model based on the RBF kernel function, the leave-one-out result shows that the standard deviation of the non-hail sample recognition accuracy is low, indicating that the model has strong stability in distinguishing non-hail scenes; and the fluctuation range of the hail sample recognition accuracy is small, indicating that the indicators have consistency in capturing hail characteristics. In addition, by comparing the cross-validation results of different kernel function models, such as L-SVM, RBF-SVM, and S-SVM, the optimal model can be selected, such as RBF-SVM, which has the highest overall accuracy, thereby avoiding model selection deviation caused by single training-test division.
[0026] Further, the method further integrates a support vector machine (SVM) model: the inversion product data is used as the input feature, and the Linear, RBF, and Sigmoid kernel functions are used to train the classification model, the hyperparameters are optimized by grid search, and finally, the RBF-SVM model is selected as the final model. The SVM model maps the low-dimensional input features, such as cloud top height and cloud top temperature, to a high-dimensional space through the kernel function, effectively solving the classification problem of nonlinear characteristics in hail recognition. For example, the optical thickness and the liquid water path have high overlap in low-dimensional space, and it is difficult to distinguish them by a linear boundary, but the RBF kernel function can calculate the similarity between samples through a Gaussian radial basis function, and can construct a nonlinear decision boundary in high-dimensional space to accurately capture the hidden feature differences between hail and non-hail samples.
[0027] Further, a sample equalization process is introduced in the training process of the SVM model: synthetic samples are generated for the few-class hail sample using the SMOTE algorithm to make the ratio of hail samples to non-hail samples reach 1:1. Although the number of hail samples and non-hail samples in the original data set is balanced, hail events in the actual meteorological scene are usually small probability events, and the real data may be naturally unbalanced, such as the number of non-hail samples being much more than that of hail samples. The SMOTE algorithm generates synthetic hail samples by interpolation, such as generating virtual samples based on k-nearest neighbor characteristics, which can simulate the feature space distribution of real hail cloud systems, avoid the model from being biased to learn the characteristics of the majority class due to the sparsity of the few-class samples, and overfit the characteristics of non-hail samples, such as the low cloud top height and high cloud top temperature. For example, in the feature space of optical thickness and other indicators, hail samples may form several isolated clusters, and the samples generated by SMOTE can fill the gaps between the clusters, enhancing the overall understanding of the model to the hail feature distribution. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 Flowchart of the present application. DETAILED DESCRIPTION
[0029] Example 1, a Gaussian distribution-based FY-2G satellite data hail quantification identification method, includes the following steps: S1. Constructing a spatio-temporal matching data set: obtaining the retrieval product data of FY-2G satellite cloud top height, cloud top temperature, supercooled layer thickness, optical thickness, effective particle radius, liquid water path, and blackbody brightness temperature, and extracting the satellite retrieval data within a period of time before and after the hail time as the hail sample with the ground observed hail point as the center; selecting the same number of satellite retrieval data in non-hail period and non-hail area in space as non-hail samples to form a data set containing spatio-temporal matching features; S2. Gaussian distribution adaptability test: using quantile-quantile diagram to test the distribution characteristics of the retrieval product data, taking the coincidence degree of quantile points and standard normal distribution straight line as the judgment basis to determine that the cloud top height, cloud top temperature, and supercooled layer thickness obey the Gaussian distribution, and the optical thickness, liquid water path, and blackbody brightness temperature approximately obey the Gaussian distribution; S3. Double-sample probability density modeling: calculating the mathematical expectation and standard deviation of the hail sample and the non-hail sample respectively , for the retrieval product data obeying or approximately obeying the Gaussian distribution to construct the double-sample Gaussian distribution probability density function:
[0030] calculating the lower bound b of the hail sample to obtain the high hail probability interval (b, +∞);
[0031] calculating the upper bound a of the non-hail sample to obtain the low hail probability interval (-∞, a);
[0032] According to the interval overlap part definition hail probability interval , forming a quantitative identification index containing a three-level probability gradient, the numerical integration method uses Simpson's rule to discretize the probability density function for calculation, and the integral threshold a and b are solved by iterative approximation to ensure that the interval covers 95% of the sample data. S5. Multi-dimensional joint identification: The inversion product data of the sample to be identified are matched with the corresponding three-level probability interval in step S4, and the hail occurrence probability is determined by the following rules:
[0033] Single-index trigger rule: If any index falls into the high hail probability interval, it is directly determined as high probability hail;
[0034] Multi-index coordination rule: If at least 3 indexes fall into the medium hail probability interval and no index falls into the low probability interval, it is determined as medium probability hail.
[0035] By constructing a spatio-temporal matching data set, collecting FY-2G satellite multi-reverse product data including cloud top height, cloud top temperature, etc., and selecting hail samples and non-hail samples with core ground observation hail points, the data is ensured to have spatio-temporal matching characteristics, laying a foundation for subsequent analysis. Secondly, the quantile-quantile diagram is used to test the Gaussian distribution adaptability of the inversion product data, and by comparing the coincidence degree of the quantile points and the standard normal distribution straight line, the data indexes conforming or approximately conforming to the Gaussian distribution are selected, providing reliable basis for probability density modeling. Then, for the selected indexes, the mathematical expectation and standard deviation of hail samples and non-hail samples are calculated, and the two-sample Gaussian distribution probability density function is constructed to quantify the data distribution characteristics of the two types of samples. Then, one-sided confidence intervals are established for hail samples and non-hail samples respectively, and high, medium and low three-level hail probability intervals are divided to form quantitative identification indexes. Finally, through single-index trigger rule and multi-index coordination rule, multi-dimensional joint identification of hail occurrence probability is realized. When any index falls into the high hail probability interval or at least 3 indexes fall into the medium hail probability interval and no index falls into the low probability interval, the hail occurrence probability can be determined.
[0036] In terms of recognition accuracy, existing methods rely on a single indicator or simple threshold judgment, which is easily disturbed by environmental factors, leading to misjudgment and missed judgment. For example, judging hail only according to cloud top temperature may misjudge low-temperature cloud layers without hail as hail. However, the present scheme uses multi-dimensional data fusion and Gaussian distribution modeling to construct a three-level probability gradient recognition indicator, uses single indicator triggering and multi-indicator coordination rules, can more comprehensively capture the characteristics of hail formation, and greatly improves the recognition accuracy. According to the actual data verification, under complex weather conditions, the hail recognition accuracy of the present scheme is about 20% to 30% higher than that of the traditional method. In terms of reliability, existing technologies lack in-depth analysis of data distribution characteristics, making it difficult to cope with data fluctuations. The present scheme uses Gaussian distribution adaptability test to scientifically select data indicators and establish probability density models, enhancing the adaptability of the recognition method to different data scenarios and reducing the risk of false judgments caused by data fluctuations. In terms of quantitative degree, existing technologies are mostly qualitative or semi-quantitative judgments, and it is difficult to provide accurate hail occurrence probability. The present scheme realizes the quantitative expression of hail occurrence probability by constructing a three-level probability interval, providing more accurate and intuitive data support for meteorological warning and decision-making, enabling meteorological departments to more reasonably arrange disaster prevention and mitigation work, and effectively reducing the loss caused by hail disasters.
[0037] The construction method of the spatiotemporal matching data set in step S1 is: taking 184 hail points on 30 hail days as the center, extracting corresponding satellite inversion data at the moment according to the latitude and longitude coordinates, and selecting non-hail point data through spatial random sampling method to ensure that the two types of samples are aligned in time resolution and spatial resolution. Extracting data from the center of the hail point can accurately capture the core characteristics of the hail cloud system, such as high cloud top height, low cloud top temperature, and other key parameters, so that the hail sample directly reflects the true environment when hail occurs; spatial random sampling selects non-hail point data, which can avoid human selection bias and ensure that non-hail samples cover a wide range of non-hail cloud types, such as stratiform clouds, ordinary cumulus clouds, etc., truly reflecting the data distribution of non-hail scenes. In terms of spatiotemporal consistency, strictly aligning the time resolution and spatial resolution can eliminate false correlations caused by spatiotemporal misalignment, such as avoiding forcibly matching non-hail data at different times or in different regions with hail points, ensuring that the two types of samples are under the same spatiotemporal reference, making subsequent statistical tests and probability modeling based on Gaussian distribution more scientific. From the model training effect, this construction method forms a balanced data set in terms of spatiotemporal characteristics, which can effectively improve the model's ability to distinguish between hail and non-hail scenes. For example, when comparing the probability density distribution of cloud top temperature indicators, it can clearly show the significant difference between hail and non-hail samples, avoiding model misjudgment caused by spatiotemporal confusion, and laying a reliable data foundation for the establishment of subsequent quantitative recognition indicators.
[0038] The determination criterion of the "Gaussian distribution fitness test" in the step S2 is: when the Pearson correlation coefficient of the quantile points of the inversion product data and the standard normal distribution straight line is greater than or equal to 0.9, it is determined that the data is subject to Gaussian distribution; when the correlation coefficient is between 0.8 and 0.9, it is determined that the data is approximately subject to Gaussian distribution. The Pearson correlation coefficient quantifies the linear correlation degree of the quantile points and the standard normal distribution straight line through a mathematical formula, avoiding subjective judgment errors, such as avoiding the ambiguous conclusions caused by visual coincidence of Q-Q plots, so that the test results are more credible. In terms of operability, the standard provides a clear numerical threshold, which is convenient for researchers to quickly judge the data distribution type. For example, when dealing with inversion product data such as cloud top height and cloud top temperature, the correlation coefficient can be directly classified by calculation, improving the test efficiency. The correlation coefficient greater than or equal to 0.9 indicates that the linear fitting degree of the data points and the standard normal distribution is very high, such as cloud top height, cloud top temperature and supercooled layer thickness, and the distribution characteristics can be directly modeled by using the probability density function and confidence interval of Gaussian distribution. For the correlation coefficient between 0.8 and 0.9, such as optical thickness and liquid water path, although there is a certain deviation, the overall distribution trend is close to Gaussian distribution, and the statistical characteristics of Gaussian distribution can still be used to establish quantitative identification indicators through approximate processing, which widens the data applicability while ensuring the model accuracy, avoids information loss caused by strictly excluding approximate distribution data, and provides a reliable basis for subsequent two-sample probability density modeling and confidence interval division.
[0039] Step S3, the dual-sample probability density modeling, further includes correction of the long-tail distribution of optical thickness, liquid water path, and blackbody brightness temperature: by introducing a truncation threshold, outliers are removed, and the expected value and standard deviation are recalculated, improving the model's fitting accuracy to non-ideal distribution data. Outliers in the long-tail distribution, such as sudden spikes in optical thickness or abnormally low values in liquid water path, may be caused by factors such as satellite observation errors and cloud boundary interference, and are not characteristic data of real hail or non-hail scenarios. By setting a truncation threshold, such as removing data points exceeding the mean ± 3 times the standard deviation based on the 3σ principle, such noise can be effectively filtered out, avoiding the distortion of overall distribution parameters, such as expected value and standard deviation, by outliers. For example, if there are a few extremely high non-hail samples in the liquid water path data, it may cause the expected value of the non-hail dataset to be artificially high, leading to subsequent confidence interval divisions deviating from the true distribution. Removing these samples makes the probability density curve of the non-hail samples closer to the actual liquid water content characteristics of the cloud system. From the model fitting results, the corrected expected value and standard deviation more accurately reflect the central tendency and dispersion of the main data. Taking blackbody brightness temperature as an example, after removing abnormally high temperature values caused by instrument malfunction in the hailfall samples, the expected value of the hailfall dataset is closer to the true low temperature characteristics. The peak position and distribution amplitude of its probability density curve better reflect the temperature characteristics of strong convective cloud tops. This enhances the ability to distinguish the distribution differences between hailfall and non-hailfall samples in dual-sample modeling, improves the rationality of subsequent confidence interval division and the reliability of identification indicators. This correction mechanism, while preserving the distribution characteristics of the main data, optimizes the applicability of the Gaussian distribution model through outlier control, especially for parameters that approximately follow a Gaussian distribution, effectively balancing data integrity and model accuracy.
[0040] The integral calculation method for the dynamic division of confidence intervals in step S4 is as follows: The probability density function is solved using numerical integration to ensure that the confidence interval for hail samples contains 95% of the hail data, and the confidence interval for non-hail samples contains 95% of the non-hail data. The specific formula is as follows: , By directly calculating the area under the probability density curve through mathematical integration, the confidence interval [b, +∞) for hail samples is ensured to strictly cover 95% of hail data. For example, for hail samples with cloud top height, b=7.20 is determined through integration, making this interval include 95% of the cloud top height values for hail points. The confidence interval (-∞, φ) for non-hail samples precisely includes 95% of the non-hail data. For example, the cloud top temperature for non-hail samples is calculated as follows: =-42.72, avoiding the arbitrariness of subjectively setting thresholds. In terms of dynamic adaptability, the numerical integration method can flexibly adapt to differences in data distribution, considering the probability density characteristics of different inversion product data, such as supercooled layer thickness and liquid water path. For example, the maximum supercooled layer thickness for non-hail samples is limited to 8 kilometers by actual observations; this can be adjusted by combining integration with data truncation. This makes the confidence intervals more closely resemble the true distribution. In terms of statistical reliability, the interval division based on a 95% confidence level meets statistical standards, and the high probability of hail intervals ( The +∞) range contains only 5% of non-hail data, which can be used as a strong criterion for determining a high probability of hail, such as directly determining a high probability of hail when the blackbody brightness temperature is ≤-37.70℃; the medium probability range of hail [ [b] Covers overlapping data areas. Through multi-indicator coordination rules, at least three indicators fall into the range and there are no low-probability indicators for comprehensive judgment, effectively reducing the risk of misjudgment caused by data overlap of a single indicator.
[0041] In step S4, the optimization of the supercooled layer thickness identification index is as follows: Based on the constraint that the maximum supercooled layer thickness at non-hail-falling points in actual observation data is 8 kilometers, the supremacy 'a' of the non-hail-falling samples is forcibly limited to 8 kilometers, forming a corrected medium hail probability interval [4.20, 8] and a high hail probability interval (8, +∞). The maximum supercooled layer thickness at non-hail-falling points in actual observation is 8 kilometers, indicating that when this index exceeds 8 kilometers, the cloud system has broken through the physical boundary of the non-hail-falling scenario. This is achieved through forced limitation. =8 km, which avoids misjudgment caused by the pure mathematical integration result deviating from actual observation. For example, it excludes the spurious data scenario of "the supercooled layer thickness of non-hail points exceeds 8 km", and makes the confidence interval (-∞, 8] of non-hail samples strictly fit the observation facts, enhancing the credibility of the indicator. The corrected high hail probability interval (8, +∞) directly corresponds to the physical characteristic of "the supercooled layer thickness exceeding the possible range of non-hail cloud system". For example, when strong convective clouds develop vigorously, the supercooled layer thickness increases significantly. When it exceeds 8 km, it can be clearly identified as a high probability hail signal, reducing the original number. The (8, 14, 18) interval in the learning solution suffers from a gap in indicators due to the lack of actual non-hailfall data. The medium hailfall probability interval [4, 20, 8] covers the overlapping area of supercooled layer thickness, including both hailfall samples (95% of the hailfall data not covered by the high probability interval) and extreme values of non-hailfall samples. False alarms can be further filtered through multi-indicator collaborative rules. For example, when other indicators, such as cloud top temperature and liquid water path, simultaneously fall into the medium-high probability interval, the result is comprehensively judged as medium-probability hailfall, improving the rigor of the identification logic.
[0042] The multi-dimensional joint recognition in the step S5 further comprises a weight distribution mechanism: the cloud top temperature and the blackbody brightness temperature are assigned a weight, and the weight coefficient is greater than or equal to 0.2; the optical thickness and the liquid water path are assigned a weight, and the weight coefficient is less than or equal to 0.1; the comprehensive hail probability value is calculated by weighted summation; and the weight distribution mechanism is determined based on the physical contribution of the indicators to hail formation: the cloud top temperature and the blackbody brightness temperature directly reflect the thermodynamic characteristics of the cloud top and are closely related to the supercooled water environment for hail growth, so a higher weight of 0.2-0.3 is assigned; and the optical thickness and the liquid water path reflect the microphysical structure of the cloud system and are assigned a lower weight of 0.05-0.1 as auxiliary criteria.
[0043] The method further comprises a model verification step: the effect of the quantitative recognition indicators is tested by using leave-one-out cross-validation (LOOCV) to calculate the hail sample recognition accuracy, the non-hail sample recognition accuracy and the overall recognition accuracy; the data set is separated into a training set and a test set one by one by LOOCV, one sample is reserved as the test set each time, and the remaining samples are used as the training set, so that each sample participates in model training and verification, and evaluation bias caused by random sampling is avoided. For example, in 368 samples, each hail point and non-hail point is reserved for testing, so that the hail sample recognition accuracy is 89.13%, the non-hail sample recognition accuracy is 91.85%, the overall recognition accuracy is 87.50%, the discrimination ability of the model for the two types of data is truly reflected, and the overall effectiveness of the indicators is comprehensively reflected. From the perspective of reliability verification, the method can quantitatively calculate the stability of the model by multiple iterations. For example, when verifying the SVM model based on the RBF kernel function, the leave-one-out result shows that the standard deviation of the non-hail sample recognition accuracy is low, indicating that the model has strong stability in discriminating non-hail scenes; and the fluctuation range of the hail sample recognition accuracy is small, indicating that the indicators have consistency in capturing hail characteristics. In addition, by comparing the cross-validation results of different kernel function models, such as L-SVM, RBF-SVM and S-SVM, the optimal model can be selected, such as RBF-SVM with the highest overall accuracy, which avoids the model selection bias caused by single training-test division.
[0044] The method further integrates a support vector machine (SVM) model: taking the inversion product data as input features, respectively using Linear, RBF, Sigmoid kernel functions to train the classification model, and optimizing the hyperparameters through grid search, finally selecting the RBF-SVM model as the final model. The SVM model maps the low-dimensional input features such as cloud top height, cloud top temperature, and other multiple inversion product data to a high-dimensional space through the kernel function, effectively solving the classification problem of non-linear features in hail identification. For example, the distribution of optical thickness and liquid water path and other indicators in low-dimensional space has a high degree of overlap, and it is difficult to distinguish by linear boundary, while the RBF kernel function calculates the similarity between samples through Gaussian radial basis function, and can construct a non-linear decision boundary in high-dimensional space to accurately capture the implicit feature differences between hail and non-hail samples.
[0045] During the training process of the SVM model, sample equalization processing is introduced: SMOTE algorithm is used to generate synthetic samples for the minority class hail sample, so that the ratio of hail sample to non-hail sample reaches 1:1. Although the number of hail samples and non-hail samples in the original data set is balanced, hail events in actual meteorological scenarios are usually small probability events, and real data may have natural imbalance such as non-hail samples much more than hail samples. The SMOTE algorithm generates synthetic hail samples by interpolation, such as generating virtual samples based on k-nearest neighbor features, which can simulate the feature space distribution of real hail cloud systems, avoiding the model from learning the characteristics of the majority class due to the sparsity of the minority class samples, such as overfitting the low cloud top height, high cloud top temperature and other characteristics of non-hail samples. For example, in the feature space of optical thickness and other indicators, hail samples may form several isolated clusters, and the samples generated by SMOTE can fill the gaps between clusters, enhancing the model's overall understanding of hail feature distribution.
[0046] Embodiment 2
[0047] Construction of spatiotemporal matching data set
[0048] Data acquisition:
[0049] FY-2G satellite inversion product data of 30 hail days from 2020 to 2022 were collected, including cloud top height (ztop), cloud top temperature (ttop), supercooled layer thickness (hsc), optical thickness (optn), effective particle radius (ref), liquid water path (lwp), and blackbody brightness temperature (tbb).
[0050] Taking 184 hail points observed on the ground as the center, satellite inversion data within 15 minutes before and after the hail time were extracted according to the latitude and longitude coordinates, as 184 groups of hail samples.
[0051] Non-hail sample selection: Through spatial random sampling method, in non-hail period and non-hail area, randomly select 184 groups of satellite retrieval data with spatial distance ≥ 50 kilometers from hail point, ensure that the time resolution is consistent with the hail sample time window, and the spatial resolution (0.05°×0.05° grid) is aligned with the hail sample.
[0052] Data preprocessing: Eliminate samples that fail to match space-time, such as missing satellite data or abnormal quality marks, finally form a space-time matching data set containing 368 samples, 184 hail + 184 non-hail.
[0053] 2. Gaussian distribution fitness test
[0054] Test method: Draw quantile-quantile graph (Q-Q graph) for 7 items of retrieval product data, compare sample quantile points with standard normal distribution straight line.
[0055] Calculate Pearson correlation coefficient (PCC) to quantify the degree of coincidence:
[0056] PCC≥0.9: Determine as "subject to Gaussian distribution", such as cloud top height, cloud top temperature, supercooled layer thickness;
[0057] 0.8≤PCC<0.9: Determine as "approximately subject to Gaussian distribution", such as optical thickness, liquid water path, blackbody brightness temperature;
[0058] PCC<0.8: Exclude use, such as effective particle radius ref.
[0059] Example:
[0060] The PCC of cloud top height (ztop) is 0.92, which is determined as subject to Gaussian distribution;
[0061] The PCC of optical thickness (optn) is 0.85, which is determined as approximately subject to Gaussian distribution.
[0062] 3. Double-sample probability density modeling and long-tail correction
[0063] Modeling steps: For the 6 indicators subject to / approximately subject to Gaussian distribution (excluding ref), calculate the mathematical expectation μ and standard deviation σ of hail sample and non-hail sample respectively, and construct double-sample Gaussian distribution probability density function: ,
[0064] For example, the mean of cloud top temperature (ttop) of hail sample is -46.30℃, the standard deviation is 14.09℃; the non-hail sample =-23.58℃, =11.67℃.
[0065] Long-tail distribution correction, for optn, lwp, tbb: remove outliers using 3σ rule: for each index, calculate the mean ± 3 times the standard deviation of hail / non-hail samples, remove data points outside the range. For example, for liquid water path lwp non-hail samples, remove data points with lwp > 201.53 + 3 x 188.45 = 766.88, recalculate the corrected = 190.21, = 175.32.
[0066] 4. Dynamic confidence interval division and index optimization
[0067] Integral calculation: numerically integrate the probability density function of each index's hail / non-hail samples to solve the one-sided 95% confidence interval:
[0068] Hail samples: calculate the lower bound b such that the interval [b,+∞) contains 95% of the hail data;
[0069] Non-hail samples: calculate the upper bound a such that the interval (-∞, a] contains 95% of the non-hail data.
[0070] Physical constraint optimization: the maximum supercooled layer thickness of non-hail points in actual observation is 8 kilometers, so the upper bound a of non-hail samples is forcibly limited to 8 kilometers, and the corrected: Low hail probability interval: [0, 4.20) kilometers, containing 5% hail data; Medium hail probability interval: [4.20, 8] kilometers, data overlap area; High hail probability interval: (8, +∞) kilometers, containing only hail data. Final three-level probability interval.
[0071] 5. Multi-dimensional joint identification and weight distribution
[0072] Identification rules:
[0073] Single-index trigger rule: if any index falls into the high probability interval, it is directly determined as high probability hail.
[0074] For example, if a sample blackbody brightness temperature is -48℃, falling into the high probability interval (-∞, -37.70)℃, it is directly determined as high probability hail.
[0075] Multi-index coordination rule: if at least 3 indexes fall into the medium probability interval and no index falls into the low probability interval, it is determined as medium probability hail; for example, if a sample satisfies cloud top temperature = -30℃, supercooled layer thickness = 6 kilometers, liquid water path = 300, and no index is in the low probability interval, it is determined as medium probability hail.
[0076] Weight distribution mechanism: for core indicators, cloud top temperature, blackbody brightness temperature, weight coefficient ≥ 0.2, such as 0.25, auxiliary indicators optical thickness, liquid water path, weight coefficient ≤ 0.1, such as 0.08, the comprehensive hail probability value is calculated by weighted summation: , wherein ∈{0,1} , whether it falls into the corresponding interval
[0077] If cloud top temperature weight 0.25, blackbody brightness temperature weight 0.25 falls into the medium probability interval, optical thickness weight 0.08 falls into the medium probability interval, and liquid water path weight 0.08 falls into the low probability interval, then P=0.25+0.25+0.08=0.58, the low probability indicator is not included in the summation, and if the threshold value P≥0.5 is set, it is determined that the hail is in the medium probability.
[0078] 6. Model verification and optimization, leave-one-out cross-validation LOOCV: 368 samples are left one by one as the test set, and the remaining 367 are used as the training set, and the following is calculated:
[0079] Hail sample recognition accuracy:
[0080] Non-hail sample recognition accuracy:
[0081] Overall recognition accuracy:
[0082] The six indicators ztop, ttop, hsc, optn, lwp and tbb are used as the input of the SVM model, and the Linear-SVM, RBF-SVM and Sigmoid-SVM models are trained respectively, and the hyperparameters are optimized through grid search, such as RBF kernel C=10, γ=0.1. The SMOTE algorithm is used for oversampling the hail samples, so that hail:non-hail=1:1, and the recognition ability of the model for the minority class is improved.
[0083] The RBF-SVM model with the highest overall accuracy in cross-validation is selected as the final classifier, such as 87.5%. Real-time retrieval product data of FY-2G satellite is received, and 7 indicators of the target area cloud system are extracted.
[0084] Using LOOCV to test samples one by one, calculate the confusion matrix, accuracy, false positive rate and other indicators to ensure the generalization ability of the model. The RBF-SVM model is integrated into the weather warning system, real-time access to FY-2G satellite retrieval data, and automatic output of hail probability level (high / medium / low).
[0085] Combined with geographic information system GIS, the hail risk area is visualized, which provides accurate guidance for artificial hail prevention and agricultural disaster prevention.
[0086] The above are only embodiments of the present application, and common knowledge of specific structures and characteristics in the scheme is not described in detail here. Those skilled in the art know all the common technical knowledge in the field of the present application before the application date or the priority date, can know all the prior art in the field, and have the ability to apply conventional experimental means before that date. Those skilled in the art can perfect and implement the present scheme based on their own abilities under the guidance of the present application. Some typical known structures or known methods should not be an obstacle for those skilled in the art to implement the present application. It should be noted that for those skilled in the art, without departing from the structure of the present application, a number of modifications and improvements can be made, which should also be considered as the protection scope of the present application, and these will not affect the implementation effect and practicality of the patent. The protection scope claimed in the present application should be subject to the content of its claims, and the specific implementation mode and the like in the specification can be used to explain the content of the claims.
Claims
1. A method for quantitative identification of hail based on Gaussian distribution in FY-2G satellite data, characterized in that, Includes the following steps: S1. Construct a spatiotemporal matching dataset: Obtain inversion product data of cloud top height, cloud top temperature, supercooled layer thickness, optical thickness, effective particle radius, liquid water path, and blackbody brightness temperature from the FY-2G satellite. Centered on the hailfall point observed on the ground, extract satellite inversion data within a period before and after the hailfall time as hailfall samples. Select an equal number of satellite inversion data from non-hailfall periods and spatially non-hailfall areas as non-hailfall samples to form a dataset containing spatiotemporal matching features. S2. Gaussian Distribution Fit Test: Quantile-quantile plots are used to test the distribution characteristics of the inverted product data. The degree of overlap between the quantile points and the standard normal distribution line is used as the criterion to determine that cloud top height, cloud top temperature, and supercooled layer thickness follow a Gaussian distribution, while optical thickness, liquid water path, and blackbody brightness temperature approximately follow a Gaussian distribution. S3. Two-Sample Probability Density Modeling: For inverted product data that follows or approximately follows a Gaussian distribution, the mathematical expectations of hail-affected samples and non-hail-affected samples are calculated separately. and standard deviation Construct the two-sample Gaussian distribution probability density function: ; S4. Dynamic Confidence Interval Division: One-sided confidence intervals are established for hail-affected and non-hail-affected samples using numerical integration. Confidence interval: Calculate the infimum b for the hail sample to obtain the high hail probability interval (b,+∞); The supremacy 'a' is calculated for non-hail samples, yielding the low hail probability interval (-∞, a). Based on the definition of the hail probability interval in the overlapping part of the interval This forms a quantitative identification index containing a three-level probability gradient; S5. Multi-dimensional Joint Identification: The inverted product data of the sample to be identified is matched with the corresponding three-level probability intervals in step S4, and the probability of hail occurrence is determined by the following rules: Single indicator trigger rule: If any indicator falls into the high probability range of hail, it is directly judged as a high probability of hail. Multi-indicator coordination rule: If at least 3 indicators fall into the medium probability range of hail and no indicator falls into the low probability range, it is judged as a medium probability hail. Background exclusion rule: If all indicators fall into the low probability range, it is determined to be non-hailfall.
2. The method for quantitative identification of hail based on Gaussian distribution in FY-2G satellite data according to claim 1, characterized in that, The method for constructing the spatiotemporal matching dataset in step S1 is as follows: taking 184 hail points on 30 hail days as the center, extracting satellite inversion data at the corresponding time according to latitude and longitude coordinates, and selecting non-hail point data through spatial random sampling to ensure that the two types of samples are aligned in terms of temporal and spatial resolution.
3. The method for quantitative identification of hail based on Gaussian distribution in FY-2G satellite data according to claim 1, characterized in that, The criteria for determining the Gaussian distribution fit test in step S2 are as follows: when the Pearson correlation coefficient between the quantile point of the inverted product data and the standard normal distribution line is ≥0.9, it is determined to follow a Gaussian distribution; when the correlation coefficient is between 0.8 and 0.9, it is determined to approximately follow a Gaussian distribution.
4. The method for quantitative identification of hail based on Gaussian distribution in FY-2G satellite data according to claim 1, characterized in that, The two-sample probability density modeling in step S3 further includes correction of the long-tail distribution of optical thickness, liquid water path, and blackbody brightness temperature: by introducing a truncation threshold, outliers are removed and the mathematical expectation and standard deviation are recalculated to improve the model's fitting accuracy for non-ideal distribution data.
5. The method for quantitative identification of hail based on Gaussian distribution in FY-2G satellite data according to claim 1, characterized in that, In step S4, the dynamic division of the confidence interval is solved by numerical integration to obtain the probability density function, satisfying the following: the one-sided confidence interval of the hail sample is [b, +∞), which satisfies... ; For non-hail samples, the one-sided confidence interval (-∞, a] satisfies ,in and These are the Gaussian probability density functions for hail and non-hail samples, respectively.
6. The method for quantitative identification of hail based on Gaussian distribution in FY-2G satellite data according to claim 1, characterized in that, In step S4, the identification index for supercooled layer thickness is optimized as follows: combined with the constraint that the maximum supercooled layer thickness at non-hail points in the actual observation data is 8 kilometers, the supremacy a of non-hail samples is forcibly limited to 8 kilometers, forming the corrected medium hail probability interval [4.20,8] and high hail probability interval (8,+∞).
7. The method for quantitative identification of hail based on Gaussian distribution in FY-2G satellite data according to claim 1, characterized in that, The multi-dimensional joint identification in step S5 also includes a weight allocation mechanism: weights are assigned to the cloud top temperature and blackbody brightness temperature indicators with a weight coefficient of 0.2-0.3, and weights are assigned to the optical thickness and liquid water path auxiliary indicators with a weight coefficient of 0.05-0.1, using a weighted summation formula. Calculate the overall probability value of hail, where This is a binary variable indicating whether the indicator falls within the corresponding probability interval; it is set to 1 if it falls within the interval, and 0 otherwise. These are the weighting coefficients for the corresponding indicators.
8. The method for quantitative identification of hail based on Gaussian distribution in FY-2G satellite data according to claim 1, characterized in that, The S5 also includes a model validation step: using leave-one-out cross-validation (LOOCV) to test the effectiveness of the quantitative identification indicators, and calculating the identification accuracy of hail samples, the identification accuracy of non-hail samples, and the overall identification accuracy.
9. The method for quantitative identification of hail based on Gaussian distribution in FY-2G satellite data according to claim 1, characterized in that, The S5 further integrates a Support Vector Machine (SVM) model: the inverted product data is used as input features, and classification models are trained using Linear, RBF, and Sigmoid kernel functions respectively. Hyperparameters are optimized through grid search, and finally the RBF-SVM model is selected as the final model.
10. The method for quantitative identification of hail based on Gaussian distribution in FY-2G satellite data according to claim 9, characterized in that, The training process of the SVM model introduces sample equalization: the SMOTE algorithm is used to generate synthetic samples for the minority class of hail samples, so that the ratio of hail samples to non-hail samples reaches 1:1.
Citation Information
Patent Citations
Hail recognition method, system and equipment based on multi-source data and storage medium
CN117055051A
Ground hail shooting identification and early warning method based on dual linear polarization radar
CN113933845A
Hail early warning method and system based on lightning jump and support vector machine
CN120233320A