Gaussian distribution-based hail quantitative identification method for FY-2G satellite data
By constructing a spatiotemporal matching data set and Gaussian distribution modeling, combining quantile-quantile graphs and multi-index collaborative rules, quantitative hail recognition of FY-2G satellite data is realized, solving the problem of insufficient regional adaptability and quantification of hail recognition methods, improving the recognition accuracy and adaptability, and supporting meteorological warning.
Patent Information
- Application Number
- CN202510972514.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-07-15
AI Technical Summary
The existing hail recognition methods have poor adaptability and insufficient quantification in different regions, which cannot meet the needs of refined early warnings, and rely on a single data source or machine learning algorithms to have high computing resources, making it difficult to adapt to the rapid early warning of small and medium-sized disasters.
The quantitative identification method of FY-2G satellite data based on Gaussian distribution is used to construct a spatio-temporal matching data set, and the data distribution adaptability is verified using quantile-quantile graphs, a two-sample Gaussian distribution probability density function is established, a three-level probability interval is divided, and a single-index trigger and multi-index collaborative rules are identified.
The hail recognition accuracy rate is improved by 20% to 30%, the regional adaptability and quantitative ability of the method are enhanced, and more accurate hail occurrence probability estimates are provided, supporting the disaster prevention and mitigation work of the meteorological department.
Smart Images

Figure CN120470271A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of meteorological disaster early warning, and relates to a method for quantitatively identifying hail using FY-2G satellite data based on Gaussian distribution. Background Art
[0002] Hail is a severe convective weather disaster that frequently occurs in my country's low-latitude plateau regions, such as Guizhou. It is characterized by strong localization, a short lifespan, and difficulty in early identification. This disaster often causes devastating damage to specialty agricultural products such as tobacco and tea, as well as protected agriculture. A single disaster can result in economic losses exceeding 100 million yuan. Satellite observations, with their wide coverage and real-time availability, have become a core technical tool for short-term hail warning. The FY-2G satellite is the primary satellite for my country's weather modification operations. The seven inversion products it provides, including cloud top height (ztop), cloud top temperature (ttop), and supercooled layer thickness (hsc), effectively reflect the macro- and microphysical characteristics of hail clouds. However, existing hail identification methods generally suffer from insufficient regional adaptability and a lack of quantitative indicators. This results in significant fluctuations in identification accuracy under varying meteorological and topographical conditions, making them inadequate for precise early warning.
[0003] Hail identification models based on stability indices, such as the CAPE and TT indices, and ground-based radar data are highly dependent on the geographic and meteorological conditions of a specific region, significantly reducing their performance when applied across regions. For example, a stability index developed by European researchers for local meteorological conditions has been found to be ineffective in convective weather forecasts in tropical and North American regions. A logistic regression model based on radar and sounding data in northern Spain, while achieving an accuracy rate of 0.87 for hail identification, has a false alarm rate as high as 18%. In areas of my country with complex terrain, traditional models often have a false alarm rate exceeding 20% due to terrain obstruction and radar coverage blind spots.
[0004] Most existing methods use statistical models such as decision trees and Bayesian models to make comprehensive multi-factor judgments (e.g., literature). However, they lack single-indicator quantitative judgment criteria based on data distribution characteristics, and threshold setting relies on subjective experience, making it difficult to probabilistically classify hail of different intensities. Traditional methods mainly rely on a single data source, such as radar or ground observations, and fail to fully explore the multidimensional characteristics of FY-2G satellite inversion product data. Although some studies have attempted to introduce machine learning algorithms such as support vector machines (SVMs), the prior analysis of data distribution is insufficient, the model interpretability is poor, and the recognition accuracy is limited in the scenario of single-source satellite data.
[0005] For example, the patent number CN117055051A is titled "Hail Identification Method, System, Equipment and Storage Medium Based on Multi-Source Data". The solution proposes to integrate data such as three-dimensional radar, FY4B satellite, and ERA5 environmental field, realize identification through the FEMU-Net network, and use neural networks to automatically extract deep features of multi-source data; however, the solution relies on auxiliary data such as radar and large-scale labeled samples, and has high computing resource requirements; no quantitative indicators are designed for single-source satellite data, making it difficult to adapt to the rapid early warning needs of small and medium-scale disasters.
[0006] This paper proposes a framework for single-indicator quantitative modeling and multi-indicator collaborative identification based on Gaussian distribution. Its core technology relies on verifying the Gaussian distribution adaptability of multiple FY-2G satellite inversion product data, excluding the effective particle radius, through quantile-quantile plots (QQ plots). Using a two-sample probability density function and 95% confidence intervals, a quantitative identification indicator with a three-level probability gradient (low-medium-high) is established. Comprehensive judgment is achieved through single-indicator triggering and multi-indicator collaborative rules. This framework offers breakthrough advantages such as quantification and interpretability, regional adaptability optimization, lightweightness, and real-time performance, effectively solving practical production problems. Summary of the Invention
[0007] The present invention provides a quantitative hail identification method based on FY-2G satellite data using Gaussian distribution. Addressing the problems of poor adaptability and insufficient quantification of existing hail identification methods in different regions, the present invention achieves quantitative determination of hail probability using single and multiple indicators by constructing a spatiotemporal matching dataset, performing Gaussian distribution modeling, and performing confidence interval partitioning.
[0008] In order to solve the above problems, the technical solution adopted by the invention is: The method for quantitatively identifying hail using FY-2G satellite data based on Gaussian distribution includes the following steps: S1. Constructing a spatiotemporal matching dataset: obtaining the inversion product data of cloud top height, cloud top temperature, supercooling layer thickness, optical depth, effective particle radius, liquid water path, and blackbody brightness temperature from the FY-2G satellite; taking the ground-observed hail point as the center, extracting the satellite inversion data within a period of time before and after the hailfall moment as hail samples; selecting the same number of satellite inversion data in the non-hailfall period and spatially non-hailfall area as non-hailfall samples; Form a data set containing spatiotemporal matching features; S2. Gaussian distribution adaptability test: Use the quantile-quantile plot to test the distribution characteristics of the inversion product data, and use the degree of coincidence between the quantile point and the standard normal distribution line as the basis for judgment to determine that the cloud top height, cloud top temperature, and supercooled layer thickness follow the Gaussian distribution, and the optical thickness, liquid water path, and blackbody brightness temperature approximately follow the Gaussian distribution; S3. Two-sample probability density modeling: For the inversion product data that obey or approximately obey the Gaussian distribution, calculate the mathematical expectation of hail samples and non-hail samples respectively. and standard deviation , construct the two-sample Gaussian distribution probability density function:
[0009] ; S4. Dynamic division of confidence intervals: establish one-sided confidence intervals for hail samples and non-hail samples respectively: Calculate the infimum b of the hail samples and obtain the high hail probability interval [b, +∞); Calculate the supremum a for the non-hail samples and obtain the low hail probability interval (-∞,a]); Based on the hail probability interval [a, b] defined in the overlapping intervals, a quantitative identification index containing three-level probability gradients is formed. S5. Multi-dimensional joint identification: For the inversion product data of the sample to be identified, the corresponding three-level probability intervals in step S4 are matched respectively, and the probability of hail occurrence is determined according to the following rules: Single indicator trigger rule: If any indicator falls into the high hail probability range, it is directly determined as a high probability hail; Multi-indicator coordination rule: If at least three indicators fall into the medium hail probability interval and no indicator falls into the low probability interval, it is judged as a medium probability hail.
[0010] The principles and advantages of this solution are: By constructing a spatiotemporal matching dataset, we collected various inversion product data from the FY-2G satellite, including cloud top height and cloud top temperature. With ground-based observations of hail points as the core, we selected hail samples and non-hail samples to ensure that the data had spatiotemporal matching characteristics, laying the foundation for subsequent analysis. Secondly, we used a quantile-quantile plot to test the Gaussian distribution adaptability of the inversion product data. By comparing the degree of overlap between the quantile points and the standard normal distribution line, we screened out data indicators that obeyed or approximately obeyed the Gaussian distribution, providing a reliable basis for probability density modeling. Next, for the screened indicators, we calculated the mathematical expectation and standard deviation of the hail samples and non-hail samples, respectively, constructed a two-sample Gaussian distribution probability density function, and quantified the data distribution characteristics of the two types of samples. Then, we established one-sided confidence intervals for the hail samples and non-hail samples, divided the probability intervals of hail into three levels: high, medium, and low, and formed a quantitative identification indicator. Finally, through single-indicator triggering rules and multi-indicator coordination rules, multi-dimensional joint identification of hail occurrence probability is achieved. When any indicator falls into the high hail probability interval or at least three indicators fall into the medium hail probability interval and no indicator falls into the low probability interval, the probability of hail occurrence can be determined.
[0011] Compared with existing technologies, existing methods often rely on single indicators or simple thresholds for identification accuracy, which are susceptible to environmental factors and can lead to misidentification and missed detections. For example, judging hail based solely on cloud top temperature may misidentify cold but hail-free clouds as hail. This solution, however, utilizes multi-dimensional data fusion and Gaussian distribution modeling to construct a three-level probability gradient identification indicator. By leveraging single-indicator triggering and multi-indicator collaborative rules, it can more comprehensively capture hail formation characteristics and significantly improve identification accuracy. Field data validation shows that under complex weather conditions, this solution improves hail identification accuracy by approximately 20% to 30% compared to traditional methods. Regarding reliability, existing technologies lack in-depth analysis of data distribution characteristics and struggle to cope with data fluctuations. This solution, through Gaussian distribution adaptability testing, scientifically screens data indicators, and establishes a probability density model. This enhances the recognition method's adaptability to diverse data scenarios and reduces the risk of misidentification due to data fluctuations. Regarding quantification, existing technologies often rely on qualitative or semi-quantitative methods, failing to provide a precise probability of hail occurrence. This plan achieves a quantitative expression of the probability of hail occurrence by constructing a three-level probability interval, providing more accurate and intuitive data support for meteorological warnings and decision-making, enabling meteorological departments to arrange disaster prevention and mitigation work more reasonably and effectively reduce the losses caused by hail disasters.
[0012] Furthermore, the method for constructing the spatiotemporal matching dataset in step S1 is as follows: With 184 hail points from 30 hail days as the center, satellite inversion data at the corresponding time is extracted according to the longitude and latitude coordinates, and non-hail point data is selected through spatial random sampling to ensure that the two types of samples are aligned in terms of temporal and spatial resolution. Extracting data centered on hail points can accurately capture the core characteristics of hail clouds, such as key parameters such as high cloud top height and low cloud top temperature, so that the hail samples directly reflect the actual environment when hail occurs. Spatially random sampling of non-hail point data can avoid artificial selection bias, ensure that the non-hail samples cover a wide range of non-hail cloud types, such as stratiform clouds and ordinary cumulus clouds, and truly reflect the data distribution of non-hail scenes. In terms of spatiotemporal consistency, strict alignment of temporal and spatial resolution can eliminate false correlations caused by spatiotemporal misalignment, for example, avoiding forced matching of non-hail data with hail points at different times or in different regions, ensuring that the two types of samples are in the same spatiotemporal reference, and making subsequent statistical tests and probability modeling based on Gaussian distribution more scientific. Judging from the model training effect, the data set formed by this construction method has balanced temporal and spatial characteristics, which can effectively improve the model's ability to distinguish between hail and non-hail scenes. For example, when comparing the probability density distribution of indicators such as cloud top temperature, it can clearly present the significant differences between hail and non-hail samples, avoiding model misjudgment due to temporal and spatial mixing, and laying a reliable data foundation for the establishment of subsequent quantitative identification indicators.
[0013] Furthermore, the judgment criteria for the Gaussian distribution adaptability test in step S2 are: when the Pearson correlation coefficient between the quantile point of the inversion product data and the standard normal distribution line is ≥0.9, it is judged to obey the Gaussian distribution; when the correlation coefficient is between 0.8-0.9, it is judged to approximately obey the Gaussian distribution. The Pearson correlation coefficient quantifies the degree of linear correlation between the quantile point and the standard normal distribution line through a mathematical formula, avoiding subjective judgment errors, such as avoiding ambiguous conclusions caused by visual overlap of the QQ graph alone, making the test results more credible. In terms of operability, the standard provides a clear numerical threshold to facilitate researchers to quickly determine the data distribution type. For example, when processing inversion product data such as cloud top height and cloud top temperature, they can be directly classified by calculating the correlation coefficient, thereby improving test efficiency. A correlation coefficient ≥0.9 indicates that the data points have a very high degree of linear fit with the standard normal distribution. For example, the distribution characteristics of cloud top height, cloud top temperature, and supercooled layer thickness can be directly modeled using the probability density function and confidence interval of the Gaussian distribution. In the correlation coefficient range of 0.8-0.9, such as optical thickness and liquid water path, although there are certain deviations, the overall distribution trend is close to the Gaussian distribution. Through approximate processing, the statistical characteristics of the Gaussian distribution can still be used to establish quantitative identification indicators, which broadens the applicability of the data while ensuring the accuracy of the model, avoids the information loss caused by the strict exclusion of approximate distribution data, and provides a reliable basis for subsequent two-sample probability density modeling and confidence interval division.
[0014] Furthermore, the dual-sample probability density modeling in step S3 further includes correcting the long-tail distributions of optical thickness, liquid water path, and blackbody brightness temperature. By introducing a cutoff threshold, outliers are removed and the mathematical expectation and standard deviation are recalculated to improve the model's fitting accuracy for non-ideally distributed data. Outliers in long-tail distributions, such as sudden surges in optical thickness or abnormally low liquid water path values, may be caused by factors such as satellite observation errors and cloud boundary interference, and are not characteristic data of actual hail or non-hail scenarios. By setting a cutoff threshold, such as removing data points exceeding ±3 standard deviations from the mean based on the 3σ principle, this type of noise can be effectively filtered out, preventing outliers from distorting overall distribution parameters such as the mathematical expectation and standard deviation. For example, if individual non-hail samples in the liquid water path data contain extremely high values, they may inflate the mathematical expectation of the non-hail dataset, causing subsequent confidence intervals to deviate from the true distribution. However, removing these outliers can make the probability density curve of the non-hail samples more closely resemble the liquid water content characteristics of actual clouds. Judging from the model fitting results, the corrected mathematical expectation and standard deviation can more accurately reflect the central tendency and degree of dispersion of the data body. Taking the blackbody brightness temperature as an example, after eliminating the abnormally high temperature values caused by instrument failure in the hail samples, the mathematical expectation of the hail dataset is closer to the actual low temperature characteristics. The peak position and distribution amplitude of its probability density curve can better reflect the temperature characteristics of the strong convective cloud top. This enhances the recognition of the distribution difference between hail and non-hail samples in dual-sample modeling, improves the rationality of subsequent confidence interval division and the reliability of identification indicators. This correction mechanism optimizes the applicability of the Gaussian distribution model through outlier control while preserving the distribution characteristics of the data body. In particular, for parameters that approximately follow a Gaussian distribution, it effectively balances data integrity and model accuracy.
[0015] Furthermore, the confidence interval dynamic division in step S4 is solved by numerical integration method to solve the probability density function, which satisfies: the one-sided confidence interval of the hail sample [b, +∞) satisfies ; One-sided confidence interval for non-hail samples (-∞,a], satisfying ,in and The Gaussian probability density functions of hail and non-hail samples are calculated directly through mathematical integration to ensure that the confidence interval [b, +∞) of hail samples strictly covers 95% of hail data. For example, the cloud top height of hail samples is determined by integration to be b=7.20, so that this interval contains 95% of the cloud top height values of hail points. The confidence interval of non-hail samples is (-∞, ] accurately contains 95% of non-hail data, such as the cloud top temperature non-hail sample calculated to be a=-42.72, avoiding the arbitrariness of subjective threshold setting. In terms of dynamic adaptability, for the probability density characteristics of different inversion product data, such as supercooling layer thickness, liquid water path, etc., the numerical integration method can flexibly adapt to the data distribution differences. For example, the maximum value of the supercooling layer thickness non-hail sample is limited to 8 kilometers by actual observation, and it is adjusted by integration combined with data truncation. , so that the confidence interval is more in line with the actual distribution. In terms of statistical reliability, the interval division based on 95% confidence level meets the statistical standards, and the high hail probability interval ( ,+∞) contains only 5% of non-hail data, which can be used as a strong basis for judging the high probability of hail triggered by a single indicator. For example, when the blackbody brightness temperature is ≤−37.70℃, it can be directly judged that there is a high probability of hail. The medium hail probability interval [ ,b] Covering the data overlapping area, through the multi-indicator collaborative rule, at least three indicators fall into this interval and there is no low-probability indicator for comprehensive judgment, which effectively reduces the risk of misjudgment of a single indicator due to data overlap.
[0016] Furthermore, the identification index of supercooled layer thickness in step S4 is optimized as follows: combining the constraint condition that the maximum supercooled layer thickness at the non-hail point in the actual observation data is 8 kilometers, the supremum a of the non-hail sample is forcibly limited to 8 kilometers, forming the revised medium hail probability interval [4.20,8] and high hail probability interval (8,+∞). The maximum supercooled layer thickness at the non-hail point in the actual observation is 8 kilometers, indicating that when this index exceeds 8 kilometers, the cloud system has broken through the physical boundary of the non-hail scene. By forcibly limiting =8 km, which can avoid misjudgment caused by the separation of pure mathematical integration results from actual observations. For example, it can exclude the false data scenario of "the thickness of the supercooled layer at non-hail points exceeds 8 km", so that the confidence interval of non-hail samples (-∞,8] strictly fits the observation facts, and enhances the credibility of the indicator. The corrected high hail probability interval (8,+∞) directly corresponds to the physical feature of "the thickness of the supercooled layer exceeds the possible range of non-hail cloud systems". For example, when strong convective clouds develop vigorously, the thickness of the supercooled layer increases significantly. When it exceeds 8 kilometers, it can be clearly determined as a high probability hail signal, reducing the original data. The interval (8, 14, 18) in the academic solution addresses the indicator window problem caused by the lack of actual non-hail data support. The medium hail probability interval [4, 20, 8] covers the overlapping area of the supercooled layer thickness, including some hail samples (95% of the hail data not covered by the high probability interval), as well as the extreme values of non-hail samples. False alarms can be further filtered out through multi-indicator collaborative rules. For example, when other indicators, such as cloud top temperature and liquid water path, fall into the medium- and high-probability intervals at the same time, a comprehensive judgment is made that hail has a medium probability, improving the rigor of the recognition logic.
[0017] Furthermore, the multi-dimensional joint identification in step S5 also includes a weight distribution mechanism: weights are assigned to the cloud top temperature and blackbody brightness temperature indicators, with a weight coefficient ≥ 0.2; weights are assigned to the optical thickness and liquid-water path auxiliary indicators, with a weight coefficient ≤ 0.1; and the comprehensive hail probability value is calculated through weighted summation.
[0018] Furthermore, the method also includes a model validation step: using leave-one-out cross-validation (LOOCV) to test the effectiveness of quantitative identification indicators, calculate the hail sample identification accuracy, non-hail sample identification accuracy, and overall identification accuracy. The leave-one-out cross-validation method separates the data set into training and test sets one by one, retaining one sample as the test set each time and the remaining samples as the training set, ensuring that each sample participates in model training and validation, avoiding evaluation bias caused by random sampling. For example, in 368 groups of samples, each hail point and non-hail point were separately reserved for testing, so that the hail sample identification accuracy rate and non-hail sample identification accuracy rate, such as 91.85%, can truly reflect the model's ability to distinguish between the two types of data, and the overall identification accuracy rate, such as 87.50%, comprehensively reflects the overall effectiveness of the indicators. From the perspective of reliability verification, this method can quantify the stability of the model by calculating the mean and variance through multiple iterations. For example, when validating an RBF kernel-based SVM model, the leave-one-out method showed a low standard deviation in the accuracy of identifying non-hail samples, indicating the model's robust discrimination of non-hail scenarios. Furthermore, the fluctuation range of the accuracy of hail samples was relatively small, demonstrating the consistency of the indicator in capturing hail characteristics. Furthermore, by comparing cross-validation results for different kernel function models, such as L-SVM, RBF-SVM, and S-SVM, the optimal model can be identified. For example, the RBF-SVM achieved the highest overall accuracy, thus avoiding model selection bias caused by a single training-testing split.
[0019] Furthermore, the method further integrates a support vector machine (SVM) model: using inversion product data as input features, the classification model is trained using Linear, RBF, and Sigmoid kernel functions, respectively. Grid search is used to optimize hyperparameters, ultimately selecting the RBF-SVM model as the preferred model. The SVM model uses kernel functions to map low-dimensional input features, such as cloud top height and cloud top temperature, to a high-dimensional space, effectively addressing the classification challenge of nonlinear features in hail identification. For example, indicators such as optical thickness and liquid water path have a high degree of distribution overlap in low-dimensional space, making them difficult to distinguish using linear boundaries. The RBF kernel, however, uses Gaussian radial basis functions to calculate the similarity between samples, constructing a nonlinear decision boundary in high-dimensional space and accurately capturing the implicit feature differences between hail and non-hail samples.
[0020] Furthermore, sample balancing is introduced during the training of the SVM model: the SMOTE algorithm is used to generate synthetic samples for minority hail samples, ensuring a 1:1 ratio between hail samples and non-hail samples. Although the number of hail samples and non-hail samples in the original dataset is balanced, hail events in actual meteorological scenarios are typically low-probability events, and real data may be naturally imbalanced, such as non-hail samples far outnumbering hail samples. The SMOTE algorithm generates synthetic hail samples through interpolation, such as generating virtual samples based on k-nearest neighbor features. This can simulate the feature space distribution of real hail clouds and prevent the model from biased towards learning majority class features due to the sparse minority class samples, such as overfitting the low cloud top height and high cloud top temperature of non-hail samples. For example, in the feature space of indicators such as optical thickness, hail samples may form several isolated clusters. The samples generated by SMOTE can fill the gaps between clusters, enhancing the model's overall understanding of the hail feature distribution. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 Flowchart of the present invention. DETAILED DESCRIPTION
[0022] Example 1, a method for quantitatively identifying hail using FY-2G satellite data based on Gaussian distribution, comprising the following steps: S1. constructing a spatiotemporal matching data set: obtaining inversion product data of cloud top height, cloud top temperature, supercooling layer thickness, optical depth, effective particle radius, liquid water path, and blackbody brightness temperature from the FY-2G satellite; taking the hail point observed on the ground as the center, extracting satellite inversion data within a period of time before and after the hailfall moment as hailfall samples; selecting satellite inversion data of the same number, in non-hailfall periods, and in spatially non-hailfall areas as non-hailfall samples; This forms a data set containing spatiotemporal matching features; S2. Gaussian distribution adaptability test: the distribution characteristics of the inversion product data are tested using the quantile-quantile diagram, and the degree of coincidence between the quantile point and the standard normal distribution line is used as the basis for judgment to determine that the cloud top height, cloud top temperature, and supercooled layer thickness obey the Gaussian distribution, and the optical thickness, liquid water path, and blackbody brightness temperature approximately obey the Gaussian distribution; S3. Two-sample probability density modeling: for the inversion product data that obey or approximately obey the Gaussian distribution, the mathematical expectations of hail samples and non-hail samples are calculated respectively. and standard deviation , construct the two-sample Gaussian distribution probability density function:
[0023] , S4. Dynamic division of confidence intervals: establish one-sided confidence intervals for hail samples and non-hail samples respectively: Calculate the infimum b of the hail samples and obtain the high hail probability interval [b, +∞); Calculate the supremum a for the non-hail samples and obtain the low hail probability interval (-∞,a]); Based on the hail probability interval [a, b] defined in the overlapping intervals, a quantitative identification indicator consisting of three levels of probability gradients is formed. The numerical integration method uses Simpson's rule to discretize the probability density function. The integration thresholds a and b are solved through iterative approximation to ensure that the interval covers 95% of the sample data. S5. Multi-dimensional Joint Identification: For the inversion product data of the sample to be identified, the corresponding three-level probability intervals in step S4 are matched, and the probability of hail occurrence is determined according to the following rules: Single indicator trigger rule: If any indicator falls into the high hail probability range, it is directly determined as a high probability hail; Multi-indicator coordination rule: If at least three indicators fall into the medium hail probability interval and no indicator falls into the low probability interval, it is judged as a medium probability hail.
[0024] By constructing a spatiotemporal matching dataset, we collected various inversion product data from the FY-2G satellite, including cloud top height and cloud top temperature. With ground-based observations of hail points as the core, we selected hail samples and non-hail samples to ensure that the data had spatiotemporal matching characteristics, laying the foundation for subsequent analysis. Secondly, we used a quantile-quantile plot to test the Gaussian distribution adaptability of the inversion product data. By comparing the degree of overlap between the quantile points and the standard normal distribution line, we screened out data indicators that obeyed or approximately obeyed the Gaussian distribution, providing a reliable basis for probability density modeling. Next, for the screened indicators, we calculated the mathematical expectation and standard deviation of the hail samples and non-hail samples, respectively, constructed a two-sample Gaussian distribution probability density function, and quantified the data distribution characteristics of the two types of samples. Then, we established one-sided confidence intervals for the hail samples and non-hail samples, divided the probability intervals of hail into three levels: high, medium, and low, and formed a quantitative identification indicator. Finally, through single-indicator triggering rules and multi-indicator coordination rules, multi-dimensional joint identification of hail occurrence probability is achieved. When any indicator falls into the high hail probability interval or at least three indicators fall into the medium hail probability interval and no indicator falls into the low probability interval, the probability of hail occurrence can be determined.
[0025] In terms of recognition accuracy, existing methods often rely on a single indicator or simple threshold judgment, which is susceptible to environmental factors and can lead to misjudgments and missed detections. For example, judging hail based solely on cloud top temperature may misidentify cold but hail-free clouds as hail. However, this solution utilizes multi-dimensional data fusion and Gaussian distribution modeling to construct a three-level probability gradient recognition indicator. By leveraging single-indicator triggering and multi-indicator collaborative rules, it can more comprehensively capture hail formation characteristics and significantly improve recognition accuracy. Field data validation shows that under complex weather conditions, this solution's hail recognition accuracy is approximately 20% to 30% higher than traditional methods. Regarding reliability, existing technologies lack in-depth analysis of data distribution characteristics and struggle to cope with data fluctuations. This solution uses Gaussian distribution adaptability testing, scientifically screens data indicators, and establishes a probability density model. This enhances the recognition method's adaptability to diverse data scenarios and reduces the risk of misjudgment due to data fluctuations. In terms of quantification, existing technologies often rely on qualitative or semi-quantitative judgments, failing to provide a precise probability of hail occurrence. This plan achieves a quantitative expression of the probability of hail occurrence by constructing a three-level probability interval, providing more accurate and intuitive data support for meteorological warnings and decision-making, enabling meteorological departments to arrange disaster prevention and mitigation work more reasonably and effectively reduce the losses caused by hail disasters.
[0026] The method for constructing the spatiotemporal matching dataset in step S1 is as follows: taking the 184 hail points of 30 hail days as the center, extracting the corresponding satellite inversion data according to the latitude and longitude coordinates, and selecting non-hail point data through spatial random sampling to ensure that the two types of samples are aligned in terms of temporal and spatial resolution. Extracting data centered on the hail points can accurately capture the core characteristics of the hail cloud system, such as key parameters such as high cloud top height and low cloud top temperature, so that the hail samples directly reflect the actual environment when the hail occurs; spatial random sampling of non-hail point data can avoid human selection bias, ensure that the non-hail samples cover a wide range of non-hail cloud types, such as stratiform clouds and ordinary cumulus clouds, and truly reflect the data distribution of non-hail scenes. In terms of spatiotemporal consistency, strictly aligning the temporal resolution and spatial resolution can eliminate false correlations caused by spatiotemporal misalignment, for example, avoiding forced matching of non-hail data with hail points at different times or in different regions, ensuring that the two types of samples are in the same spatiotemporal reference, and making subsequent statistical tests and probability modeling based on Gaussian distribution more scientific. Judging from the model training effect, the data set formed by this construction method has balanced temporal and spatial characteristics, which can effectively improve the model's ability to distinguish between hail and non-hail scenes. For example, when comparing the probability density distribution of indicators such as cloud top temperature, it can clearly present the significant differences between hail and non-hail samples, avoiding model misjudgment due to temporal and spatial mixing, and laying a reliable data foundation for the establishment of subsequent quantitative identification indicators.
[0027] The judgment criteria for the "Gaussian distribution suitability test" in step S2 are: when the Pearson correlation coefficient between the quantile point of the inversion product data and the standard normal distribution line is ≥0.9, it is judged to obey the Gaussian distribution; when the correlation coefficient is between 0.8-0.9, it is judged to approximately obey the Gaussian distribution. The Pearson correlation coefficient quantifies the degree of linear correlation between the quantile point and the standard normal distribution line through a mathematical formula, avoiding subjective judgment errors, such as avoiding ambiguous conclusions caused by visual overlap of the QQ graph, making the test results more credible. In terms of operability, the standard provides a clear numerical threshold to facilitate researchers to quickly determine the data distribution type. For example, when processing inversion product data such as cloud top height and cloud top temperature, they can be directly classified by calculating the correlation coefficient, thereby improving test efficiency. A correlation coefficient ≥0.9 indicates that the data points have a very high degree of linear fit with the standard normal distribution. For example, the distribution characteristics of cloud top height, cloud top temperature, and supercooled layer thickness can be directly modeled using the probability density function and confidence interval of the Gaussian distribution. In the correlation coefficient range of 0.8-0.9, such as optical thickness and liquid water path, although there are certain deviations, the overall distribution trend is close to the Gaussian distribution. Through approximate processing, the statistical characteristics of the Gaussian distribution can still be used to establish quantitative identification indicators, which broadens the applicability of the data while ensuring the accuracy of the model, avoids the information loss caused by the strict exclusion of approximate distribution data, and provides a reliable basis for subsequent two-sample probability density modeling and confidence interval division.
[0028] The dual-sample probability density modeling in step S3 further includes correcting the long-tail distributions of optical thickness, liquid water path, and blackbody brightness temperature. By introducing a cutoff threshold, the mathematical expectation and standard deviation are recalculated after removing outliers, improving the model's fitting accuracy for non-ideally distributed data. Outliers in the long-tail distribution, such as sudden surges in optical thickness or abnormally low liquid water path values, may be caused by factors such as satellite observation errors and cloud boundary interference, and are not characteristic data of actual hail or non-hail scenarios. By setting a cutoff threshold, such as removing data points exceeding ±3 standard deviations from the mean based on the 3σ principle, this type of noise can be effectively filtered out, preventing outliers from distorting overall distribution parameters such as the mathematical expectation and standard deviation. For example, if individual non-hail samples in the liquid water path data contain extremely high values, they may inflate the mathematical expectation of the non-hail dataset, causing subsequent confidence intervals to deviate from the true distribution. However, removing these outliers can make the probability density curve of the non-hail samples more closely resemble the liquid water content characteristics of actual clouds. Judging from the model fitting results, the corrected mathematical expectation and standard deviation can more accurately reflect the central tendency and degree of dispersion of the data body. Taking the blackbody brightness temperature as an example, after eliminating the abnormally high temperature values caused by instrument failure in the hail samples, the mathematical expectation of the hail dataset is closer to the actual low temperature characteristics. The peak position and distribution amplitude of its probability density curve can better reflect the temperature characteristics of the strong convective cloud top. This enhances the recognition of the distribution difference between hail and non-hail samples in dual-sample modeling, improves the rationality of subsequent confidence interval division and the reliability of identification indicators. This correction mechanism optimizes the applicability of the Gaussian distribution model through outlier control while preserving the distribution characteristics of the data body. In particular, for parameters that approximately follow a Gaussian distribution, it effectively balances data integrity and model accuracy.
[0029] The integral calculation method for the dynamic division of the confidence interval in step S4 is: solving the probability density function by numerical integration method so that the confidence interval of the hail sample contains 95% of the hail data and the confidence interval of the non-hail sample contains 95% of the non-hail data. The specific formula is: , , the area under the probability density curve is directly calculated by mathematical integration to ensure that the confidence interval of hail samples [b, +∞) strictly covers 95% of hail data. For example, the cloud top height hail sample is determined by integration to be b=7.20, so that this interval contains 95% of the cloud top height values of hail points. The confidence interval of non-hail samples (-∞, 𝑎] accurately contains 95% of non-hail data, such as the cloud top temperature non-hail sample is calculated to be =-42.72, thus avoiding the arbitrariness of subjective threshold setting. In terms of dynamic adaptability, the numerical integration method can flexibly adapt to the data distribution differences for the probability density characteristics of different inversion product data, such as supercooling layer thickness, liquid water path, etc. For example, the maximum value of supercooling layer thickness for non-hail samples is limited to 8 kilometers by actual observation, and the integration is combined with data truncation adjustment. , so that the confidence interval is more in line with the actual distribution. In terms of statistical reliability, the interval division based on 95% confidence level meets the statistical standards, and the high hail probability interval ( ,+∞) contains only 5% of non-hail data, which can be used as a strong basis for judging the high probability of hail triggered by a single indicator. For example, when the blackbody brightness temperature is ≤−37.70℃, it can be directly judged that there is a high probability of hail. The medium hail probability interval [ ,b] Covering the data overlapping area, through the multi-indicator collaborative rule, at least three indicators fall into this interval and there is no low-probability indicator for comprehensive judgment, which effectively reduces the risk of misjudgment of a single indicator due to data overlap.
[0030] The identification index of supercooled layer thickness in step S4 is optimized as follows: combining the constraint condition that the maximum supercooled layer thickness at the non-hail point in the actual observation data is 8 kilometers, the supremum a of the non-hail sample is forcibly limited to 8 kilometers, forming the revised medium hail probability interval [4.20,8] and high hail probability interval (8,+∞). The maximum supercooled layer thickness at the non-hail point in the actual observation is 8 kilometers, indicating that when this index exceeds 8 kilometers, the cloud system has broken through the physical boundary of the non-hail scene. By forcibly limiting =8 km, which can avoid misjudgment caused by the separation of pure mathematical integration results from actual observations. For example, it can exclude the false data scenario of "the thickness of the supercooled layer at non-hail points exceeds 8 km", so that the confidence interval of non-hail samples (-∞,8] strictly fits the observation facts, and enhances the credibility of the indicator. The corrected high hail probability interval (8,+∞) directly corresponds to the physical feature of "the thickness of the supercooled layer exceeds the possible range of non-hail cloud systems". For example, when strong convective clouds develop vigorously, the thickness of the supercooled layer increases significantly. When it exceeds 8 kilometers, it can be clearly determined as a high probability hail signal, reducing the original data. The interval (8, 14, 18) in the academic solution addresses the indicator window problem caused by the lack of actual non-hail data support. The medium hail probability interval [4, 20, 8] covers the overlapping area of the supercooled layer thickness, including some hail samples (95% of the hail data not covered by the high probability interval), as well as the extreme values of non-hail samples. False alarms can be further filtered out through multi-indicator collaborative rules. For example, when other indicators, such as cloud top temperature and liquid water path, fall into the medium- and high-probability intervals at the same time, a comprehensive judgment is made that hail has a medium probability, improving the rigor of the recognition logic.
[0031] The multi-dimensional joint identification in step S5 also includes a weight allocation mechanism: weights are assigned to the cloud top temperature and blackbody brightness temperature indicators, with a weight coefficient ≥ 0.2; weights are assigned to the optical thickness and liquid water path auxiliary indicators, with a weight coefficient ≤ 0.1; and a comprehensive hail probability value is calculated through weighted summation. The weight allocation mechanism is determined based on the physical contribution of the indicators to hail formation: cloud top temperature and blackbody brightness temperature directly reflect the thermodynamic characteristics of cloud tops and are closely related to the supercooled water environment in which hail grows, so they are assigned higher weights of 0.2-0.3; optical thickness and liquid water path reflect the microphysical structure of the cloud system and are assigned lower weights of 0.05-0.1 as auxiliary criteria.
[0032] The method also includes a model validation step: using leave-one-out cross-validation (LOOCV) to test the effectiveness of quantitative identification indicators, calculate the hail sample identification accuracy, non-hail sample identification accuracy, and overall identification accuracy. The leave-one-out cross-validation method separates the data set into training and test sets one by one, retaining one sample as the test set each time and the remaining samples as the training set, ensuring that each sample participates in model training and validation, avoiding evaluation bias caused by random sampling. For example, in 368 groups of samples, each hail point and non-hail point was separately reserved for testing, so that the hail sample identification accuracy of 89.13% and the non-hail sample identification accuracy of 91.85% can truly reflect the model's ability to distinguish between the two types of data, and the overall identification accuracy of 87.50% comprehensively reflects the overall effectiveness of the indicators. From the perspective of reliability verification, this method can quantify the stability of the model by calculating the mean and variance through multiple iterations. For example, when validating an RBF kernel-based SVM model, the leave-one-out method showed a low standard deviation in the accuracy of identifying non-hail samples, indicating the model's robust discrimination of non-hail scenarios. Furthermore, the fluctuation range of the accuracy of hail samples was relatively small, demonstrating the consistency of the indicator in capturing hail characteristics. Furthermore, by comparing cross-validation results for different kernel function models, such as L-SVM, RBF-SVM, and S-SVM, the optimal model can be identified. For example, the RBF-SVM achieved the highest overall accuracy, thus avoiding model selection bias caused by a single training-testing split.
[0033] The method further integrates a support vector machine (SVM) model: using inversion product data as input features, the classification model is trained using linear, RBF, and sigmoid kernel functions. Grid search is used to optimize hyperparameters, ultimately selecting the RBF-SVM model as the preferred model. The SVM model uses kernel functions to map low-dimensional input features, such as cloud top height and cloud top temperature, into a high-dimensional space, effectively addressing the classification challenges of nonlinear features in hail identification. For example, indicators such as optical thickness and liquid water path have high distribution overlap in low-dimensional space, making them difficult to distinguish using linear boundaries. The RBF kernel, however, uses Gaussian radial basis functions to calculate the similarity between samples, constructing a nonlinear decision boundary in high-dimensional space and accurately capturing the implicit feature differences between hail and non-hail samples.
[0034] Sample balancing is introduced during the training of the SVM model. The SMOTE algorithm is used to generate synthetic samples for minority hail samples, ensuring a 1:1 ratio between hail samples and non-hail samples. While the number of hail samples and non-hail samples in the original dataset is balanced, hail events in real meteorological scenarios are typically low-probability events, and real data may be naturally imbalanced, such as non-hail samples far outnumbering hail samples. The SMOTE algorithm generates synthetic hail samples through interpolation, such as generating virtual samples based on k-nearest neighbor features. This can simulate the spatial distribution of the features of real hail clouds and prevent the model from biased towards learning majority features due to the sparse minority samples, such as overfitting the low cloud top height and high cloud top temperature of non-hail samples. For example, in the feature space of indicators such as optical thickness, hail samples may form several isolated clusters. SMOTE-generated samples can fill the gaps between clusters, enhancing the model's overall understanding of the hail feature distribution.
[0035] Example 2 Construction of spatiotemporal matching dataset Data acquisition: The FY-2G satellite inversion product data for 30 hail days from 2020 to 2022 were collected, including seven indicators: cloud top height (ztop), cloud top temperature (ttop), supercooled layer thickness (hsc), optical thickness (optn), effective particle radius (ref), liquid water path (lwp), and blackbody brightness temperature (tbb).
[0036] Taking the 184 hail points observed on the ground as the center, the satellite inversion data within 15 minutes before and after the hailfall were extracted according to the longitude and latitude coordinates, and a total of 184 groups of hail samples were used.
[0037] Selection of non-hail samples: Using the spatial random sampling method, 184 groups of satellite inversion data were randomly selected in the non-hail period and non-hail area with a spatial distance of ≥50 km from the hail point to ensure that their temporal resolution was consistent with the hail sample time window and that the spatial resolution (0.05°×0.05° grid) was aligned with the hail samples.
[0038] Data preprocessing: Samples that failed to match the time and space were eliminated, such as missing satellite data or abnormal quality marks. Finally, a time and space matching dataset containing 368 groups of samples (184 hail and 184 non-hail) was formed.
[0039] 2. Gaussian distribution suitability test Verification method: Draw quantile-quantile plots (QQ plots) for the seven inversion product data respectively, and compare the sample quantile points with the standard normal distribution line.
[0040] Calculate the Pearson correlation coefficient (PCC) to quantify the degree of coincidence: PCC ≥ 0.9: Determined to "obey Gaussian distribution", such as cloud top height, cloud top temperature, and supercooled layer thickness; 0.8≤PCC<0.9: Determined to be "approximately following Gaussian distribution", such as optical thickness, liquid water path, and blackbody brightness temperature; PCC<0.8: Excluded from use, such as effective particle radius ref.
[0041] Example: The PCC of cloud top height (ztop) is 0.92, which is determined to be Gaussian distributed; The PCC of the optical thickness (optn) is 0.85, which is determined to be approximately Gaussian distributed.
[0042] 3. Two-sample probability density modeling and long-tail correction Modeling steps: For the six indicators that obey / approximately obey Gaussian distribution (excluding ref), calculate the mathematical expectation μ and standard deviation σ of hail samples and non-hail samples respectively, and construct the probability density function of the two-sample Gaussian distribution: ,
[0043] For example, the cloud top temperature (ttop) of hail samples has a mean μ=-46.30℃ and a standard deviation σ=14.09℃; while that of non-hail samples has a μ=-23.58℃ and a σ=11.67℃.
[0044] Long-tail distribution correction: For OPTN, LWP, and TBB, the 3σ principle was used to eliminate outliers. For each indicator, the mean ±3 standard deviations of hail / non-hail samples were calculated, and data points outside this range were removed. For example, in the non-hail samples for the liquid water path LWP, data points with LWP > 201.53 + 3 × 188.45 = 766.88 were removed, and the corrected μ and σ were recalculated to 190.21 and 175.32, respectively.
[0045] 4. Dynamic division of confidence intervals and indicator optimization Integration calculation: Numerically integrate the probability density function of hail / non-hail samples for each indicator and calculate the one-sided 95% confidence interval: Hailfall sample: Calculate the infimum b so that the interval [b,+∞) Contains 95% hail data; Non-hail samples: Calculate the supremum a so that the interval (-∞, a] contains 95% of the non-hail data.
[0046] Physical constraint optimization: The maximum supercooled layer thickness at non-hail points in actual observations is 8 kilometers, so a=8 kilometers is imposed as the supremum of the non-hail samples. After correction, the following three probability intervals are used: low hail probability interval: [0, 4.20) kilometers, including 5% hail data; medium hail probability interval: [4.20, 8] kilometers, including the data overlap area; high hail probability interval: (8, +∞) kilometers, including only hail data. The final three-level probability intervals are used.
[0047] 5. Multi-dimensional joint identification and weight allocation Identification rules: Single indicator trigger rule: If any indicator falls into the high probability interval, it is directly judged as a high probability hail.
[0048] If the blackbody brightness temperature of a sample is -48℃, which falls into the high probability interval (-∞,−37.70)℃, it is directly judged as a high probability hail.
[0049] Multi-indicator collaborative rule: If at least three indicators fall into the medium probability interval and no indicator falls into the low probability interval, it is judged as medium probability hail; if the sample simultaneously meets the cloud top temperature = -30℃, supercooled layer thickness = 6 kilometers, liquid water path = 300, and no indicator is in the low probability interval, it is judged as medium probability hail.
[0050] Weight allocation mechanism: The core indicators, cloud top temperature and blackbody brightness temperature, are assigned a weight coefficient ≥ 0.2, such as 0.25. The auxiliary indicators, optical thickness and liquid water path, are assigned a weight coefficient ≤ 0.1, such as 0.08. The comprehensive hail probability value is calculated through weighted summation: ,in ∈{0,1} , whether it falls into the corresponding interval If the cloud top temperature weight of 0.25 and the blackbody brightness temperature weight of 0.25 fall into the medium probability interval, the optical thickness weight of 0.08 falls into the medium probability interval, and the liquid water path weight of 0.08 falls into the low probability interval, then P=0.25+0.25+0.08=0.58. The low probability indicators are not included in the sum. If the threshold P is set ≥0.5, it is judged as a medium probability hail.
[0051] 6. Model validation and optimization, leave-one-out cross-validation (LOOCV): leave one out of each of the 368 samples as a test set, and the remaining 367 as a training set. Calculate: Accuracy of hail sample recognition:
[0052] Accuracy of identifying non-hail samples:
[0053] Overall recognition accuracy:
[0054] Six metrics, ztop, ttop, hsc, optn, lwp, and tbb, were used as inputs for the SVM model. Linear-SVM, RBF-SVM, and Sigmoid-SVM models were trained, respectively. Grid search was used to optimize hyperparameters, such as C=10 and γ=0.1 for the RBF kernel. The SMOTE algorithm was used to oversample hail samples to a 1:1 ratio of hail to non-hail, improving the model's ability to identify minority classes.
[0055] The RBF-SVM model with the highest overall accuracy in cross-validation, such as 87.5%, is selected as the final classifier. Real-time inversion product data from the FY-2G satellite is received, and seven cloud indicators of the target area are extracted.
[0056] LOOCV was used to rotate samples one by one, calculating metrics such as the confusion matrix, accuracy, and false alarm rate to ensure model generalization. The RBF-SVM model was integrated into the meteorological warning system, which automatically outputs the hail probability level (high / medium / low) with real-time access to FY-2G satellite inversion data.
[0057] Combined with the geographic information system (GIS), hail risk areas can be visualized to provide precise guidance for artificial hail prevention operations and agricultural disaster prevention.
[0058] The above are only embodiments of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme is not described in detail here. Those of ordinary skill in the art are aware of all common technical knowledge in the technical field to which the invention belongs before the application date or priority date, can obtain all existing technologies in the field, and have the ability to apply conventional experimental means before that date. Those of ordinary skill in the art can improve and implement this scheme in combination with their own abilities under the enlightenment given by this application. Some typical known structures or known methods should not become obstacles for those of ordinary skill in the art to implement this application. It should be pointed out that for those skilled in the art, several variations and improvements can be made without departing from the structure of the present invention. These should also be regarded as the scope of protection of the present invention. These will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.
Claims
1. A quantitative identification method for hail from FY-2G satellite data based on Gaussian distribution is characterized by: The following steps are involved: S1. Construct a spatiotemporal matching dataset: Obtain inversion product data for cloud top height, cloud top temperature, supercooled layer thickness, optical depth, effective particle radius, liquid water path, and blackbody brightness temperature from the FY-2G satellite. Centered on the ground-observed hailfall point, extract satellite inversion data for a period of time before and after the hailfall as hail samples. Select an equal number of satellite inversion data from non-hailfall periods and spatially non-hailfall areas as non-hail samples, forming a dataset with spatiotemporal matching features. S2. Gaussian distribution fitness test: The distribution characteristics of the inversion product data are tested using a quantile-quantile plot. The degree of coincidence between the quantile points and the standard normal distribution line is used as the basis for judgment. It is determined that the cloud top height, cloud top temperature, and supercooled layer thickness follow a Gaussian distribution, while the optical thickness, liquid water path, and blackbody brightness temperature approximately follow a Gaussian distribution. S3. Two-sample probability density modeling: For inversion product data that follow or approximately follow a Gaussian distribution, the mathematical expectations of hail samples and non-hail samples are calculated separately. and standard deviation , construct the two-sample Gaussian distribution probability density function: ; S4. Dynamic division of confidence intervals: establish one-sided confidence intervals for hail samples and non-hail samples respectively: Calculate the infimum b of the hail samples and obtain the high hail probability interval [b, +∞); Calculate the supremum a for the non-hail samples and obtain the low hail probability interval (-∞,a]); According to the hail probability interval [a, b] defined in the overlapping part of the interval, a quantitative identification index containing three levels of probability gradient is formed; S5. Multi-dimensional joint identification: For the inversion product data of the sample to be identified, match them with the corresponding three-level probability intervals in step S4 and determine the probability of hail occurrence using the following rules: Single indicator trigger rule: If any indicator falls into the high hail probability range, it is directly determined as a high probability hail; Multi-indicator coordination rule: If at least three indicators fall into the medium hail probability range and no indicator falls into the low probability range, it is determined to be a medium probability hail. Background exclusion rule: If all indicators fall into the low probability range, it is determined that there is no hail.
2. The FY-2G satellite data hail quantification identification method based on Gaussian distribution according to claim 1, wherein The method for constructing the spatiotemporal matching dataset in step S1 is as follows: taking the 184 hail points on 30 hail days as the center, extracting the satellite inversion data at the corresponding time according to the longitude and latitude coordinates, and selecting non-hail point data through spatial random sampling to ensure that the two types of samples are aligned in terms of temporal resolution and spatial resolution.
3. The FY-2G satellite data hail quantification identification method based on Gaussian distribution according to claim 1, wherein The judgment criteria for the Gaussian distribution adaptability test in step S2 are: when the Pearson correlation coefficient between the quantile point of the inversion product data and the standard normal distribution line is ≥0.9, it is judged to obey the Gaussian distribution; when the correlation coefficient is between 0.8-0.9, it is judged to approximately obey the Gaussian distribution.
4. The FY-2G satellite data hail quantification identification method based on Gaussian distribution according to claim 1, wherein The double-sample probability density modeling in step S3 further includes correcting the long-tail distribution of optical thickness, liquid-water path, and blackbody brightness temperature: by introducing a truncation threshold, recalculating the mathematical expectation and standard deviation after removing outliers, the model's fitting accuracy for non-ideal distribution data is improved.
5. The FY-2G satellite data hail quantitative identification method based on Gaussian distribution according to claim 1 is characterized in that, In step S4, the dynamic division of the confidence interval is solved by numerical integration method to solve the probability density function, which satisfies: the one-sided confidence interval of the hail sample [b, +∞), satisfies ; The one-sided confidence interval of the non-hail sample (-∞,a], which satisfies ,in and are the Gaussian probability density functions of hail and non-hail samples, respectively.
6. The method for quantitatively identifying hail using FY-2G satellite data based on Gaussian distribution according to claim 1, wherein: The identification index for the supercooled layer thickness in step S4 is optimized as follows: based on the constraint that the maximum supercooled layer thickness at the non-hailfall point in the actual observation data is 8 kilometers, the supremum a of the non-hailfall samples is forcibly limited to 8 kilometers, forming a revised medium hail probability interval [4.20, 8] and a high hail probability interval (8, +∞).
7. The method for quantitatively identifying hail using FY-2G satellite data based on Gaussian distribution according to claim 1, wherein: The multi-dimensional joint identification in step S5 also includes a weight distribution mechanism: weights are assigned to the cloud top temperature and blackbody brightness temperature indicators, with a weight coefficient of 0.2-0.3; weights are assigned to the optical thickness and liquid water path auxiliary indicators, with a weight coefficient of 0.05-0.1; and the weighted summation formula is used. Calculate the comprehensive hail probability value, where Is a binary variable that indicates whether the indicator falls into the corresponding probability interval. If it does, it takes 1, otherwise it takes 0. is the weight coefficient of the corresponding indicator.
8. The method for quantitatively identifying hail using FY-2G satellite data based on Gaussian distribution according to claim 1, wherein: The S5 also includes a model validation step: using leave-one-out cross validation (LOOCV) to test the effectiveness of the quantitative identification indicators, and calculating the hail sample identification accuracy, non-hail sample identification accuracy and overall identification accuracy.
9. The method for quantitatively identifying hail using FY-2G satellite data based on Gaussian distribution according to claim 1, wherein: The S5 further integrates a support vector machine (SVM) model: the inversion product data is used as input features, and the classification model is trained using Linear, RBF, and Sigmoid kernel functions respectively. The hyperparameters are optimized through grid search, and the RBF-SVM model is finally selected as the preferred model.
10. The method for quantitatively identifying hail using FY-2G satellite data based on Gaussian distribution according to claim 9, wherein: Sample balancing is introduced during the training of the SVM model: the SMOTE algorithm is used to generate synthetic samples for the minority hail samples, so that the ratio of hail samples to non-hail samples reaches 1:1.
Citation Information
Patent Citations
Ground hail shooting identification and early warning method based on dual linear polarization radar
CN113933845A
Short-time rainstorm weather background automatic discrimination method, system, device and terminal
CN115902812A
Data feature enhanced artificial intelligence dual-polarization radar quantitative rainfall estimation method
CN120143161A
Hail early warning method and system based on lightning jump and support vector machine
CN120233320A
Hail Frequency Predictions Using Artificial Intelligence
US20240125972A1