Hail probability and size identification algorithm based on fuzzy logic and semi-supervised learning
By combining fuzzy logic and semi-supervised learning algorithms, the probability and size of hail are identified, and the problem of high false alarm rate in the existing technology is solved, which improves the accuracy and efficiency of hail recognition.
Patent Information
- Application Number
- CN202510038806.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, hail recognition algorithms have the problem of high false alarm rates, and there is a lack of attempts to use fuzzy logic and semi-supervised learning algorithms for hail recognition.
The hail probability and size recognition algorithm based on fuzzy logic and semi-supervised learning is used to calculate the heights of 0℃ and -20℃ layers, and the hail probability is comprehensively identified using the fuzzy logic algorithm, and the self-training semi-supervised learning algorithm is used to identify the size of hail.
The hail recognition rate has been improved, the false alarm rate has been significantly reduced, the application of dual polarization radar data has been strengthened, and the forecast and early warning of hail weather has been provided.
Smart Images

Figure CN119939370A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of weather phenomenon analysis and prediction, and in particular to a hail probability and size recognition algorithm based on fuzzy logic and semi-supervised learning. Background Art
[0002] Hail is one of the meteorological disasters. Hail develops and changes rapidly, has a short life cycle, is sudden and destructive, and is difficult to predict. It often causes great harm to industry, agriculture, and people’s production and life safety. Weather radar data plays a key role in the monitoring, analysis, and short-term warning of severe convective weather. Domestic and foreign scholars have done a lot of research on hail weather, but there has been no attempt to identify hail by integrating fuzzy logic and semi-supervised machine learning algorithms. There are few studies on the probability identification of hail based on dual-polarization radar and sounding data in the existing technology, especially the use of radar dual-polarization parameters, and the hail identification algorithm has a high false alarm rate. Summary of the invention
[0003] The technical problem to be solved by the present invention is to provide a hail probability and size recognition algorithm based on fuzzy logic and semi-supervised learning, which can effectively solve the problems raised in the above background technology.
[0004] To solve the above problems, the technical solution adopted by the present invention is: a hail probability and size recognition algorithm based on fuzzy logic and semi-supervised learning, comprising the following steps:
[0005] S1. Calculate the height of the 0℃ layer and the -20℃ layer. The polynomial is defined as:
[0006]
[0007] Where y is the dependent variable temperature, x is the independent variable height, M is the order of the polynomial, w0,...wM are the coefficients of the polynomial, by determining the value of the coefficient, a polynomial function polynomial fitting model can be obtained, and then by fitting the known data points, the value of the unknown data points can be predicted;
[0008] S2. Use fuzzy logic algorithm to comprehensively identify the probability of hail based on multiple characteristic parameters of radar base data and sounding data, and determine the membership function and weight of each characteristic parameter;
[0009] S3. Calculate the probability of hail occurrence by using fuzzy logic algorithm and determined weight coefficients;
[0010] S4. Based on the hail probability recognition results, the size of the hail is identified using the self-training semi-supervised learning algorithm and label classification is performed;
[0011] S5. Verify and evaluate the algorithm results, and evaluate the accuracy of the algorithm recognition results by comparing them with the actual hail locations.
[0012] As a further preferred embodiment of the present invention, the characteristic parameters in step S2 include combined reflectivity (CR), the height difference between the basic reflectivity 55dBZ and the 0°C layer (H0), the height difference between the basic reflectivity 45dBZ and the -20°C layer (H-20), the vertically integrated liquid water volume (VIL), the vertically integrated liquid water density (VILD) and the echo top height (ET).
[0013] As a further preferred embodiment of the present invention, the characteristic parameters also include a differential reflectivity factor (ZDR), a differential propagation phase shift rate (KDP) and a correlation coefficient (CC). In step S4, the combined reflectivity (CR), the height difference between the basic reflectivity 55dBZ and the 0°C layer (H0), the height difference between the basic reflectivity 45dBZ and the -20°C layer (H-20), the vertically integrated liquid water volume (VIL), the vertically integrated liquid water density (VILD), the echo top height (ET), the differential reflectivity factor (ZDR), the differential propagation phase shift rate (KDP), and the correlation coefficient (CC) are used as inputs for classification to obtain hail size labels.
[0014] As a further preferred embodiment of the present invention, in step 4, a semi-supervised learning algorithm Self-training is used to identify the size of hailstones, and three size labels of ≤0.5cm, 0.5cm-2cm, and ≥2cm are set for classification. For labels that cannot be determined as false alarm data, they are recorded as unlabeled data.
[0015] As a further preferred embodiment of the present invention, the prediction using the self-training semi-supervised learning algorithm includes the following steps:
[0016] a. Train a weak classifier using a small number of labeled samples;
[0017] b. Use this weak classifier to predict unlabeled data samples, and formulate a strategy to use the predicted partial labels as the true labels of the samples, and combine them with the labeled data to obtain new expanded samples
[0018] c. Use all the currently labeled augmented samples to train the classifier. If the number of iterations is reached or there are no new labels to add, proceed to step 4. Otherwise, continue to step 2 to predict the unlabeled samples.
[0019] d. Use the classifier trained in step c to evaluate on the test sample set.
[0020] As a further preferred solution of the present invention, the weak classifier includes a support vector machine SVM and a K-nearest neighbor algorithm KNN.
[0021] Compared with the prior art, the present invention provides a hail probability and size recognition algorithm based on fuzzy logic and semi-supervised learning, which has the following beneficial effects:
[0022] This algorithm improves the recognition rate of hail and significantly reduces the false alarm rate, further strengthening the application of dual-polarization radar data and providing support for the forecast and warning of hail weather. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is the membership function of the six characteristic parameters;
[0024] Figure 2 It is the process of self-training semi-supervised learning algorithm;
[0025] Figure 3 The accuracy of using KNN or SVM as the basic weak classifier;
[0026] Figure 4 The relationship between hail identification results and actual location; DETAILED DESCRIPTION
[0027] If "and / or" or "and / or" appears in the full text, its meaning includes three parallel options. Taking "A and / or B" as an example, it includes option A, or option B, or a option in which both A and B are satisfied.
[0028] In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the fact that ordinary technicians in the field can implement it. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0029] Reference Figure 1-4 , as a specific embodiment of the present invention:
[0030] The present invention provides a hail probability and size recognition algorithm based on fuzzy logic and semi-supervised learning, taking Henan Province as an example, comprising the following steps:
[0031] Step 1: Calculate the height of the 0℃ layer and the -20℃ layer
[0032] Real-time sounding temperature data of Zhengzhou, Nanyang and Lushi in Henan Province are obtained, and the polynomial function is used to fit and calculate the height of the 0℃ layer and the -20℃ layer. The polynomial function is defined as:
[0033]
[0034] Where y is the dependent variable temperature, x is the independent variable height, M is the order of the polynomial, and w0,...wM are the coefficients of the polynomial. By determining the values of the coefficients, a polynomial function polynomial fitting model can be obtained, and then the values of unknown data points can be predicted by fitting known data points.
[0035] Step 2: Comprehensive identification of hail probability. Fuzzy logic algorithm is used to determine the membership function and weight of each characteristic parameter to conduct comprehensive identification of hail probability. Based on expert experience, the membership functions of six characteristic parameters are determined as follows: Figure 1 shown.
[0036] Step 3
[0037] Determine the weight coefficient and calculate the probability of hail
[0038] After statistical analysis, the weight coefficient combination is 0.3, 0.05, 0.09, 0.20, 0.22, and 0.14. A weighted value is obtained from the weight coefficient combination, which is recorded as CP. In order to reduce the false alarm rate, the present invention subtracts 0.1 from the CP value and defines the hail probability formula as shown in formula (2).
[0039]
[0040] Step 4: Set the system ECU to obtain the current angle of the steering wheel from the electric power steering system (EPS) through the CAN bus;
[0041] Hail size identification
[0042] Based on the above results, machine learning-semi-supervised learning algorithm is used to further identify the size of hail, and three labels are set: ≤0.5cm, 0.5cm-2cm, and ≥2cm, which are recorded as 0, 1, and 2 respectively. Among them, false alarm data refers to data where the fuzzy logic algorithm identifies that there is hail but there is no actual hail. However, the actual hail comes from stations or manual observations, which makes the data incomplete. In this way, there is a situation where hail has actually occurred somewhere but there is no actual hail data. In the false alarm data, there is a situation where the fuzzy logic algorithm identifies that the presence of hail is "true". Therefore, the label of the false alarm data cannot be determined and is recorded as unlabeled data.
[0043] The above situation needs further analysis. For the situation with both labeled data and unlabeled data, the semi-supervised learning algorithm is the most suitable, that is, the semi-supervised learning algorithm contains a large amount of unlabeled data and a small amount of labeled data, and mainly uses the information in the unlabeled data to assist the labeled data for supervised learning.
[0044] The present invention adopts a self-training semi-supervised learning algorithm, such as Figure 2As shown, a weak classifier is first trained with a small number of labeled samples;
[0045] This weak classifier is then used to predict the unlabeled data samples, and according to a certain strategy (satisfying the prediction probability is greater than or equal to 90%), the predicted partial labels are used as the true labels of the samples, and combined with the labeled data to obtain new expanded samples;
[0046] The third step is to use all the current labeled expansion samples (including the samples processed in the second step) to train the classifier. If the stopping condition is reached (the number of iterations is greater than 1000 or there are no new labels to be added), the fourth step is entered. Otherwise, the second step is continued to predict the unlabeled samples.
[0047] The fourth step is to evaluate the classifier trained in the third step on the test sample set.
[0048] Among them, the first step needs to train a weak classifier. The present invention selects support vector machine (SVM) and K-nearest neighbor algorithm (KNN) for classification and recognition. The input feature quantity is expanded to 9 feature parameters such as CR, H0, H-20, VIL, VILD, ET, ZDR, KDP, and CC.
[0049] According to the hail probability hit results obtained by the above fuzzy logic, 3404 grid points with labels and 180894 grid points without labels were obtained. Among them, the test samples selected 20% of the total number of labels, and the marks that met the prediction probability of more than 90% were selected as pseudo labels as pre-labeled samples, and combined with the labeled data to obtain the expanded samples. The following results were finally obtained Figure 3 As shown in the results, the accuracy of using KNN as the basic weak classifier is 83%, which is higher than that of SVM.
[0050] Finally, the prediction was performed on 180,894 unmarked grid points, and the results were shown in Table 1. It can be seen that through further identification by machine learning, 10,850 grid points were judged to have hail, of which 9,168 grid points were of size "≤0.5 cm", 1,682 grid points were of size "0.5 cm-2 cm", and the remaining 170,044 grid points were judged to have no hail.
[0051] Table 1 Self-training algorithm recognition results
[0052] Classification ≤0.5cm 0.5cm-2cm ≥ 2 cm quantity 9168 1682 0
[0053] Step 5: Inspection and Evaluation
[0054] Experimental cases were selected for testing and evaluation. Figure 4The figure shows the recognition results of this algorithm. The blue circles are the locations of the three hailstorms that actually occurred and their 20km range. The red dots are the locations where the probability of this algorithm being recognized is greater than 50%.
[0055] Table 2 Evaluation results
[0056]
[0057] From the evaluation results shown in Table 2, it can be seen that the recognition rate of the algorithm for the hail process reached 99.33%, and the accuracy was 68.66%, which is 22.51% higher than the accuracy of the PUP product data (UAM) (46.15%).
[0058] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. All equivalent structural changes made by using the contents of the present invention specification and drawings under the inventive concept of the present invention, or directly / indirectly applied in other related technical fields are included in the patent protection scope of the present invention.
Claims
1. A hail probability and size recognition algorithm based on fuzzy logic and semi-supervised learning, characterized in that: The following steps are involved: S1. Calculate the height of the 0℃ layer and the -20℃ layer. The polynomial is defined as: Where y is the dependent variable temperature, x is the independent variable height, M is the order of the polynomial, w0,...wM are the coefficients of the polynomial, by determining the value of the coefficient, a polynomial function polynomial fitting model can be obtained, and then by fitting the known data points, the value of the unknown data points can be predicted; S2. Use fuzzy logic algorithm to comprehensively identify the probability of hail based on multiple characteristic parameters of radar base data and sounding data, and determine the membership function and weight of each characteristic parameter; S3. Calculate the probability of hail occurrence by using fuzzy logic algorithm and determined weight coefficients; S4. Based on the hail probability recognition results, the size of the hail is identified using the self-training semi-supervised learning algorithm and label classification is performed; S5. Verify and evaluate the algorithm results, and evaluate the accuracy of the algorithm recognition results by comparing them with the actual hail locations.
2. The hail probability and size recognition algorithm based on fuzzy logic and semi-supervised learning according to claim 1 is characterized in that: The characteristic parameters in step S2 include combined reflectivity (CR), the height difference between the basic reflectivity 55dBZ and the 0°C layer (H0), the height difference between the basic reflectivity 45dBZ and the -20°C layer (H-20), the vertically integrated liquid water volume (VIL), the vertically integrated liquid water density (VILD) and the echo top height (ET).
3. The hail probability and size recognition algorithm based on fuzzy logic and semi-supervised learning according to claim 2 is characterized in that: The characteristic parameters also include a differential reflectivity factor (ZDR), a differential propagation phase shift rate (KDP) and a correlation coefficient (CC). In step S4, the combined reflectivity (CR), the height difference between the basic reflectivity 55dBZ and the 0°C layer (H0), the height difference between the basic reflectivity 45dBZ and the -20°C layer (H-20), the vertically integrated liquid water volume (VIL), the vertically integrated liquid water density (VILD), the echo top height (ET), the differential reflectivity factor (ZDR), the differential propagation phase shift rate (KDP), and the correlation coefficient (CC) are used as inputs for classification to obtain the size label of the hail.
4. The hail probability and size recognition algorithm based on fuzzy logic and semi-supervised learning according to claim 3 is characterized in that: In step 4, a semi-supervised learning algorithm Self-training is used to identify the size of hailstones, and three size labels of ≤0.5cm, 0.5cm-2cm, and ≥2cm are set for classification. Labels that cannot be determined as false alarm data are recorded as unlabeled data.
5. The hail probability and size recognition algorithm based on fuzzy logic and semi-supervised learning according to claim 4 is characterized in that: The prediction using the Self-training semi-supervised learning algorithm includes the following steps: a. Train a weak classifier using a small number of labeled samples; b. Use this weak classifier to predict unlabeled data samples, and formulate a strategy to use the predicted partial labels as the true labels of the samples, and combine them with the labeled data to obtain new expanded samples c. Use all the currently labeled augmented samples to train the classifier. If the number of iterations is reached or there are no new labels to add, proceed to step 4. Otherwise, continue to step 2 to predict the unlabeled samples. d. Use the classifier trained in step c to evaluate on the test sample set.
6. The hail probability and size recognition algorithm based on fuzzy logic and semi-supervised learning according to claim 5 is characterized in that: The weak classifiers include support vector machine SVM and K-nearest neighbor algorithm KNN.
Citation Information
Cited By
Hail prediction method and device, medium and program product
CN120352956A
Hail prediction method, device, medium and program product
CN120352956B