Gale event classification method based on iterative clustering analysis
By combining iterative K-means clustering and random forest algorithms, strong wind events are automatically identified and classified, solving the problems of low efficiency and limited accuracy in existing technologies, and achieving efficient and interpretable classification of strong wind events.
Patent Information
- Application Number
- CN202610079024.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-21
- Publication Date
- 2026-02-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies are inefficient and inconsistent in identifying and classifying strong wind events. They are heavily influenced by researchers' experience and are prone to missing short-lived or small-scale strong wind events. Traditional methods rely on the spatial distribution of stations and radar features, which limits the accuracy of identification.
By employing iterative K-means clustering analysis combined with the random forest algorithm, and through environmental physical quantity calculation and dual evaluation, the system automatically identifies and classifies strong wind events into types such as cold air, typhoons, and convective systems, and performs secondary clustering optimization using key feature variables.
It improves the efficiency and accuracy of gale event classification, enhances the meteorological interpretability of classification results, directly relates to the intrinsic physical characteristics of weather systems, and avoids the limitations of traditional methods.
Smart Images

Figure CN121542865A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of strong wind event technology, and specifically to a strong wind event classification method based on iterative clustering analysis. Background Technology
[0002] Currently, the identification and classification of severe wind events mainly rely on manual case screening and subjective judgment, a method particularly prevalent in studies of small- and medium-scale weather processes such as thunderstorms and strong winds. However, this method has significant limitations, including low efficiency, poor consistency, heavy reliance on researcher experience, and a tendency to miss some transient or small-scale wind events. In recent years, artificial intelligence technology has developed rapidly and has already achieved initial applications in severe convective wind warnings. Existing technologies for identifying typhoon gale events based on meteorological station data primarily involve screening stations with daily maximum wind speeds of at least 10.8 m / s (equivalent to a Force 6 gale) from meteorological observation station data and defining these stations as gale stations. Based on the spatial distribution characteristics of the selected gale stations, several independent wind belts are identified, and the weighted center of gravity of each wind belt is calculated. Based on the distance between the typhoon center and each major gale station, as well as the distance between the typhoon center and the center of gravity of each wind belt, a threshold is set to comprehensively determine whether a station belongs to a typhoon wind belt. If the number of valid stations within a identified typhoon wind belt is no less than three, then the gale events recorded by all stations within that wind belt are determined to be caused by a typhoon system. This scheme is an improvement on the typhoon rain belt identification method, and its effectiveness in identifying gale events largely depends on the selection of key parameters D0 and D1. In practical applications, this scheme has limited ability to identify typhoons affecting land from a distance, and its identification accuracy is easily constrained by the spatial density of ground observation stations. The method of automatically identifying thunderstorms and strong winds using radar data by employing machine learning and deep learning primarily identifies sites with instantaneous wind speeds of at least 17.2 m / s observed by automatic weather stations, along with lightning activity within a 10 km radius and a radar reflectivity factor greater than 30 dBZ, as thunderstorm and strong wind events. A dataset is constructed based on the selected thunderstorm and strong wind records and divided into training, validation, and test sets in a 7:2:1 ratio. Three types of models—decision tree, CNN, and YOLO—are introduced, each using different radar feature factors as input, to train the three models, ultimately achieving intelligent identification and prediction of thunderstorms and strong winds. While this approach extracts key identification features from radar data, it does not deeply integrate the physical causes of thunderstorms and strong winds, which may limit the model's interpretability and generalization ability to some extent. Summary of the Invention
[0003] Purpose of the invention: The purpose of this invention is to provide a method for classifying strong wind events based on iterative clustering analysis. By using an attribution scheme based on iterative K-means clustering analysis and random forest algorithm, it can automatically identify strong wind events in a region and classify them into several specific types, thus solving the problems existing in the background technology.
[0004] Technical solution: The present invention provides a method for classifying strong wind events based on iterative clustering analysis, comprising the following steps:
[0005] (1) Collect observation data, reanalysis data and precipitation data from automatic weather stations in the target area, screen and extract samples of gale events, and calculate various environmental physical quantities corresponding to each gale event;
[0006] (2) Based on environmental physical quantities, an unsupervised clustering algorithm was used to perform the first clustering analysis on all samples of strong wind events to obtain preliminary results of strong wind type classification;
[0007] (3) Based on the initial classification results, the importance of each environmental physical quantity to different wind types is evaluated by the random forest algorithm, and key feature variables are selected. Using the key feature variables, secondary clustering analysis is performed on wind events that were misjudged or had ambiguous boundaries in the initial classification results, and their types are reclassified to obtain the final classification results of wind weather systems.
[0008] Furthermore, in step (1), based on whether the gale event is accompanied by precipitation, the gale event is divided into two categories: no precipitation type and precipitation type, and the corresponding environmental physical quantities are calculated respectively.
[0009] Furthermore, in step (2), the optimal number of clusters for the unsupervised clustering algorithm is determined using a silhouette coefficient-based method.
[0010] Furthermore, step (2) also includes: conducting a dual evaluation of the initial classification results, including: objectively scoring the comprehensive evaluation index calculated based on accuracy, precision and recall, and conducting a rationality analysis of the environmental physical quantity characteristics of each type of strong wind; if the dual evaluation results meet the preset conditions, then proceed to step (3) for secondary classification; otherwise, return to step (2) to adjust the parameters and re-perform the initial classification.
[0011] Furthermore, the comprehensive evaluation index is derived by calculating the weighted inverse average of accuracy, precision, and recall.
[0012] Furthermore, in step (2), the unsupervised clustering algorithm is the K-means clustering algorithm.
[0013] Furthermore, in step (3), the random forest algorithm uses the average impurity reduction method based on the Gini index to calculate the feature importance score of each environmental physical quantity, and selects key feature variables based on the score.
[0014] Furthermore, in step (3), the environmental physical quantities include one or more of the following: low-level vertical wind shear, convective effective potential energy, frontogenetic function, distance between the high wind station and the nearest precipitation area, precipitation area and maximum rainfall intensity, surface dew point temperature difference, air pressure change and temperature change.
[0016] Furthermore, in step (3), the final classification results of the gale weather system include cold air activity type, typhoon type and convective system type.
[0017] Beneficial Effects: Compared with existing technologies, this invention has the following significant advantages: The strong wind attribution scheme constructed in this invention differs from traditional methods that only identify a single type. This scheme can simultaneously output classification results for different causal types, such as cold air, typhoons, and convective systems, for all strong wind events. This improves the efficiency of strong wind classification and avoids the cumbersome process of using different methods for different weather systems. This invention overcomes the limitations of traditional strong wind identification methods that rely on indirect indicators such as station spatial distribution, fixed thresholds, and radar characteristics, using environmental physical quantities calculated based on conventional data as classification indicators. This not only improves the standardization and accessibility of data sources but, more importantly, allows the classification results to directly correlate with the intrinsic physical characteristics of the weather system, significantly enhancing the meteorological interpretability of the classification. Attached Figure Description
[0018] Figure 1 This is a flowchart of the present invention;
[0019] Figure 2 This is an objective score after two classifications in this invention; wherein Figure 2 (a) in the text represents the initial classification score; Figure 2 (b) in the table represents the second classification score. Detailed Implementation
[0020] This invention provides a method for classifying strong wind events based on iterative clustering analysis, comprising the following steps:
[0021] Step 1: Data Preparation. Automatic weather station data, CMORPH hourly precipitation data, and ERA5 reanalysis data collected from the study area between 2020 and 2023 were collected. Data Preprocessing. From the automatic weather station data, observations of stations with a 2-minute maximum wind speed greater than or equal to 17.2 m / s in the study area between 2020 and 2023 were selected and recorded as all gale events. Quality control was performed on the automatic weather station data, and anomaly checks were conducted based on the individual station wind field observation data from 2020 to 2023, removing abnormal stations. Construction of the Gale Dataset. Based on the latitude and longitude of the gale stations, meteorological elements and precipitation amounts of the corresponding latitude and longitude ERA5 reanalysis data and fused precipitation data were obtained through spatial interpolation. The temporal resolution of the ERA5 reanalysis data was four times a day, at 00, 06, 12, and 18 UTC, with an interval of 6 hours; the temporal resolution of the AWS and fused precipitation data was 1 hour. By searching backwards from the moment the strong wind occurred, the nearest ERA5 data point is used to extract upper-level data representing the environmental factors at the time of the strong wind. If strong winds occur consecutively at a certain station within a 6-hour interval of ERA5 data, the nearest ERA5 data point is searched backwards from the time of the first strong wind occurrence to represent the environmental factors. Based on automatic weather station data, precipitation data and reanalysis data are integrated to construct a strong wind dataset that includes environmental factors related to strong winds.
[0022] Step 2: Calculate environmental physical quantities. Based on the strong wind dataset, strong winds are classified into two categories: winds without precipitation and winds with precipitation, depending on whether precipitation accompanies the occurrence of strong winds at the strong wind stations. Key environmental physical quantities corresponding to each strong wind station are calculated, including the 0-3km vertical wind shear, 850hPa frontogene function, convective available potential energy of the mixing layer, and convective suppression energy. For winds without precipitation, the distance between the strong wind station and the nearest precipitation area, the area of that precipitation area, and the maximum hourly rainfall intensity are further calculated; for winds with precipitation, the maximum hourly rainfall intensity and precipitation area of the precipitation area where the station is located are recorded.
[0023] Step 3: Initial Classification. After constructing the high wind dataset and calculating the corresponding environmental physical quantities for each station, K-means clustering is used to classify high wind events with and without precipitation into separate categories. Details are as follows:
[0024] Selection of the number of clusters: The choice of the number of clusters K is crucial to the results of K-means clustering analysis. Therefore, the silhouette coefficient is usually chosen to determine the optimal number of clusters K. SC reflects the similarity between a sample point and its cluster. The specific calculation method is as follows:
[0025] (1);
[0026] in This represents the average distance between the i-th sample point and other members of the same cluster. SC represents the average distance from the i-th sample point to its nearest cluster. The value of SC ranges from -1 to 1: a higher SC value indicates better classification performance. The optimal number of clusters corresponds to the maximum average SC value across all sample points. In this invention, when classifying strong wind types, the maximum average SC value corresponding to different numbers of clusters is 0.357, corresponding to a cluster number of 3.
[0027] After determining the number of clusters, the K-means algorithm will be used to classify the wind types. The K-means clustering algorithm is a computationally efficient non-hierarchical clustering algorithm in unsupervised learning. In each iteration, the algorithm randomly searches for K cluster centers for initialization, and then calculates the data points (…). ) and cluster center ( The distance between () ), A smaller value indicates a higher degree of similarity among data points within the same cluster. The objective is to minimize the following objective function:
[0028] (2);
[0029] A feedback mechanism based on dual evaluation: To verify the necessity and effectiveness of the classification results, this invention introduces a feedback mechanism based on dual evaluation. First, the classification results are evaluated by both objective scoring and differences in physical characteristics; then, a decision is made based on the evaluation results—if the evaluation results are reasonable, a secondary classification process is initiated; otherwise, the initial classification is returned to and re-executed, thus forming a closed-loop quality control process.
[0030] For objective scoring, based on weather records from 2020 to 2023, accuracy, precision, and recall are introduced to quantitatively evaluate the results. Their definitions are as follows:
[0031] (3);
[0032] (4);
[0033] (5);
[0034] In this model, positive samples (P) represent the number of samples belonging to a certain wind type, while negative samples (N) represent the number of samples not belonging to that type. The evaluation metrics based on the confusion matrix are as follows: True Positive (TP) refers to the number of observations that actually belong to that wind type and are correctly classified as such; False Negative (FN) refers to the number of observations that actually belong to that type but are incorrectly classified as other types; False Positive (FP) refers to the number of observations that do not actually belong to that type but are incorrectly classified as such; and True Negative (TN) refers to the number of observations that do not actually belong to that type but are correctly classified as not belonging to that type. Therefore, accuracy measures the proportion of all samples (including positive and negative samples) correctly classified for that wind type; precision reflects the proportion of samples classified as belonging to that type that actually belong to that wind type; and recall represents the proportion of observations that actually belong to that type that are correctly identified.
[0035] Based on accuracy, precision, and recall, the formula for calculating the comprehensive evaluation index G is as follows:
[0036] (6);
[0037] In addition to objective scoring, it is also necessary to analyze the physical characteristics of each type of strong wind, observe whether there are significant differences in the physical characteristics of each type of strong wind, and combine the two evaluation methods to complete the assessment of the classification results.
[0038] After the initial classification and evaluation of the classification results, the final classification of strong winds is determined. At this point, the type of strong winds is classified into three categories: the weather systems that cause strong winds are mainly cold air activity, typhoons, and thunderstorm systems.
[0039] Step 4: Secondary Classification. After the initial classification identifies the weather system, a case-by-case analysis of the classification results reveals a certain degree of misclassification. To reduce this proportion, this invention employs a random forest method. For different types of strong winds, the environmental variables used in the initial clustering are ranked by feature importance. Representative meteorological variables from cold air, typhoon, and convective wind types are selected for secondary classification based on the initial classification, improving the accuracy of wind classification. The clustering elements are shown in Table 1.
[0040] Table 1. Representative meteorological elements of K-means clustering in primary and secondary classification. ;
[0041] Step 5: Random Forest Evaluation: This invention utilizes the Mean Decrease Impurity (MDI) method in random forests to calculate the Variable Importance Measures (VIM) scores for different features based on the Gini index. The VIM calculation steps are as follows: Calculation features Importance of node q in the i-th tree : (7); in The Gini index, and These represent the Gini indexes of the two new nodes before and after the node.
[0042] Calculation features The importance of the i-th tree :
[0043] (8); Where Q is a feature The set of nodes that appear in the i-th tree.
[0044] Statistical characteristics Importance rating among all trees And perform normalization:
[0045] (9); Where I is the number of decision trees in the random forest.
[0046] For misclassified sites, typhoon-type gales mixed with convective system gales were removed and reclassified based on the typhoon's onset time. For misclassified sites within cold air gales and typhoon systems, K-means clustering was used to separate them based on their key physical characteristics, and their types were reclassified. The corresponding clustering elements are shown in Table 1. A small category appeared in the secondary classification, lacking characteristics of cold air activity, typhoons, and convective gales, and was classified as other types of gales.
[0047] like Figure 2 As shown, the objective scoring results of the three types of strong winds after the initial and secondary classifications are presented. Figure 2 In (a) of the initial classification, the precision rates for cold air-type gales and typhoon-type gales were 0.67 and 0.65, respectively, while the recall rate for convective gales was 0.49. The G values for all three types of gales were above 0.6. Figure 2In (b) of the study, the accuracy, precision, recall, and G of the three types of strong winds were significantly improved after the secondary classification. The precision of cold air-type and typhoon-type strong winds improved to 0.94 and 0.98, respectively, and the recall of convective strong winds improved to 0.83. Furthermore, after analyzing the results of the secondary classification, in addition to the above three types of strong wind events, there were some samples in the study area with complex causes that were difficult to classify into the above three categories. These samples were classified as other types of strong winds. Based on this scheme, the strong winds in the study area were divided into four types according to their causal system through two classifications: cold air-type (40.3%), typhoon-type (27.9%), convective (22.2%), and other types (9.6%). In conclusion, the strong wind attribution scheme can be effectively used for the type identification of strong wind events.
Claims
1. A method for classifying a strong wind event based on iterative cluster analysis, characterized in that, The method comprises the following steps: (1) collecting automatic weather station observation data, reanalysis data and precipitation data in a target area, screening and extracting gale event samples, and calculating a plurality of environmental physical quantities corresponding to each gale event; (2) based on the environmental physical quantities, using an unsupervised clustering algorithm to perform first clustering analysis on all gale event samples, and obtaining a preliminary gale type classification result; (3) based on the preliminary classification result, using a random forest algorithm to evaluate the importance of each environmental physical quantity to different gale types, and screening out key feature variables; Using the key feature variables, performing secondary clustering analysis on the gale events misjudged or with fuzzy boundaries in the preliminary classification result, and reclassifying the types thereof, to obtain a final gale weather system classification result.
2. The method according to claim 1, wherein, In step (1), according to whether precipitation is accompanied when a gale event occurs, the gale event is divided into two types of no precipitation and precipitation, and the corresponding environmental physical quantities are calculated.
3. The method according to claim 2, wherein, In step (2), the method based on the contour coefficient is used to determine the optimal clustering number of the unsupervised clustering algorithm.
4. The method according to claim 2, wherein, Step (2) further comprises double evaluation of the preliminary classification result, including objective scoring based on a comprehensive evaluation index calculated by accuracy, precision and recall rate, and reasonable analysis of the environmental physical quantity characteristics of each type of gale; if the double evaluation result meets the preset condition, proceed to step (3) for secondary classification; otherwise, return to step (2) to adjust the parameters and perform preliminary classification again.
5. The method according to claim 4, wherein, The comprehensive evaluation index is obtained by calculating the weighted reciprocal average of accuracy, precision and recall rate.
6. The method according to claim 1, wherein, In step (2), the unsupervised clustering algorithm is a K-means clustering algorithm.
7. The method according to claim 1, wherein, In step (3), the random forest algorithm uses the average impurity reduction method based on the Gini index to calculate the feature importance score of each environmental physical quantity, and screens the key feature variables according to the score.
8. The method according to claim 1, wherein, In step (3), the environmental physical quantities include one or more of low-level vertical wind shear, convective available potential energy, frontogenesis function, distance between the gale site and the nearest precipitation area, precipitation area and maximum rain intensity, ground dew point temperature difference, pressure change and temperature change.
9. The method according to claim 1, wherein, In step (3), the final gale weather system classification result includes cold air activity type, typhoon type and convective system type.
Citation Information
Patent Citations
Forecasting wind speed correction method and system based on adaptive clustering, and storage medium
CN117436028A