Clustering calculation method for determining gust coefficient in strong wind based on convective environment parameters

Through the clustering calculation method based on convection environment parameters, combined with dynamic weight adjustment and improved OPTICS clustering algorithm, the prediction error and data sparsity of traditional gust coefficient determination methods in strong convective weather is solved, and high-precision and high-adaptive gust coefficient prediction are achieved.

CN120123796AActive Publication Date: 2025-06-10HUAFENG METEOROLOGICAL MEDIA GRP LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510607749.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-10
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The traditional method of determining gust coefficient ignores the dynamic influence of convective environment parameters on strong winds and gusts, resulting in a significant increase in prediction error in strong convective weather, and poor data sparsity and algorithm adaptability.

Method used

A clustering calculation method is proposed to determine the gust coefficient during strong winds based on convection environment parameters. By collecting multiple convection environment parameters, combining historical weight dynamic adjustment, the improved OPTICS clustering algorithm and dynamic weight adjustment function are used to perform refined data classification and real-time prediction.

Benefits of technology

It significantly improves the calculation accuracy, effectively captures the nonlinear change characteristics of the gust coefficient, reduces prediction deviation, improves the adaptability and response speed of the model, and meets the timeliness of disaster prevention and early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123796A_ABST
    Figure CN120123796A_ABST
Patent Text Reader

Abstract

The invention relates to a clustering calculation method for determining a gust coefficient in strong wind based on convective environment parameters, which comprises the following steps of: collecting observation data of a target station, and calculating the convective environment parameters; the average wind direction is divided into eight orientation grades, the wind speed is divided into five grades, the four convection environment parameters are subjected to grading assignment according to the numerical value range, the total weight score is calculated by combining the weight proportion of all the parameters in historical convection events, and five convection grades are divided according to the total score; dividing the data into 200 classes based on the wind speed, the wind direction and the convection level, calculating the average value of gust coefficients class by class, and constructing a gust coefficient library; an improved clustering algorithm is adopted, a distance measurement function of dynamic weight adjustment is combined, real-time meteorological data is mapped to a feature space, adjacent clusters are screened, and gust coefficients are fused. According to the method, through multi-flow parameter comprehensive grading and dynamic clustering optimization, the calculation error of the gust coefficient in strong wind is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of electronic digital data processing, and in particular relates to a clustering calculation method for determining a gust coefficient in strong winds based on convective environmental parameters. Background Art

[0002] Traditional methods for determining gust coefficients are mostly based on static statistical models or empirical lookup tables. For example, a linear relationship between the average gust coefficient and the average wind speed is established through historical wind speed data. However, such methods have significant limitations: the environmental parameters are simplistic. Existing methods usually only consider wind speed and wind direction, while ignoring the dynamic effects of convective environmental parameters on strong wind gusts, such as convective effective potential energy and vertical wind shear; the intensity of convective activity is closely related to the suddenness of gusts, but traditional models fail to incorporate such key parameters into the calculation system, resulting in a significant increase in prediction errors in severe convective weather; the classification granularity is insufficient. Traditional methods are relatively rough in the classification of wind speed and wind direction, such as only dividing into 4 wind direction grades, and not combining convective intensity levels for refined classification, especially in strong winds. During different periods of time, the gust coefficients under different convection levels vary greatly, and static classification is difficult to adapt to changes in complex meteorological conditions; there is a data sparsity problem. For rare meteorological combinations, such as strong convection accompanied by specific wind directions and high wind speeds, the historical sample size is insufficient, resulting in unstable prediction results of traditional lookup tables or regression models in such scenarios, and even serious deviations; the algorithm has poor adaptability. The existing clustering methods perform poorly in dealing with the temporal periodicity and spatial heterogeneity of meteorological data, and no dynamic weight adjustment mechanism is introduced, which cannot effectively capture the nonlinear characteristics of the gust coefficient changing with environmental parameters.

[0003] In response to the above problems, some studies have tried to improve the prediction accuracy by introducing machine learning models, but such methods rely on a large amount of labeled data and have high computational complexity, making them difficult to apply in real time. In addition, there is no solution in the existing technology that comprehensively classifies and weights multiple convective environmental parameters and combines them with improved clustering algorithms to dynamically optimize predictions. Summary of the invention

[0004] The present invention proposes a clustering calculation method for determining a gust coefficient in strong winds based on convective environment parameters, which comprises the following steps: Collect observation data of the target station, including 10-minute average wind speed, wind direction and the maximum 3-second gust data of the corresponding period, and obtain meteorological reanalysis data to calculate convective effective potential energy (CAPE), Sandia index (SI), temperature dew point difference and vertical wind shear; The wind direction is divided into 8 azimuth levels, the wind speed is divided into 5 levels, and the 4 convective environmental parameters are assigned values ​​according to the numerical range. The weighted total score is calculated based on the weight ratio of each parameter in historical convective events, and the convective levels are divided into 5 levels according to the total score. Divide the data into 200 categories based on wind speed, wind direction, and convection level, calculate the average value of the gust factor for each category, and construct a gust factor library. Adopt a clustering algorithm, combine a distance metric function with dynamic weight adjustment, map real-time meteorological data to the feature space, screen neighboring clusters, and fuse the benchmark gust factors of neighboring clusters to predict the gust wind speed at the target site.

[0005] Specifically, the distance metric function is: , where D is the wind direction, V is the wind speed, Y is the convection level, and ω D , ω V , ω Y are the dynamic weights of wind direction, wind speed, and convection level respectively.

[0006] Specifically, the dynamic weights are as follows: where ω D , ω V , ω Y are the dynamic weights of wind direction, wind speed, and convection level respectively, t is the UTC time, and V avg is the sliding 24-hour average wind speed.

[0007] Specifically, the clustering algorithm is the improved OPTICS algorithm.

[0008] Specifically, for the sample point i, its core distance , where Q3 represents the third quartile and IQR represents the interquartile range. The reachable distance from sample point i to j is .

[0009] Specifically, the benchmark gust factor of the neighboring cluster is as follows: ; where v k is the maximum 3-second gust of the kth sample in the cluster Cm, V k is the corresponding 10-minute average wind speed, Y m is the convection level of the cluster center, I(·) is the indicator function, taking 1 when the condition holds and 0 otherwise, and |Cm| is the number of samples contained in the cluster Cm.

[0010] Specifically, screening neighboring clusters and fusing the benchmark gust factors of neighboring clusters are as follows: map real-time data to the feature space; calculate the similarity between the mapped data and each cluster, and screen neighboring clusters with large similarities; generate the final value by weighted fusion of the gust factors of neighboring clusters and introducing a convection level correction term.

[0011] Specifically, the similarity calculation formula is: ; Among them, ω D , ω V , ω Y are the dynamic weights of wind direction, wind speed, and convection level respectively, and ΔD new,m , ΔV new,m , ΔY new,m are the distances between the real-time wind direction, wind speed, convection level and the m-th cluster.

[0012] Specifically, the four convection environment parameters are divided into grades and assigned values according to the numerical range. Specifically: the convective available potential energy CAPE is divided into five grades of [0, 500), [500, 1000), [1000, 2500), [2500, 4000), ≥4000 joules per kilogram, and the corresponding assigned values are 1 to 5; the Showalter index SI is divided into five grades of [0, ∞), [-3, 0), [-6, -3), [-9, -6), (-∞, -9), and the corresponding assigned values are 1 to 5; the temperature dew point difference is divided into five grades of ≥10, [6, 10), [4, 6), [2, 4), [0, 2) °C, and the corresponding assigned values are 1 to 5; the vertical wind shear is divided into five grades, and the corresponding assigned values are 1 to 5.

[0013] Beneficial technical effects: By comprehensively grading the four environmental parameters of convective available potential energy (CAPE), Showalter index (SI), temperature dew point difference, and vertical wind shear, and dynamically adjusting in combination with historical weights, the calculation accuracy is significantly improved; the improved clustering algorithm introduces a time periodicity factor and a dynamic distance metric, effectively capturing the non-linear change characteristics of the gust factor. Especially when the wind direction suddenly changes or the convection rapidly increases, the prediction stability is improved; through the probabilistic coefficient fusion strategy and the adjacent cluster screening mechanism, the data sparsity problem of rare meteorological combinations (such as severe convection + specific wind direction + high wind speed) is solved, and the prediction deviation is reduced compared with the traditional look-up table method; the combination of the dynamic weight adjustment function and the convection level correction term enables the model to adaptively optimize according to real-time meteorological data, and the response speed is improved when the wind speed fluctuates or the convection intensity changes, meeting the timeliness requirements of disaster prevention and warning; through 200 categories of refined data classification and incremental clustering algorithm, while ensuring the prediction accuracy, the computational complexity is reduced compared with traditional machine learning models. Description of the Drawings

[0014] The present invention will be further described below with reference to the accompanying drawings.

[0015] Figure 1 is a flowchart of the clustering calculation method for determining the gust factor during strong winds based on convection environment parameters; Figure 2 is a schematic diagram of the root mean square error of the gust wind speed determined for the full wind speed sample using the method of the present invention; Figure 3Schematic diagram of the normalized root mean square error of gust wind speed determined for the full wind speed sample using the method of the present invention; Figure 4 Schematic diagram of the root mean square error of gust wind speed determined for the strong wind speed sample using the method of the present invention; Figure 5 Schematic diagram of the normalized root mean square error of gust wind speed determined for the strong wind speed sample using the method of the present invention. Detailed implementation manners

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.

[0017] As Figure 1 shown, the present invention discloses a clustering calculation method for gust coefficients in strong winds based on convective environmental parameters, specifically as follows: I. Collect the observation data at the point where the gust coefficient needs to be calculated, including: 10-minute average wind speed and direction, and the maximum 3-second gust within 10 minutes; divide the wind direction into 8 equal azimuths, namely 0°, 45°, 90°, 135°, 180°, 225°, 270°, 315°, determine the belonging azimuth of the wind direction according to the closest principle, and divide the 10-minute average wind speed into 5 wind speed ranges according to (0, 5.4), [5.4, 7.9), [7.9, 10.8), [10.8, 13.8), ≥13.8; the longer the time period for collecting data, the better, but at least a complete 1-year time period should be satisfied.

[0018] II. Based on the reanalysis data, calculate the convective environmental parameters at the point, including convective available potential energy CAPE, Showalter index SI, temperature dew point difference, and vertical wind shear; divide the convective available potential energy CAPE, Showalter index SI, temperature dew point difference, and vertical wind shear into 5 ranges according to their numerical values, assign values of 1, 2, 3, 4, 5, add the weights assigned to the 4 environmental parameters respectively to obtain the total score, and divide it into 5 ranges according to the total score, that is, corresponding to 5 different convective levels.

[0019] Collect the three-dimensional reanalysis grid meteorological data of the site that can include, and calculate the convective available potential energy CAPE, Showalter index SI, temperature dew point difference, and vertical wind shear. The calculation formula of the convective available potential energy CAPE is as follows: ; where g is the acceleration due to gravity, z LFC is the free convection height, z EL is the equilibrium height, Tvparcel is the virtual temperature of the air parcel, T venv is the virtual temperature of the environment. The calculation formula of Convective Available Potential Energy (CAPE) can be simplified as: ; The Showalter Index (SI) is T 500 -T s , where T 500 is the environmental temperature on the 500 hPa isobaric surface (unit: °C), and T s is the temperature (unit: °C) of the air parcel when it rises along the dry adiabat from 850 hPa, reaches the condensation level and then rises along the wet adiabat to 500 hPa.

[0020] The temperature dew point depression is T - Td, where T is the air temperature (unit: °C) and Td is the dew point temperature (unit: °C).

[0021] The vertical wind shear is the ratio of the difference in wind speed (unit: m / s) in the vertical direction to the corresponding height difference (unit: m). Generally, the vertical wind shear is calculated from the lower troposphere (such as the 850 hPa height level, approximately equivalent to about 1500 m) to the middle troposphere (such as the 500 hPa height level, approximately equivalent to about 5500 m).

[0022] The classification and assignment of convective environmental parameters are as follows: The Convective Available Potential Energy (CAPE) is divided into five grades: [0, 500), [500, 1000), [1000, 2500), [2500, 4000), ≥4000 J / kg, and the corresponding assignments are 1 to 5; the Showalter Index (SI) is divided into five grades: [0, ∞), [-3, 0), [-6, -3), [-9, -6), (-∞, -9), and the corresponding assignments are 1 to 5; the temperature dew point depression is divided into five grades: ≥10, [6, 10), [4, 6), [2, 4), [0, 2) °C, and the corresponding assignments are 1 to 5; the vertical wind shear is divided into five grades, and the corresponding assignments are 1 to 5.

[0023] The Convective Available Potential Energy (CAPE), the Showalter Index (SI), the temperature dew point depression, and the vertical wind shear are numerically divided into 5 grades and assigned values of 1, 2, 3, 4, and 5 respectively. The four environmental parameters are weighted and added to obtain the total score. According to the total score, it is divided into 5 grades, that is, corresponding to 5 different convective grades. The value-taking method of the weights of the four convective environmental parameters is calculated according to the proportion of the respective assignments of these four environmental parameters in the total sum in the historical convective events that have occurred. After rounding the total weight score to the nearest integer, the values are 1, 2, 3, 4, and 5 respectively, and the convective grades at the station are divided into 5 grades.

[0024] III. According to the overall wind speed sequence, the wind direction is divided into 8 grades, the wind speed is divided into 5 grades, and combined with 5 different convection levels, a total of 200 categories are obtained. The ratio of the maximum 3-second wind speed to the 10-minute average wind speed is calculated for each moment, that is, the gust factor at each moment is obtained, and then the gust factors corresponding to each category are statistically averaged. In actual forecasting, for the average wind speed, average wind direction, and 4 convection environment parameter data predicted at a certain moment at the obtained station, according to this information, it is mapped to 200 groups of gust factor data, and the clustering analysis method is used to obtain the possible gust factor at this moment. The average wind speed is multiplied by the gust factor to obtain the predicted gust wind speed result at this station at this moment. Specifically: 1. Feature preprocessing Definition of input features: Wind direction D: Discretized into 8 directions (0°, 45°,..., 315°), encoded as D ∈ {1, 2,..., 8}; Wind speed V: 5-level standardized value, encoded as V ∈ {1, 2, 3, 4, 5}; Convection level Y: The result of rounding the total weighted score, Y ∈ {1, 2, 3, 4, 5}.

[0025] 2. Improved distance metric Define the mixed distance function: ; where: ω D 、ω V 、ω Y are the dynamic weights of wind direction, wind speed, and convection level respectively, , that is, the normalized circular distance, .

[0026] Dynamic weight adjustment: ; where t is the current UTC time (hour), used to introduce the diurnal cycle change, and V avg is the historical average wind speed of the current period (sliding 24-hour window). A time periodic factor is introduced into the wind direction weight to reflect the diurnal variation law of the wind direction, such as the sea-land breeze effect, and enhance the model's ability to capture the differences in the wind direction patterns between day and night. For the wind speed weight, through the exponential decay function, considering that the gust factor tends to be stable at high wind speeds, the influence of the wind speed difference in high wind speed periods on classification is reduced. The setting of the convection level weight can strengthen the correction effect on classification when the convection intensity changes suddenly.

[0027] 3. Density-based clustering algorithm Adopt the density-based clustering algorithm, specifically the improved OPTICS algorithm, to process the data. The specific steps are as follows: (1) Definition of core object: If the number of samples within the neighborhood (with the core distance as the radius) of a sample point exceeds the dynamic threshold, then this sample is determined as a core object.

[0028] (2) Calculation of sample point distance: For sample point i, its core distance ϵ core is defined as the minimum neighborhood radius for sample point i to become a core object. Specifically, where Q3 represents the third quartile, that is, the value at the 75% position after arranging the historical sample distance values in ascending order, and IQR represents the interquartile range, which is the difference between Q3 and the first quartile Q1. When determining the core neighborhood radius, take the typical distance value of most samples (75%) and add half of the discrete range. The reachable distance from sample point i to j is . In the calculation of sample point distance, through adaptive adjustment, the distance calculation can be automatically optimized considering the historical data distribution.

[0029] (3) Cluster generation: Through density reachability, that is, starting from the core object, connecting adjacent samples according to the reachable distance, dividing the historical data into several clusters, so that the samples within the same cluster have similar combinations of wind direction, wind speed, and convection level.

[0030] The clustering algorithm of the present invention can overcome the dependence of traditional clustering algorithms on spherical clusters, be applicable to the spatial heterogeneity of meteorological data, automatically filter samples in sparse areas, improve the stability of the model, and moreover, when new data appears, only local adjustment of the cluster structure is required, avoiding global reclustering and reducing the computational complexity.

[0031] 4. Calculation of benchmark gust factor For each cluster Cm, calculate the benchmark gust factor α of each cluster m : ; where v k is the maximum 3 - second gust of the k - th sample in the cluster Cm to be calculated, V k is the corresponding 10 - minute average wind speed, Y m is the convection level of the cluster center, I(·) is the indicator function, taking 1 when the condition holds and 0 otherwise, used to strengthen the influence of strong convection, |C m | is the number of samples included in the cluster C m .

[0032] 5. Real - time prediction (1) Feature mapping: Map the real - time data (D new , V new , Y new ) to the feature space.

[0033] (2) Similarity calculation: ; Among them, ΔD new,m is the distance between the real-time wind direction data D new and the wind direction D m of the m-th cluster. , ΔV new,m is the distance between the real-time wind speed level and the wind speed level V m of the m-th cluster. , ΔY new,m is the distance between the real-time convection level and the convection level Y m of the m-th cluster. . D m , V m , Y m can be obtained by taking the mode (the level with the highest occurrence frequency) or weighted average (calculated according to the sample time weight). For example, D m takes the circular average of the sample wind directions in the cluster. If the sample wind directions in a cluster are [1, 1, 2], then D m ≈1.3.

[0034] (3) Proximity cluster screening: Screen clusters with high similarity. Specifically, clusters that satisfy and can be selected.

[0035] (4) Gust factor fusion, specifically: , where s m is the similarity between the real-time data and the m-th cluster, α m is the benchmark gust factor of the m-th cluster, and Y new is the real-time convection level. Through weighted fusion, the influence of clusters with high similarity is strengthened, and the low-similarity noise is suppressed. By introducing a convection correction term, the coefficient is dynamically adjusted, so that the higher the convection level, the greater the increase in the gust factor, and the prediction accuracy in strong convection scenarios is improved.

[0036] Through dynamic clustering, the data sparsity problem of small sample categories (such as the combination of strong convection + northwest wind + high wind speed) can be solved, and the prediction stability of sudden weather (such as the rapid change of wind direction during a squall line passage) can be improved. Compared with the static look-up table method, the prediction error of the gust factor during strong convection can be greatly reduced.

[0037] Using the 10-minute observation data sequence of national meteorological stations from January 1, 2022 to December 31, 2023, the gust factors at all observation sites are obtained by the above method; and the 10-minute observation data sequence from January 1 to October 31, 2024 is used to verify and test this method.

[0038] Statistically analyze the errors during the full wind speed and strong wind speed periods respectively. The full wind speed refers to all gust observation samples, and the strong wind speed refers to the samples when the gust observation value is greater than 17.2 m / s. Statistically analyze their root mean square error (RMSE) and normalized root mean square error (NMSE) respectively, where the vertical axis is the number of stations and the horizontal axis is RMSE (m / s) or NMSE (%) respectively. As Figure 2-5 shown, during the full wind speed period, the RMSE mainly concentrates between 1.2 and 2.5 m / s, and the normalized root mean square error mainly concentrates between 27% and 60%; during the strong wind speed period, the RMSE mainly concentrates between 4.0 and 9.5 m / s, and the normalized root mean square error mainly concentrates between 25% and 50%. It can be seen that the newly established gust model has good performance when used for calculating gust wind speeds, especially for strong wind speeds.

Claims

1. A clustering calculation method for determining gust coefficients in strong winds based on convective environment parameters, characterized in that: The following steps are involved: Collect observation data of the target station, including 10-minute average wind speed, wind direction and the maximum 3-second gust data of the corresponding period, and obtain meteorological reanalysis data to calculate convective effective potential energy (CAPE), Sandia index (SI), temperature dew point difference and vertical wind shear; The wind direction is divided into 8 azimuth levels, the wind speed is divided into 5 levels, and the 4 convective environmental parameters are assigned values ​​according to the numerical range. The weighted total score is calculated based on the weight ratio of each parameter in historical convective events, and the convective levels are divided into 5 levels according to the total score. The data is divided into 200 categories based on wind speed, wind direction and convection level, and the average value of gust coefficient is calculated for each category to build a gust coefficient library; A clustering algorithm is used in combination with a distance measurement function with dynamic weight adjustment to map real-time meteorological data into the feature space, screen neighboring clusters and fuse the benchmark gust coefficients of neighboring clusters to predict the gust wind speed at the target site.

2. The method according to claim 1, characterized in that The distance metric function is: , where D is wind direction, V is wind speed, Y is convection level, ω D ,ω V ,ω Y They are the dynamic weights of wind direction, wind speed and convection level respectively.

3. The method according to claim 1, characterized in that The dynamic weights are as follows: Among them, ω D ,ω V ,ω Y are the dynamic weights of wind direction, wind speed, and convection level, t is UTC time, V avg It is the sliding 24-hour average wind speed.

4. The method according to claim 1, characterized in that: The clustering algorithm is the improved OPTICS algorithm.

5. The method according to claim 4, characterized in that For sample point i, its core distance , where Q3 represents the third quartile, IQR represents the interquartile range, and the accessible distance from sample point i to j is .

6. The method according to claim 1, characterized in that The baseline gust coefficient for the neighboring cluster is: Among them, v k is the maximum 3-second gust of the k-th sample in cluster Cm, V k is the 10-minute average wind speed, Y m is the convection level of the cluster center, I(·) is the indicative function, which takes 1 when the condition is met and 0 otherwise, and |Cm| is the number of samples contained in cluster Cm.

7. The method according to claim 1, characterized in that Filter neighboring clusters and fuse the benchmark gust coefficients of neighboring clusters. Specifically, the real-time data is mapped to the feature space; the similarity between the mapped data and each cluster is calculated, and the neighboring clusters with large similarity are filtered; the gust coefficients of neighboring clusters are weightedly fused, and the convection level correction term is introduced to generate the final value.

8. The method according to claim 7, characterized in that The similarity calculation formula is: Among them, ω D ,ω V ,ω Y are the dynamic weights of wind direction, wind speed, and convection level, ΔD new,m , ΔV new,m , ΔY new,m is the distance between the real-time wind direction, wind speed, convection level and the mth cluster.

9. The method according to claim 1, characterized in that: The four convective environmental parameters are assigned values ​​according to the numerical range, specifically: the convective effective potential energy CAPE is divided into five levels: [0,500), [500,1000), [1000,2500), [2500,4000), ≥4000 J / kg, and the corresponding values ​​are 1 to 5; the Sandy index SI is divided into five levels: [0,∞), [-3,0), [-6,-3), [-9,-6), (-∞,-9), and the corresponding values ​​are 1 to 5; the temperature dew point difference is divided into five levels: ≥10, [6,10), [4,6), [2,4), [0,2)℃, and the corresponding values ​​are 1 to 5; the vertical wind shear is divided into five levels: There are five levels, with corresponding values ​​ranging from 1 to 5.

Citation Information

Patent Citations

  • Method and system for judging strong convection gust

    CN111783821A

  • Troposphere low-layer gust calculation method based on turbulence intensity and average wind speed

    CN115455368A

  • High-dimensional density peak value clustering method based on improved equidistant mapping

    CN116304768A

  • Photovoltaic output prediction method and device, computer equipment, readable storage medium and program product

    CN118607724A

  • Ground wind speed measurement and calculation correction method considering convective complex environment characteristics

    CN119538099A