Method for accurate recognition and tracking of thunderstorm clouds based on machine learning

CN122815575APending Publication Date: 2026-09-25内蒙古自治区气候中心(内蒙古自治区气候变化中心内蒙古自治区雷电防护中心)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611008826.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

采用固定阈值进行一刀切式聚类,极易导致夏季广域雷暴被过度分割为多个孤立小簇或春季密集弱雷暴被错误合并,严重降低了雷暴云识别的准确率和边界重构的保真度

Benefits of technology

本发明通过历史多普勒天气雷达组合反射率数据测算历史每场雷暴的最优ST-DBSCAN参数,进而确定训练标签。众所周知,基于ST-DBSCAN参数的ST-DBSCAN聚类算法中,ST-DBSCAN参数包括空间邻域阈值和时间邻域阈值,其中,空间邻域阈值用于判定闪电点之间的空间邻近性,时间邻域阈值用于判定闪电点之间的时间邻近性,二者共同作为ST-DBSCAN聚类引擎的核心输入,对离散的闪电数据进行时空密度聚类。本发明以ST-DBSCAN参数为标签,引入热力学特征和动力学特征作为第一环境特征进行模型训练,使得ST-DBSCAN参数与环境特征之间的映射关系更加合理、准确。在模型的推理过程中,实时输入测得的第一环境特征数据,即可实时输出动态变化的动态空间邻域阈值和动态时间邻域阈值,并据此进行时空密度聚类提取出同源闪电簇,从而精准识别雷暴单体(雷暴云),重构雷暴云轮廓,根据轮廓测算雷暴云在当前生命史中连续时间窗口的质心位置,从而得到雷暴云的移动轨迹,实现雷暴云的精准识别与追踪。

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The application discloses a thunderstorm cloud accurate identification and tracking method based on machine learning, and belongs to the cross technical field of meteorological disaster monitoring and early warning and spatiotemporal data mining. The method divides a target monitoring area into spatial grids, extracts a first environmental feature before a historical lightning event in each grid to construct a feature vector set; calculates optimal ST-DBSCAN parameters as a training label based on historical radar combined reflectivity data, and trains a machine learning model in combination with the vector set; collects the first environmental feature in real time to input the model, infers dynamic spatial and temporal neighborhood threshold values of each grid, and accordingly performs adaptive ST-DBSCAN spatiotemporal density clustering on discrete lightning data, extracts homologous lightning clusters, and thereby reconstructs a thunderstorm cloud spatial profile and outputs a moving track. Compared with the prior art, the application realizes adaptive adjustment of clustering parameters on meteorological environment, effectively avoids over-segmentation or false merging of thunderstorm clouds, and improves the accuracy of lightning attribution and tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of meteorological disaster monitoring and early warning, spatiotemporal data mining and machine learning, and more specifically, to a method for accurate identification and tracking of thunderstorm clouds based on machine learning. Background Technology

[0002] Accurate identification of thunderstorm clouds and aggregation of discrete lightning location data, also known as lightning clustering, are the core foundation for severe convective weather early warning and lightning disaster prevention. Traditional thunderstorm cloud identification often employs the standard ST-DBSCAN spatiotemporal density clustering algorithm. However, existing methods suffer from the following significant technical shortcomings: (1) Clustering parameters are fixed and lack environmental adaptability: When performing spatiotemporal clustering, the spatial threshold (Eps1, such as 20 km) and time threshold (Eps2, such as 30 minutes) of the traditional ST-DBSCAN algorithm are often globally fixed. Moreover, the parameters are not adaptive and often need to be manually adjusted for different regions.

[0003] (2) Clustering failure due to differences in thunderstorm scale: In reality, the development speed and diffusion range of thunderstorms vary greatly in different seasons and under different meteorological backgrounds. For example, weak thunderstorms in spring and strong convective thunderstorms in summer are completely different in terms of spatiotemporal scale. Using a fixed threshold for one-size-fits-all clustering can easily lead to the over-segmentation of wide-area thunderstorms in summer into multiple isolated small clusters or the incorrect merging of dense weak thunderstorms in spring, which seriously reduces the accuracy of thunderstorm cloud identification and the fidelity of boundary reconstruction.

[0004] In view of the above, the present invention is proposed to solve the aforementioned technical problems. Summary of the Invention The purpose of this invention is to provide a machine learning-based method for accurate identification and tracking of thunderstorm clouds, in order to solve the aforementioned technical problems.

[0005] To achieve the above objectives, the present invention provides the following technical solution: A machine learning-based method for accurate identification and tracking of thunderstorm clouds includes the following steps: S100 divides the target monitoring area into several spatial grids according to latitude and longitude. Based on discrete lightning location data, it collects and extracts the first environmental features within the first preset time before the occurrence of historical lightning events in each grid, and constructs a feature vector set. The first environmental features include thermodynamic features and dynamic features. S200, construct and train a machine learning model, where: S210: After constructing the machine learning model, the optimal ST-DBSCAN parameters for each historical thunderstorm under radar boundary constraints are calculated based on historical Doppler weather radar combined reflectivity data. The obtained optimal ST-DBSCAN parameters are used as training labels for the machine learning model and combined with the feature vector set of S100 for model training. S220 collects the first environmental features of each grid in real time and inputs them into a pre-trained machine learning model for forward inference, and outputs the dynamic spatial neighborhood threshold and dynamic temporal neighborhood threshold of each grid in real time. The S300 inputs the dynamic spatial neighborhood threshold and dynamic temporal neighborhood threshold output in real time from the machine learning model into the ST-DBSCAN clustering engine to perform ST-DBSCAN dynamic spatiotemporal density clustering on the lightning data of all collected grids and extract homogeneous lightning clusters in real time. S400, based on the homogeneous lightning cluster data generated by S300, extracts and outputs the current life cycle duration of the thunderstorm cloud based on the occurrence time of the first and last ground lightning within the cluster, divides its existing life cycle into several time windows, extracts the spatial cluster boundary based on the spatial coordinate set of the lightning within each time window based on the homogeneous lightning cluster data within each time window, reconstructs the spatial outline boundary of the thunderstorm cloud within each time window, and thus extracts and outputs the movement trajectory of the thunderstorm cloud.

[0006] Optionally, in S210, historical Doppler weather radar combined reflectivity data is selected as the baseline fact of the true physical boundary of thunderstorm clouds, and the optimal ST-DBSCAN parameters of each historical thunderstorm under radar boundary constraints are derived in reverse through a heuristic search algorithm.

[0007] Optionally, in step S400, based on the homogeneous lightning cluster data generated in step S300, the current lifespan duration of the thunderstorm cloud is extracted and output based on the occurrence time of the first and last ground lightning within the cluster. Its existing lifespan is divided into several time windows. Based on the homogeneous lightning cluster data within each time window, spatial clustering boundaries are extracted based on the spatial coordinate set of lightning within the cluster. The spatial outline boundary of the thunderstorm cloud within each time window is reconstructed and its coverage area is calculated and output. The movement trajectory of the thunderstorm cloud is extracted and output.

[0008] Optionally, the S400 also synthesizes and outputs the life history profile and disaster-causing characteristics of thunderstorm clouds.

[0009] Optionally, S100, the first preset duration is 0.5-2 hours.

[0010] Optionally, the first environmental feature may also include underlying surface features.

[0011] Optionally, the underlying surface features of a grid include the topographic slope and elevation of the surface location corresponding to that grid.

[0012] Optionally, thermodynamic characteristics include ambient temperature, relative humidity, convective available potential energy, and K exponent, while kinetic characteristics include vertical wind shear and water vapor flux divergence.

[0013] Optionally, S400, based on the data of the same lightning clusters within each time window, and based on the set of spatial coordinates of the lightning within the cluster, constructs a two-dimensional convex hull polygon using the convex hull algorithm, which is used as the spatial outline boundary of the reconstructed thunderstorm cloud within the time window and outputs it, and calculates its coverage area and movement trajectory accordingly and outputs them.

[0014] Optionally, in step S400, based on the data of homogeneous lightning clusters within each time window, and using the set of spatial coordinates of lightning within the cluster, an Alpha Shape polygon is constructed using the Alpha Shape algorithm. This polygon serves as the spatial outline boundary of the reconstructed thunderstorm cloud within that time window, and its coverage area and movement trajectory are calculated and output accordingly.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention calculates the optimal ST-DBSCAN parameters for each historical thunderstorm using historical Doppler weather radar combined reflectivity data, thereby determining the training labels. As is well known, in the ST-DBSCAN clustering algorithm based on ST-DBSCAN parameters, the ST-DBSCAN parameters include spatial neighborhood thresholds and temporal neighborhood thresholds. The spatial neighborhood threshold is used to determine the spatial proximity between lightning points, and the temporal neighborhood threshold is used to determine the temporal proximity between lightning points. Both serve as the core input of the ST-DBSCAN clustering engine, performing spatiotemporal density clustering on discrete lightning data. This invention uses ST-DBSCAN parameters as labels and introduces thermodynamic and kinetic features as the first environmental features for model training, making the mapping relationship between ST-DBSCAN parameters and environmental features more reasonable and accurate. During the model's inference process, the measured first environmental feature data is input in real time, and the dynamically changing dynamic spatial neighborhood threshold and dynamic temporal neighborhood threshold are output in real time. Based on this, spatiotemporal density clustering is performed to extract homogeneous lightning clusters, thereby accurately identifying thunderstorm cells (thunderstorm clouds), reconstructing the thunderstorm cloud outline, and calculating the centroid position of the thunderstorm cloud in the continuous time window of its current life history based on the outline, thus obtaining the movement trajectory of the thunderstorm cloud and achieving accurate identification and tracking of thunderstorm clouds.

[0016] This invention uses training labels determined by optimal ST-DBSCAN parameters under radar reflectivity data constraints, a mapping relationship between historical first environmental features and labels, and dynamic spatial and temporal neighborhood thresholds output from real-time first environmental features. It incorporates real-time environmental conditions as a criterion in the lightning homing process, thus effectively solving the technical shortcomings of traditional ST-DBSCAN algorithms, such as fixed clustering parameters and lack of environmental adaptability. Specifically, this invention eliminates the need for manual setting of clustering parameters. Under different meteorological conditions, the model can autonomously adjust the spatial and temporal neighborhood thresholds based on historical training results and the current environmental state, demonstrating adaptability to thunderstorm processes in different regions, seasons, and weather backgrounds, without requiring tedious manual parameter tuning for different regions.

[0017] Based on this, the model can learn the true physical scale patterns of thunderstorms of different seasons and intensities under radar boundary constraints. This allows the ST-DBSCAN clustering engine to generate adaptive neighborhood thresholds that match the spatiotemporal scale of both weak spring thunderstorms and strong summer convective thunderstorms. This effectively avoids the situation where wide-area summer thunderstorms are over-segmented into multiple isolated clusters due to excessively small fixed thresholds, or dense weak spring thunderstorms are incorrectly merged due to excessively large fixed thresholds. This significantly improves the accuracy of thunderstorm cloud identification and the physical consistency of clustering results. Furthermore, after extracting homologous lightning clusters, this invention extracts the life history by the time of the first and last ground flashes, divides the life cycle into several time windows, and extracts spatial clustering boundaries based on spatial coordinate sets within each time window. This reconstructs a fine spatial contour boundary of the thunderstorm cloud and calculates the centroid to obtain the movement trajectory. This further ensures that regardless of the size of the individual thunderstorm or the length of its life cycle, high-fidelity contour reconstruction and continuous trajectory tracking can be achieved. This reduces the clustering distortion caused by improper fixed threshold settings and improves the reliability and accuracy of lightning homing. Detailed Implementation

[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0019] This invention provides a method for accurate identification and tracking of thunderstorm clouds based on machine learning. It belongs to the interdisciplinary fields of meteorological disaster monitoring and early warning, spatiotemporal data mining, and machine learning. The method uses machine learning to dynamically predict clustering parameters based on multidimensional environmental features, and then utilizes the adaptive ST-DBSCAN algorithm for accurate identification of thunderstorm clouds and intelligent lightning aggregation (lightning aggregation). Specifically, it includes the following steps: S100, a gridded spatiotemporal segmentation and multimodal environmental feature construction, aims to quantify the physical impact of different meteorological backgrounds on thunderstorm spread velocity. Specifically: This invention breaks through the limitations of traditional algorithms that extract data globally. It divides the target monitoring area into several high-resolution spatial grids of N×N kilometers based on latitude and longitude, with each grid corresponding to both longitude and latitude coordinates. Lightning data for each grid is collected independently. As is well known in the field of thunderstorm cloud lightning tracking, the aforementioned lightning data refers to the recorded data obtained by continuously monitoring lightning electrical parameters within this grid space (latitude and longitude coordinates) to help identify lightning events and trigger discrete lightning event recordings. This recorded data should at least include the start timestamp of the lightning event, waveform data, and lightning electrical parameters extracted from the waveform. Because multiple lightning events within a grid are discrete, the obtained lightning data is discrete data. This invention uses discrete lightning location data as a benchmark to collect and extract the first environmental features within a first preset time period before the occurrence of a historical lightning event (flashpoint event start timestamp) in each grid, constructing a feature vector set. The first environmental features include thermodynamic and dynamic features.

[0020] This invention uses thermodynamic characteristics to characterize the energy state and instability of the atmosphere containing thunderstorm clouds. These characteristics directly determine the discharge activity of thunderstorm clouds, i.e., the frequency density of lightning occurrences per unit time. Specifically, the higher the atmospheric energy and the stronger the convective instability, the more vigorous the thunderstorm development, the higher the frequency of lightning occurrences, and the shorter the time interval between adjacent lightning strikes. Conversely, when the atmospheric energy is low or the instability is insufficient, thunderstorm development is weak, lightning is sparse, and the time interval between adjacent lightning strikes is longer. Therefore, thermodynamic characteristics provide a direct physical basis for determining the temporal neighborhood threshold in the ST-DBSCAN clustering engine.

[0021] Dynamic characteristics are used to characterize the motion state and dynamic structure of the atmosphere containing thunderstorm clouds. These characteristics directly determine the motion morphology and spatial diffusion range of thunderstorm clouds. Specifically, when atmospheric dynamics are strong, thunderstorm clouds are subjected to shearing and transport effects from airflows at different altitudes, causing the cloud body to stretch, tilt, and even break apart. This significantly expands the spatial dispersion range of lightning generated by the same thunderstorm cell on the horizontal plane. Conversely, in environments with weaker dynamics, the cloud structure is compact, and lightning is concentrated within a smaller spatial area. Therefore, dynamic characteristics provide a direct physical basis for determining the spatial neighborhood threshold in the ST-DBSCAN clustering engine.

[0022] Therefore, the first environmental feature, composed of thermodynamic and kinetic characteristics, can characterize the physical state of thunderstorm clouds in both temporal and spatial dimensions. Together, they form the physical basis for all the external driving information required by the ST-DBSCAN clustering engine: thermodynamic characteristics determine the degree of lightning aggregation along the time axis, thus constraining the direction of the temporal neighborhood threshold; kinetic characteristics determine the degree of lightning dispersion along the spatial axis, thus constraining the direction of the spatial neighborhood threshold. Based on these two types of features, the machine learning model can determine the optimal ST-DBSCAN parameters for each historical thunderstorm under radar boundary constraints from historical Doppler weather radar combined reflectivity data. This establishes a complete mapping relationship from the environmental physical state to the optimal ST-DBSCAN parameters, thereby achieving adaptive adjustment of the clustering threshold to the meteorological background.

[0023] Optionally, lightning electrical parameters include, but are not limited to, current parameters (e.g., peak current amplitude, current steepness, amount of transferred charge, etc.), lightning type (cloud lightning or ground lightning), polarity parameters (positive ground lightning or negative ground lightning), waveform and structural parameters (e.g., waveform parameters, number of return strokes and return stroke interval, etc.), radiation field parameters (radiation field peak value, spectral characteristics), and derived and statistical parameters used to assess long-term risks and engineering design (e.g., ground lightning density, cumulative probability distribution of lightning current amplitude, etc.).

[0024] S200 constructs a closed-loop driven machine learning model and dynamically infers clustering parameters, where: S210: Construct and train a lightweight supervised machine learning model (such as LightGBM, XGBoost, or deep neural network DNN) with meteorological and physical significance. Then, proceed to the training phase of automatic labeling: Calculate the optimal ST-DBSCAN parameters of each historical thunderstorm under radar boundary constraints based on historical Doppler weather radar combined reflectivity data. Use the obtained optimal ST-DBSCAN parameters as the training label for the machine learning model and combine them with the feature vector set of S100 for model training.

[0025] It should be noted that the historical Doppler weather radar combined reflectivity data should have a time span sufficient to cover thunderstorm samples under different environmental conditions that need to be overcome. For example, if the target monitoring area is located in a mid-to-high latitude region with significant seasonal variations, and involves lightning distortion that may be caused by differences in thermal conditions between different seasons (spring and summer), i.e., weak thunderstorms in spring are over-segmented due to sparse lightning and strong thunderstorms in summer are incorrectly merged due to dense lightning, then the historical data should include historical thunderstorm sample data for both spring and summer in this area. If the target monitoring area is a region alternately affected by the edge of the subtropical high and the cold vortex system, and involves the lightning distortion that may be caused by the intensity difference between the thermal thunderstorms on the edge of the subtropical high and the strong convective thunderstorms under the background of the cold vortex, that is, the thermal thunderstorms on the edge of the subtropical high are over-segmented due to the sparse lightning and wide distribution range, and the strong thunderstorms under the background of the cold vortex are incorrectly merged due to the dense lightning, then the historical data should simultaneously include historical sample data of thermal thunderstorms on the edge of the subtropical high and strong thunderstorms under the cold vortex that occurred in this area at different times; If the target monitoring area is a region with fragmented terrain and variable weather systems (such as the downstream areas of rivers, which are affected by multiple weather systems including river systems, lake and land breezes, and cyclones entering the sea, resulting in significant differences in the lifespan of thunderstorms: in spring, due to the convergence of cold air and warm and humid air currents, thunderstorm activity is scattered and lasts for less than 1 hour, while in summer, due to the influence of the edge of the subtropical high and the periphery of tropical cyclones, it is easy to form highly organized squall lines or mesoscale convective systems with lifespans of several hours or even longer; or in areas where the piedmont areas of plateau mountains are interspersed with plains, due to the influence of local topographic forcing and afternoon thermal convection, thunderstorms are mostly short-lived single-point triggered and rapidly dissipating cells, but they will frequently encounter the passage of synoptic-scale systems, and thunderstorms can develop into long-lived multi-cell storms lasting for several hours under favorable environmental conditions), then the historical data should include historical sample data of both short-lived and long-lived thunderstorms to avoid short-lived thunderstorms being incorrectly extended due to excessively large time thresholds and long-lived thunderstorms being truncated due to excessively small time thresholds.

[0026] If the target monitoring area is located in a mid-latitude region affected by multiple weather systems, involving different directions of steering airflow, varying intensities of vertical wind shear, and different moisture transport conditions, the scale of thunderstorms in this region will vary significantly. When the region is affected by strong deep vertical wind shear and abundant warm and humid advection, thunderstorms are prone to organize into squall lines or mesoscale convective systems, with horizontal scales reaching tens to hundreds of kilometers. When the region is controlled by weak wind shear and local thermal convection, thunderstorms are mostly isolated entities with a horizontal scale of only a few kilometers. In addition, if the region is affected by the passage of different weather systems, such as frontal systems, low-vortex shear lines, warm and humid airflows at the edge of the subtropical high, and the outer circulation of tropical cyclones, the scale of thunderstorms in the region will also vary greatly from a few kilometers to hundreds of kilometers. In this case, historical data should include both small-scale and large-scale historical thunderstorm sample data to avoid small-scale thunderstorms being incorrectly merged due to excessively large spatial thresholds and large-scale systems being over-segmented due to excessively small spatial thresholds.

[0027] If the underlying surface type of the target monitoring area is diverse, involving the possible distortion of lightning return due to the modulation of the lightning triggering location and spatial distribution by different terrains (mountains and plains) and different surface covers (cities, water bodies, vegetation), that is, the difference between the abnormal spatial distribution of lightning caused by the local uplift effect of mountain thunderstorms and the uniform distribution caused by the flat terrain of plain thunderstorms, or the different effects of the urban heat island effect and the cooling effect of water bodies on the lightning movement path, then the historical data should simultaneously include historical sample data of thunderstorms under different underlying surface backgrounds, and the first environmental feature should include underlying surface features.

[0028] If multiple scenarios as described above exist simultaneously, the historical data should cover a combination of samples from all of these scenarios to ensure that the diversity of the training samples is sufficient to cover and improve the distortion problem.

[0029] S220, enter the inference stage of dynamic model output, collect the first environmental features of each grid in real time and input them into the pre-trained machine learning model for forward inference, and output the dynamic spatial neighborhood threshold (Eps1) and dynamic temporal neighborhood threshold (Eps2) of each grid in the specific time and space in real time. The S300 inputs the dynamic spatial neighborhood threshold and dynamic temporal neighborhood threshold output in real time from the machine learning model into the ST-DBSCAN clustering engine to perform ST-DBSCAN dynamic spatiotemporal density clustering on the lightning data of all collected grids and extract homogeneous lightning clusters in real time. S400, based on the homogeneous lightning cluster data generated by S300, extracts and outputs the current life cycle duration of the thunderstorm cloud based on the occurrence time of the first and last ground flashes within the cluster. It divides the existing life cycle into several time windows. Based on the homogeneous lightning cluster data within each time window, it can be determined that the homogeneous lightning cluster data includes all lightning data of homogeneous lightning within the target monitoring area, specifically including the latitude and longitude coordinates of the corresponding grid of homogeneous lightning (the grid for homogeneous lightning events), the start timestamp of the lightning event, waveform data, and lightning electrical parameters extracted from the waveform. Based on the spatial coordinate set of lightning within the cluster, spatial cluster boundary extraction is performed to reconstruct the spatial contour boundary of the thunderstorm cloud within each time window, thereby extracting and outputting the movement trajectory of the thunderstorm cloud.

[0030] Data within the last period of a time window with less than a full time window remaining can be disregarded. Alternatively, if the number of valid grids (grids for performing same-source lightning events) in the same lightning cluster data within that period is sufficient for spatial cluster boundary extraction and reconstruction of the spatial outline boundary of thunderstorm clouds within that period, then the data can be used, and the spatial outline boundary, coverage area, and movement trajectory of thunderstorm clouds can be extracted and output. You can design the data yourself according to actual needs and site conditions.

[0031] This invention calculates the optimal ST-DBSCAN parameters for each historical thunderstorm using historical Doppler weather radar combined reflectivity data, thereby determining the training labels. As is well known, in the ST-DBSCAN clustering algorithm based on ST-DBSCAN parameters, the ST-DBSCAN parameters include spatial neighborhood thresholds and temporal neighborhood thresholds. The spatial neighborhood threshold is used to determine the spatial proximity between lightning points, and the temporal neighborhood threshold is used to determine the temporal proximity between lightning points. Both serve as the core input of the ST-DBSCAN clustering engine, performing spatiotemporal density clustering on discrete lightning data. This invention uses ST-DBSCAN parameters as labels and introduces thermodynamic and kinetic features as the first environmental features for model training, making the mapping relationship between ST-DBSCAN parameters and environmental features more reasonable and accurate. During the model's inference process, the measured first environmental feature data is input in real time, and the dynamically changing dynamic spatial neighborhood threshold and dynamic temporal neighborhood threshold are output in real time. Based on this, spatiotemporal density clustering is performed to extract homogeneous lightning clusters, thereby accurately identifying thunderstorm cells (thunderstorm clouds), reconstructing the thunderstorm cloud outline, and calculating the centroid position of the thunderstorm cloud in the continuous time window of its current life history based on the outline, thus obtaining the movement trajectory of the thunderstorm cloud and achieving accurate identification and tracking of thunderstorm clouds.

[0032] This invention uses training labels determined by optimal ST-DBSCAN parameters under radar reflectivity data constraints, a mapping relationship between historical first environmental features and labels, and dynamic spatial and temporal neighborhood thresholds output from real-time first environmental features. It incorporates real-time environmental conditions as a criterion in the lightning homing process, thus effectively solving the technical shortcomings of traditional ST-DBSCAN algorithms, such as fixed clustering parameters and lack of environmental adaptability. Specifically, this invention eliminates the need for manual setting of clustering parameters. Under different meteorological conditions, the model can autonomously adjust the spatial and temporal neighborhood thresholds based on historical training results and the current environmental state, demonstrating adaptability to thunderstorm processes in different regions, seasons, and weather backgrounds, without requiring tedious manual parameter tuning for different regions.

[0033] Based on this, the model can learn the true physical scale patterns of thunderstorms of different seasons and intensities under radar boundary constraints. This allows the ST-DBSCAN clustering engine to generate adaptive neighborhood thresholds that match the spatiotemporal scale of both weak spring thunderstorms and strong summer convective thunderstorms. This effectively avoids the situation where wide-area summer thunderstorms are over-segmented into multiple isolated clusters due to excessively small fixed thresholds, or dense weak spring thunderstorms are incorrectly merged due to excessively large fixed thresholds. This significantly improves the accuracy of thunderstorm cloud identification and the physical consistency of clustering results. Furthermore, after extracting homologous lightning clusters, this invention extracts the life history by the time of the first and last ground flashes, divides the life cycle into several time windows, and extracts spatial clustering boundaries based on spatial coordinate sets within each time window. This reconstructs a fine spatial contour boundary of the thunderstorm cloud and calculates the centroid to obtain the movement trajectory. This further ensures that regardless of the size of the individual thunderstorm or the length of its life cycle, high-fidelity contour reconstruction and continuous trajectory tracking can be achieved. This reduces the clustering distortion caused by improper fixed threshold settings and improves the reliability and accuracy of lightning homing.

[0034] This invention is applicable to intelligent identification of thunderstorm clouds, automated collection of lightning strike events, and tracking of severe convective weather on a wide-area scale. It is particularly suitable for large-scale thunderstorm data mining and lightning-induced fire disaster tracing scenarios requiring data to span different seasons and climate zones. Compared to existing technologies, This invention breaks through the barrier of fixed parameters and introduces meteorological environmental characteristics (the first environmental characteristic of this application) into the spatiotemporal clustering process, solving the technical defects of fixed parameters in traditional spatial clustering algorithms and realizing dynamic adaptive clustering of different lightning.

[0035] Furthermore, by leveraging the dynamic thresholds provided by the machine learning model, this invention effectively avoids the over-segmentation of strong summer thunderstorms and the erroneous merging of weak spring thunderstorms, significantly improving the accuracy of the "flash return" operation and enhancing clustering fidelity. Moreover, this invention cleverly integrates supervised learning (predicting parameters based on first environmental features) with unsupervised learning (ST-DBSCAN density clustering), retaining the advantages of density clustering in processing spatial point clouds while endowing it with intelligent adjustment capabilities based on meteorological mechanisms, providing more scientific underlying algorithmic support for meteorological and forestry disaster prevention and control.

[0036] In one possible implementation, S210, historical Doppler weather radar combined reflectivity data is selected as the baseline fact of the true physical boundary of thunderstorm clouds, and the optimal ST-DBSCAN parameters of each historical thunderstorm under radar boundary constraints are derived in reverse through a heuristic search algorithm.

[0037] This invention introduces a heuristic search algorithm (such as Bayesian optimization) that enables the approximation of the globally optimal parameter combination with a finite number of searches within a two-dimensional parameter space composed of spatial and temporal neighborhood thresholds. This heuristic search algorithm uses the real spatial morphology of thunderstorm clouds reflected by radar combined reflectivity data as the objective function, iteratively updating candidate parameters and dynamically adjusting the search direction and step size. Compared to manual parameter tuning, it has higher parameter calibration efficiency and repeatability, significantly reducing computational overhead. Furthermore, this algorithm automatically adapts to the scale characteristics of various thunderstorm samples without requiring repeated design of search strategies for different regions and seasons, thus providing consistent and reliable training labels for subsequent machine learning models and ensuring the effectiveness of model training.

[0038] In one possible implementation, S400, based on the homogeneous lightning cluster data generated in S300, extracts and outputs the current lifespan duration of the thunderstorm cloud based on the occurrence time of the first and last ground lightning within the cluster, divides its existing lifespan into several time windows, extracts spatial cluster boundaries based on the homogeneous lightning cluster data within each time window and the spatial coordinate set of the lightning within the cluster, reconstructs the spatial outline boundary of the thunderstorm cloud within each time window and calculates its coverage area for output, and determines the centroid of the thunderstorm cloud from the spatial outline boundary to extract and output the movement trajectory of the thunderstorm cloud.

[0039] This invention transforms homogeneous lightning clusters into temporal profile sequences with clear geometric meaning and area attributes, elevating the intensity assessment and impact range determination of thunderstorm clouds from qualitative description to quantitative analysis. This coverage area can intuitively reflect the horizontal impact range and intensity evolution trend of thunderstorm clouds within each time window, providing directly quantifiable basic data support for subsequent thunderstorm cloud disaster level classification, disaster warning threshold setting, and quantitative retrospective analysis of historical thunderstorm events, demonstrating intuitive application value.

[0040] Furthermore, the S400 also synthesizes and outputs the life history profile and disaster-causing attribute characteristics of thunderstorm clouds. The life history profile includes the timestamps of the first and last discharges of the thunderstorm cell, the movement trajectory during the current life cycle, the temporal changes of the spatial contour boundary, etc. It can also extract the total lightning frequency of the thunderstorm cell in all effective grids within each time window as one of the components of the life history profile to characterize the temporal changes of lightning frequency (in different time windows).

[0041] As is well known, the output of the spatial outline boundary, coverage area, and movement trajectory of thunderstorm clouds within each time window, including timestamps, and the temporal changes of related data, all belong to the content of disaster-causing attribute features. In addition, the total lightning frequency of thunderstorm cells across all effective grids within each time window can be extracted, representing the temporal changes in lightning frequency across different time windows, and this can be output as a disaster-causing attribute feature. It can also include the maximum lightning frequency of thunderstorm cells in the current lifecycle and the ground flash density of thunderstorm cells within each time window.

[0042] In one possible implementation, S100, the first preset duration is 0.5-2 hours, preferably 1 hour. The time window duration of S400 can be set to a fixed value to determine the time resolution, and is within the permissible range of time resolution. It is designed according to the current thunderstorm cloud life cycle length so that the current thunderstorm individual life cycle is exactly divided into an integer number of time windows. The specific number can be set according to the actual situation.

[0043] In one possible implementation, the first environmental feature further includes underlying surface features. The underlying surface features of a grid include the topographic slope and elevation of the surface location corresponding to that grid. Preferably, it also includes the slope aspect, land cover type (land use type, which determines heat capacity, roughness, and evaporation efficiency), soil moisture, and surface roughness. Thermodynamic features include convective available potential energy and convective inhibition energy. Preferably, it also includes ambient temperature, relative humidity, K-index, lifting index, atmospheric precipitable water, vertical distribution of ambient temperature, and / or vertical distribution of dew point temperature. Dynamic features include vertical wind shear. Preferably, it also includes water vapor flux divergence, vertical velocity, and divergence / vorticity.

[0044] In one possible implementation, S400, based on the data of the same lightning clusters within each time window, constructs a two-dimensional convex hull polygon using the convex hull algorithm based on the set of spatial coordinates of the lightning within the cluster. This polygon serves as the spatial outline boundary of the reconstructed thunderstorm cloud within that time window and is output accordingly. Based on this, its coverage area and movement trajectory are calculated and output.

[0045] In one possible implementation, S400, based on the data of the same lightning clusters within each time window, constructs an Alpha Shape polygon using the Alpha Shape algorithm based on the set of spatial coordinates of the lightning within the cluster. This polygon serves as the spatial outline boundary of the reconstructed thunderstorm cloud within that time window and is output accordingly. Based on this, its coverage area and movement trajectory are calculated and output.

[0046] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for accurate identification and tracking of thunderstorm clouds based on machine learning, characterized in that, Includes the following steps: S100, the target monitoring area is divided into several spatial grids according to latitude and longitude. Based on discrete lightning location data, the first environmental features within the first preset time before the occurrence of historical lightning events in each grid are collected and extracted to construct a feature vector set. The first environmental features include thermodynamic features and dynamic features. S200, construct and train a machine learning model, where: S210, after constructing the machine learning model, calculate the optimal ST-DBSCAN parameters of each historical thunderstorm under radar boundary constraints based on historical Doppler weather radar combined reflectivity data, and use the obtained optimal ST-DBSCAN parameters as the training labels for the machine learning model, and combine them with the feature vector set of S100 to train the model. S220 collects the first environmental features of each grid in real time and inputs them into a pre-trained machine learning model for forward inference, and outputs the dynamic spatial neighborhood threshold and dynamic temporal neighborhood threshold of each grid in real time. S300, the dynamic spatial neighborhood threshold and dynamic temporal neighborhood threshold output in real time by the machine learning model are input into the ST-DBSCAN clustering engine to perform ST-DBSCAN dynamic spatiotemporal density clustering on the lightning data of all collected grids and extract homogeneous lightning clusters in real time. S400, based on the homogeneous lightning cluster data generated by S300, extracts the spatial cluster boundary based on the spatial coordinate set of lightning within the cluster, reconstructs the spatial outline boundary of thunderstorm clouds within each time window, and thus extracts and outputs the movement trajectory of thunderstorm clouds.

2. The method for accurate identification and tracking of thunderstorm clouds based on machine learning according to claim 1, characterized in that, S210 selects historical Doppler weather radar combined reflectivity data as the benchmark fact of the true physical boundary of thunderstorm clouds, and uses a heuristic search algorithm to reverse derive the optimal ST-DBSCAN parameters of each historical thunderstorm under radar boundary constraints.

3. The method for accurate identification and tracking of thunderstorm clouds based on machine learning according to claim 1, characterized in that, S400: Based on the homogeneous lightning cluster data generated by S300, extract and output the current life cycle duration of the thunderstorm cloud based on the occurrence time of the first and last ground lightning within the cluster. Divide its existing life cycle into several time windows. Based on the homogeneous lightning cluster data within each time window, extract the spatial cluster boundary based on the spatial coordinate set of the lightning within the cluster. Reconstruct the spatial outline boundary of the thunderstorm cloud within each time window and calculate its coverage area for output. Extract and output the movement trajectory of the thunderstorm cloud.

4. The method for accurate identification and tracking of thunderstorm clouds based on machine learning according to claim 3, characterized in that, S400 also synthesizes and outputs the life history archives and disaster-causing attributes of thunderstorm clouds.

5. The method for accurate identification and tracking of thunderstorm clouds based on machine learning according to claim 1, characterized in that, S100, the first preset duration is 0.5-2 hours.

6. The method for accurate identification and tracking of thunderstorm clouds based on machine learning according to claim 1, characterized in that, The first environmental feature also includes underlying surface features.

7. The method for accurate identification and tracking of thunderstorm clouds based on machine learning according to claim 6, characterized in that, The underlying surface features of a grid include the topographic slope and elevation of the surface location corresponding to that grid.

8. The method for accurate identification and tracking of thunderstorm clouds based on machine learning according to claim 1, characterized in that, The thermodynamic characteristics include ambient temperature, relative humidity, convective available potential energy, and K exponent, while the kinetic characteristics include vertical wind shear and water vapor flux divergence.

9. The method for accurate identification and tracking of thunderstorm clouds based on machine learning according to claim 1, characterized in that, S400: Based on the data of the same lightning clusters within each time window, and using the spatial coordinate set of lightning within the cluster, a two-dimensional convex hull polygon is constructed using the convex hull algorithm. This polygon serves as the spatial outline boundary of the reconstructed thunderstorm cloud within the time window and is output accordingly. The coverage area and movement trajectory are then calculated and output.

10. The method for accurate identification and tracking of thunderstorm clouds based on machine learning according to claim 1, characterized in that, S400: Based on the data of the same lightning clusters within each time window, and using the set of spatial coordinates of the lightning within the cluster, an Alpha Shape polygon is constructed using the Alpha Shape algorithm. This polygon serves as the spatial outline boundary of the reconstructed thunderstorm cloud within the time window and is output accordingly. The coverage area and movement trajectory are then calculated and output.