Lightning density construction system and method based on multi-source meteorological factors and machine learning

By employing multi-source meteorological factors and machine learning methods, a lightning density construction system was developed, which addresses the issues of insufficient spatial coverage and temporal consistency in existing datasets. This system generates a high-quality global lightning density dataset, supporting climate diagnostics and risk management applications.

CN122020172APending Publication Date: 2026-05-12NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV
Filing Date
2026-01-30
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies struggle to generate lightning density datasets that are globally comprehensive, have long-term time series, are spatially continuous, high-resolution, and temporally consistent, thus failing to meet the needs of applications such as climate diagnostics, extreme event assessment, and risk management.

Method used

A lightning density construction system was built using multi-source meteorological factors and machine learning methods. The system includes modules for data acquisition, preprocessing, feature construction, machine learning modeling, accuracy evaluation, and model integration. High-quality lightning density data is generated through parallel modeling and model fusion of multiple machine learning models.

Benefits of technology

A global lightning density dataset with wide coverage, high resolution, and continuous time was generated, which reduced the cost of data acquisition, improved the availability of data in long-term climate change analysis and historical comparison studies, and provided quality assessment and physical interpretation information for the data generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020172A_ABST
    Figure CN122020172A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of meteorological data processing and earth system information, provides a lightning density construction system and method based on multi-source meteorological factors and machine learning, and realizes spatial continuous construction of lightning density in historical periods by constructing a mapping relation between lightning observation data and meteorological factors. According to the method, ground lightning observation data is used as a modeling target, thermal conditions, dynamic conditions and cloud micro-physical related variables in re-analyzed meteorological data are used as input features, multiple machine learning models are constructed for training, prediction results of different models are fused by adopting an integration method, and the prediction accuracy is improved. And outputting global lightning density data under the unified spatial resolution and time scale. Meanwhile, the stability and the consistency of a model prediction result are quantitatively evaluated by calculating a precision evaluation index of a grid point scale, so that a global lightning density data set with continuous space, consistent time and relatively high stability is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of meteorological data processing and earth system information technology, specifically relating to a lightning density construction system and method based on multi-source meteorological factors and machine learning. Background Technology

[0002] Lightning activity is a crucial and direct indicator of severe convective weather processes, closely related to precipitation, ice-phase microphysical processes, deep convection latent heat release, upper tropospheric charge structure, and extreme disasters such as thunderstorms, hail, and short-duration heavy rainfall. For applications in climate diagnostics, extreme event assessment, and risk management, there is an urgent need to construct a global lightning density dataset with long-term series, spatial continuity, high resolution, and consistent caliber. This dataset would characterize the interannual or decadal variations of lightning and their regional differences, providing data support for improving convective parameterization in climate models, statistical analysis, training and validation of data-driven models, as well as operational applications such as meteorological services, insurance assessment, and infrastructure lightning protection.

[0003] However, existing lightning observation methods and related data products cannot simultaneously meet the above requirements in terms of spatial coverage, temporal continuity, and data consistency, mainly in the following aspects: (a) Limited spatial coverage and sampling conditions; Optical lightning detection based on low-Earth orbit satellites (such as OTD / LIS) suffers from significant temporal and spatial sampling inhomogeneity due to orbital characteristics and instantaneous field of view, particularly in high-latitude regions where effective sample data is insufficient. Existing LIS / OTD climatological products are primarily presented in gridded statistical form with multi-year averages. Their product attributes are essentially climatological descriptions of lightning activity, making it difficult to directly provide multi-year, monthly global lightning density gridded data with continuous temporal variations. Furthermore, their spatial resolution is relatively coarse, failing to meet the needs of refined climate diagnostics and long-term change analysis.

[0004] (ii) The ground-based global lightning location network has detection biases; The Global Ground-Based Lightning Location Network (WWLLN), based on very low frequency radio signals, enables long-term, near real-time global observation of lightning activity and can generate high-resolution lightning density products. However, this type of network exhibits some selectivity regarding lightning type and intensity, and its detection efficiency varies with time, region, and sensor site configuration. Particularly during its early operational phase (2005-2012), the gradual increase in the number of base stations easily introduced non-physical temporal variation trends. Although subsequent products (such as WGLC) have made some corrections to factors such as detection efficiency, their time series length remains relatively limited, making it difficult to cover earlier historical periods.

[0005] (iii) The reanalysis data lacks directly available lightning density products; While existing meteorological reanalysis data provide a wealth of variables related to convective environment, dynamic conditions, and cloud microphysics, they typically do not directly include long-term global lightning density gridded data that is strictly consistent with observational data in terms of scope. Related studies largely rely on empirical parameterization or statistical methods, using reanalysis variables as proxies for lightning activity estimation, but a unified and systematic technique for constructing lightning density datasets is still lacking.

[0006] In summary, there is an urgent need for a generation technology that can comprehensively utilize multi-source meteorological, cloud microphysical, and dynamic environmental factors to construct a lightning density dataset with a unified spatial grid and temporal resolution, covering the globe and possessing long-term series characteristics, so as to better support related scientific research and operational applications. Summary of the Invention

[0007] To address the aforementioned technical problems, this invention provides a lightning density construction system and method based on multi-source meteorological factors and machine learning, thereby resolving the issues in the prior art. The technical solution adopted by this invention is as follows: A lightning density construction system based on multi-source meteorological factors and machine learning includes: a data acquisition module, a data preprocessing module, a feature construction module, a machine learning modeling module, an accuracy evaluation module, a key factor analysis module, and a model integration and data output module; among which: The data acquisition module is used for collecting lightning observation data; The data preprocessing module is used to preprocess lightning observation data and outputs standardized meteorological environmental variable data and lightning target data in a unified spatiotemporal grid. The feature construction module is used to encode standardized meteorological environmental variable data and lightning target data into learnable inputs for the model, resulting in a standardized feature matrix and corresponding lightning density label data; The machine learning modeling module is used to capture the evolution of lightning density by learning biases of different algorithms, and obtain prediction results; The accuracy assessment module is used to generate grid-level assessment fields and statistical tables based on the prediction results, and output assessment reports and log records. The key factor analysis module is used to output a list of key factors, contribution measurement results, and explanatory analysis graphs based on the prediction results and the output of the accuracy assessment module. The model integration and data output module is used to save the prediction results of multiple models and the final output data.

[0008] The lightning density construction method based on multi-source meteorological factors and machine learning includes the following steps: Step 1: The data acquisition module is used for collecting lightning observation data; Step 2: The data preprocessing module is used to preprocess lightning observation data and outputs standardized meteorological environmental variable data and lightning target data in a unified spatiotemporal grid. Step 3: The feature construction module is used to encode standardized meteorological environmental variable data and lightning target data into learnable inputs for the model, resulting in a standardized feature matrix and corresponding lightning density label data. Step 4: The machine learning modeling module is used to capture the evolution pattern of lightning density through the learning bias of different algorithms to obtain prediction results; Step 5: The accuracy assessment module is used to generate grid-level assessment fields and statistical tables based on the prediction results, and output assessment reports and log records. Step 6: The key factor analysis module is used to output a list of key factors, contribution measurement results, and explanatory analysis graphs based on the prediction results and the output of the accuracy assessment module. Step 7: The model integration and data output module is used to save the multi-model prediction results and the final output data.

[0009] Furthermore, step 4 includes: The multi-algorithm parallel modeling steps for heterogeneous model grouping are as follows: construct multiple machine learning models with different structural types to perform lightning density prediction modeling in parallel, and classify them into different model categories according to the structural characteristics of the models, including ensemble learning models based on gradient boosting decision trees, random forest models based on bagging, and deep learning models based on multi-layer neural network structures. The automated parameter optimization steps based on validation feedback are as follows: During the training of each model, the key parameter set is determined for different model structures, and automated parameter optimization is introduced within the preset parameter space. The model parameters are iteratively adjusted based on the prediction performance feedback on the validation data. The model selection and overfitting control steps based on stability assessment are as follows: During the model training phase, the changes in the model's predictive performance under different sample partitioning conditions are evaluated through cross-validation; the generalization stability of the model is assessed based on the performance fluctuations under multiple training or different validation conditions. Model solidification and call preparation steps: After completing model training, parameter optimization and stability determination, the trained model is solidified by serializing and storing its model structure information and corresponding parameters, and establishing model index relationships.

[0010] Furthermore, step 4 includes: Step 6 includes: Model selection and interpretation path selection steps: After completing model training and ensemble prediction, models with relatively low prediction performance are first excluded; based on the evaluation results, the corresponding feature contribution analysis path is determined for different models according to the structural type of the machine learning model; among them, for models based on gradient boosting decision tree structure, the feature contribution decomposition method is selected; for random forest models based on bagging method, the feature importance measurement method based on splitting or permutation mechanism is selected. Preliminary evaluation steps for feature contributions based on model structure: For different models, calculate the overall contribution of their input features to the prediction results: For gradient boosting decision tree models, use a feature contribution decomposition method based on game theory to quantify the feature contribution; for random forest models, use a feature importance measurement method based on tree structure statistics or feature permutation to evaluate the relative importance of each feature in the model prediction. The steps for determining the cross-model key feature set are as follows: Based on the feature contribution ranking results obtained from different models, the features with high contribution in each model are summarized and cross-compared, and the features that show high contribution in multiple models are selected to construct a cross-model consistent key feature set; Key feature response relationship analysis steps: After determining the model with the best prediction performance and its corresponding set of key features, select the key features with the highest feature contribution. Based on the sample-level feature contribution decomposition results, analyze the changes in the contribution of key features to the prediction results in different value ranges. Ultimately, the dependencies and response trends between key features and prediction results are obtained, which are used to identify the impact characteristics of key meteorological factors on lightning density changes under different conditions.

[0011] The present invention has the following beneficial effects: (1) Generate a global lightning density dataset with wide coverage, high resolution and temporal continuity; Existing lightning data products based on LIS / OTD mainly exist in the form of multi-year averages and climatological data. Their effective coverage is mainly concentrated within about 38° north and south latitude, making it difficult to directly use them to construct continuous time series of global lightning density that varies from year to month. Moreover, their time series products have relatively coarse spatial resolution.

[0012] This invention introduces multi-source meteorological environmental factors and employs machine learning modeling methods to fully mine information from existing observation and reanalysis data without relying on new lightning observation hardware. This enables the effective expansion of lightning density in both time and space dimensions, generating a lightning density dataset with global coverage, high spatial resolution, and temporal continuity. This provides data support for global-scale lightning change analysis that differs from traditional climate products.

[0013] (2) Improve the availability of lightning density data in the time dimension with lower data acquisition costs; This invention utilizes periods of relatively stable WWLLN data quality to construct a machine learning prediction model. The statistical relationship learned by the model between lightning activity and meteorological environmental factors is then applied to the reconstruction of historical periods on longer timescales. This allows for the continuous reconstruction of historical lightning density without relying on long-term continuous lightning observations. Through this method, the invention can generate lightning density data covering a long historical period with lower data acquisition and processing costs, thereby significantly improving the usability of lightning density data in long-term climate change analysis and historical comparative studies.

[0014] (3) Reduce the uncertainty caused by a single model by integrating multiple models and improve the stability of the reconstruction results; This invention employs multiple machine learning models to model lightning density separately, and utilizes a ridge regression ensemble method to fuse the prediction results of different models. This effectively reduces the impact of single model structural assumptions and parameter selection on the prediction results, thereby reducing model uncertainty. Through the multi-model ensemble strategy, the lightning density reconstruction results generated by this invention exhibit higher stability and robustness in spatial distribution and temporal evolution, which is beneficial for conducting consistency analysis in different regions and time periods.

[0015] (4) Provide quality assessment and physical interpretation information simultaneously during the data generation process; This invention generates a lightning density reconstruction dataset and outputs a grid-scale prediction accuracy evaluation index. It also combines feature importance analysis and the SHAP method to analyze the relative contributions of different meteorological factors in lightning density prediction. This design not only helps in evaluating the quality of the reconstruction results but also provides auxiliary information for analyzing the physical mechanisms of lightning activity changes, thereby improving the interpretability and application value of the dataset in scientific research and engineering applications. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the overall framework of the present invention. Detailed Implementation

[0017] The following will be described in conjunction with embodiments of the present invention. Figure 1 The technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.

[0018] This invention proposes a lightning density construction system and method based on multi-source meteorological factors and machine learning. By constructing a mapping relationship between lightning observation data and meteorological factors, it achieves spatial continuity in the construction of historical lightning density. This invention uses ground-based lightning observation data as the modeling target and thermodynamic conditions, dynamic conditions, and cloud microphysical variables from reanalysis meteorological data as input features. Multiple machine learning models are constructed and trained, and the prediction results of different models are fused using an ensemble method to output global lightning density data at a unified spatial resolution and temporal scale. Simultaneously, by calculating a grid-scale accuracy evaluation index, the stability and consistency of the model prediction results are quantitatively evaluated, thereby obtaining a spatially continuous, temporally consistent, and highly stable global lightning density dataset.

[0019] Specifically, such as Figure 1 This invention proposes a lightning density construction system based on multi-source meteorological factors and machine learning, comprising: a data acquisition module, a data preprocessing module, a feature construction module, a machine learning modeling module, an accuracy evaluation module, a key factor analysis module, and a model integration and data output module; wherein: 1. The data acquisition module is used for collecting lightning observation data; specifically: The hardware of the data acquisition module is based on a communication interface (network card / wireless communication module) and a memory, used to receive or read external data resources; the software is based on a data access service / data acquisition program, which may include sub-components such as a downloader, API caller, file scanner and data verifier.

[0020] The data acquisition module takes WWLLN lightning observation data (Zenodo database records) and ERA5 reanalysis data (NetCDF files) as input.

[0021] The data acquisition module processes the following steps: parsing the time information, spatial grid information, and variable metadata of the data source; verifying file integrity, missing fragments, and time coverage; archiving the raw data by time period and variable type, and writing it to local storage or distributed storage.

[0022] The data acquisition module outputs: raw data file / data block index and metadata list (time range, spatial range, resolution, etc.) for the preprocessing module to use.

[0023] The data acquisition module enables stable access and traceable management of multi-source data through communication and storage capabilities, ensuring that subsequent processing uses consistent data versions as input.

[0024] 2. The data preprocessing module is used for preprocessing lightning observation data and outputs standardized meteorological environmental variable data and lightning target data in a unified spatiotemporal grid; specifically: The hardware / electronic components of the data preprocessing module are implemented in collaboration with the processor and memory. When necessary, parallel computing resources (multi-core CPU / cluster nodes) can be called to improve the efficiency of raster computing. Its software implementation is based on the preprocessing pipeline program and may include a time alignment submodule, a spatial interpolation submodule, a region masking submodule, and a missing value processing submodule.

[0025] The input to the data preprocessing module is the multi-source raw data and its metadata output from the data acquisition module.

[0026] The data preprocessing module's processes include: (1) Time alignment processing: Obtain the original time series from different data sources, and resample and time-match various types of data according to a preset time scale to make the multi-source data correspond to each other at the same time node or time interval. The time alignment processing includes time scale conversion and timestamp matching to ensure the consistency of multi-source data in the time dimension.

[0027] (2) Spatial interpolation and grid unification: To address the differences in spatial resolution and grid structure among different data sources, various types of data are converted to a unified spatial grid. Specifically, based on the spatial resolution and range of the target grid (0.25°), the original data (WWLLN, 0.1°) is interpolated to represent it on a unified latitude and longitude grid, thereby eliminating the impact of spatial scale differences on subsequent modeling.

[0028] (3) Region masking: Based on preset geographical regions or physical conditions, invalid regions in the spatial grid are identified and masked. The region masking is based on surface type (ocean) and geographical location (Antarctica, Greenland) to exclude regions that do not participate in modeling, so as to ensure that the data participating in modeling meets the physical rationality requirements.

[0029] (4) Handling missing and outlier values: The preprocessed data is checked for missing values ​​and the missing data is removed to ensure the continuity and numerical stability of the input data.

[0030] The data preprocessing module outputs: standardized meteorological environmental variable data and lightning target data in a unified spatiotemporal grid (training phase).

[0031] The data preprocessing module eliminates biases caused by differences in data sources by unifying spatiotemporal benchmarks and quality control, enabling subsequent feature construction and modeling to be carried out on a comparable and stable data basis.

[0032] 3. The feature construction module encodes standardized meteorological environmental variable data and lightning target data into learnable inputs for the model, obtaining a standardized feature matrix and corresponding lightning density label data; specifically: The hardware / electronic components of the feature construction module are implemented based on processors and memory, and can improve efficiency by using vectorized operations and batch read / write; its software implementation is based on feature engineering programs and can include thermal feature submodules, dynamic feature submodules, cloud microphysical feature submodules, geographic and temporal description feature submodules, and interactive / composite feature submodules.

[0033] The input to the feature construction module is the standardized meteorological variable field and its temporal / spatial index, which is output by the preprocessing module.

[0034] The feature construction module processes the following: Thermal instability feature construction extracts relevant variables characterizing atmospheric thermal instability and convective potential from meteorological reanalysis data, and constructs composite features that comprehensively reflect convective energy conditions and precipitation processes to characterize the thermal environment conducive to lightning occurrence. In one implementation, the thermal instability features can be obtained by combining or coupling convective available potential energy indices with precipitation-related variables to address the insufficient response of single thermal indices to actual convection and lightning activity. These composite features characterize the lightning potential under the combined influence of convective energy conditions and actual precipitation processes, providing more physically meaningful thermal input information for subsequent machine learning models. Environmental and geographic descriptive feature construction describes the environmental and geographic characteristics of geographical location (absolute latitude) and temporal attributes (month). These environmental and geographic features characterize the background differences in lightning activity across different regions and time periods, improving the model's generalization ability in both spatial and temporal dimensions.

[0035] The feature building module outputs: a feature matrix and a list of features (name, source variable, description) for model training / prediction.

[0036] The feature construction module encodes complex meteorological process information into learnable inputs for the model through systematic feature representation, thereby improving the physical rationality and generalization ability of lightning density reconstruction.

[0037] 3. The machine learning modeling module is used to capture the evolution of lightning density through the learning biases of different algorithms, and obtain prediction results; specifically: The hardware / electronic components of the machine learning modeling module are implemented using high-performance processors (CPU / GPU) to perform intensive matrix operations and gradient calculations; RAM is used to store large-scale feature matrices and intermediate model parameters; and a cache is used to accelerate the iterative training process. The software implementation of the machine learning modeling module is based on a machine learning modeling engine, including algorithm libraries (XGBoost, Random Forest, LightGBM, and Deep Neural Networks (DNN)), an automatic hyperparameter tuner (Optuna integration), and model serialization tools.

[0038] The inputs to the machine learning modeling module are: the standardized feature matrix (sample set) output by the feature construction module, the corresponding lightning density label data (y), etc.

[0039] 4. The processing steps of the machine learning modeling module include: (1) Parallel modeling steps of heterogeneous model grouping: Construct multiple machine learning models with different structural types to perform lightning density prediction modeling in parallel, and classify them into different model categories according to the structural characteristics of the models, including ensemble learning models based on gradient boosting decision trees, random forest models based on bagging, and deep learning models based on multi-layer neural network structures. Through the above parallel modeling method of heterogeneous model grouping, the complementarity of different models in feature representation ability is fully utilized, providing diversified prediction inputs for subsequent model fusion.

[0040] (2) Automated parameter optimization steps based on validation feedback: During the training of each model, the key parameter set is determined for different model structures, and an automated parameter optimization step is introduced within a preset parameter space. The model parameters are iteratively adjusted based on the prediction performance feedback from the validation data to obtain a parameter combination with better performance. The parameter optimization process can be implemented using the automated hyperparameter search method commonly used in this field. Its implementation can be based on the Optuna framework, but is not limited to this specific implementation method.

[0041] (3) Model screening and overfitting control steps based on stability judgment: During the model training stage, a model stability verification step is introduced. Through cross-validation, the changes in the model's predictive performance under different sample partitioning conditions are evaluated. The generalization stability of the model is judged based on the performance fluctuations of the model under multiple training or different verification conditions. When the model's predictive performance remains stable under different training conditions, the model is judged to have reliable generalization ability.

[0042] (4) Model solidification and call preparation steps: After completing model training, parameter optimization and stability determination, the trained model is solidified, its model structure information and corresponding parameters are serialized and stored, and a model index relationship is established for unified call in the subsequent model integration and lightning density prediction stages.

[0043] The output of the machine learning modeling module is: the candidate model file (y1-4) after training, the hyperparameter evaluation log, and the convergence curve of the training process.

[0044] The machine learning modeling module constructs a heterogeneous model pool with strong fitting capabilities. By learning biases from different algorithms, it captures the complex evolutionary patterns of lightning density, providing a high-performance foundational prediction source for subsequent multi-model ensembles. 5. The accuracy assessment module is used to generate grid-level assessment fields and statistical tables based on the prediction results, and output assessment reports and log records; specifically: Hardware / electronic implementation of the accuracy assessment module: Statistical calculations are performed by a processor, and memory is used to store intermediate statistics and assessment results. Software implementation of the accuracy assessment module: Assessment and report generation components, which may include an indicator calculator, a spatial statistician, and a result visualization / exporter.

[0045] The accuracy assessment module takes the following inputs: model prediction results, corresponding observation / label data, spatial mask, and sample index.

[0046] The accuracy assessment module processing includes: Calculate the coefficient of determination (R²) 2 ): Measures the model’s ability to explain data variability.

[0047] ; Root mean square error (RMSE): Measures the error between the predicted value and the actual value.

[0048] ; Correlation coefficient (r): measures the linear correlation between the actual value and the predicted value.

[0049] ; All models uniformly summarize the above indicators, generate grid-level evaluation fields and statistical tables, and output evaluation reports and log records.

[0050] The accuracy assessment module outputs: grid-level assessment index files, model comparisons, and global-scale visualizations.

[0051] The accuracy assessment module provides a quantitative basis for dataset quality control and model selection, ensuring that the final data output has traceable accuracy information. 6. The key factor analysis module outputs a list of key factors, contribution measurement results, and explanatory analysis graphs based on the prediction results and the output of the accuracy assessment module; specifically: Key factor analysis module hardware / electronic implementation: Interpretive computations are performed by a processor and memory, and parallel computing can be used to accelerate the interpretation of large-scale samples when necessary. Key factor analysis module software implementation: Interpretability analysis component, which may include a feature importance calculator, SHAP interpreter, dependency analyzer, and result exporter.

[0052] Key factor analysis module inputs: trained model, feature matrix, prediction results, and sample index.

[0053] The key factor analysis module process includes: (1) Model screening and interpretation path selection steps: After completing model training and ensemble prediction, models with relatively low prediction performance are first excluded, and their prediction results are not included in the key factor interpretation analysis to avoid interference from low-reliability models with feature interpretation results. Based on the above evaluation results, the corresponding feature contribution analysis path is determined for different models according to the structural type of the machine learning model. Among them, for models based on gradient boosting decision tree structure, the feature contribution decomposition method suitable for tree models is selected; for random forest models based on bagging method, the feature importance measurement method based on splitting or permutation mechanism is selected.

[0054] (2) Preliminary evaluation steps of feature contribution based on model structure: For different models, calculate the overall contribution of their input features to the prediction results: For the model based on gradient boosting decision tree, use the feature contribution decomposition method based on game theory to quantify the feature contribution. In one implementation, the SHAP method can be used; For the random forest model, use the feature importance measurement method based on tree structure statistics or feature permutation to evaluate the relative importance of each feature in the model prediction. In one implementation, the Feature Importance calculation method can be used.

[0055] (3) Steps for determining the cross-model key feature set: Based on the feature contribution ranking results obtained from different models, the features with higher contribution in each model are summarized and cross-compared, and the features that show high contribution in multiple models are selected to construct a cross-model consistent key feature set.

[0056] (4) Key Feature Response Relationship Analysis Steps: After determining the model with the best prediction performance and its corresponding set of key features, select the key features with the highest feature contribution ranking. Based on the sample-level feature contribution decomposition results, analyze the changes in the contribution of key features to the prediction results within different value ranges. Through the above steps, obtain the dependency relationship and response trend between key features and prediction results, which can be used to identify the influence characteristics of key meteorological factors on lightning density changes under different conditions.

[0057] (5) Key factor results output and interpretation support steps: The final determined key factors and their ranking and interpretation charts are output as interpretation analysis results to support the credibility assessment and physical mechanism analysis of the lightning density reconstruction results.

[0058] Key factor analysis module outputs: a list of key factors, contribution measurement results, and explanatory analysis graphs.

[0059] The key factor analysis module provides physical interpretability evidence while outputting prediction results, supporting mechanism analysis and model credibility assessment.

[0060] 7. The model integration and data output module is used to save the prediction results of multiple models and the final output data; specifically: Hardware / electronic implementation of the model integration and data output module: The processor performs the fusion calculations, and the memory stores the multi-model prediction results and the final output data. Software implementation of the model integration and data output module: It integrates fusion components and data publishing components, and may include a weight solver, a fusion predictor, a formatter writer, and a metadata generator.

[0061] The inputs to the model integration and data output module are: prediction results from multiple machine learning models (y1-4), prediction target (lightning density), and output grid definition.

[0062] The model integration and data output module processing includes: Fusion Form: The final prediction is generated using weighted fusion, which can be expressed as follows: ; Weight Calculation: In one implementation, ridge regression can be used to calculate the weights, with the objective function being: ; Data output of the model integration and data output module: The final lightning density results are written in a standardized format (NetCDF file) according to a unified grid and time scale, and metadata such as variable descriptions, units, time coverage, and evaluation indicators are written.

[0063] Output of the model integration and data output module: final global lightning density dataset file from 1979 to the present and its metadata file.

[0064] The model integration and data output module reduces the uncertainty of a single model through integration and fusion, and outputs data products in a standardized format that can be directly used for scientific analysis and engineering applications.

[0065] Overall hardware and software implementation of this invention: The system described in this invention can be deployed on computer equipment, servers, or cloud computing platforms. The computer equipment includes at least a processor (CPU), memory (including RAM and non-volatile memory), a communication interface (network interface card / wireless communication module), and input / output interfaces. The memory stores data processing programs for executing each functional module. The processor calls and executes the programs, enabling each functional module to run as a software module, service component, or containerized service, and achieving data interaction between modules through a data bus or inter-process communication mechanism. To facilitate scalable processing, the system may further include a task scheduling component and a data caching component to support batch processing or parallel computing of multi-time, multi-region data. These components can be implemented as thread pools, process pools, distributed task queues, or workflow scheduling engines, but the specific implementation method is not limited.

[0066] In specific implementation of this invention, such as Figure 1 The system comprises two modules: a data acquisition module and a data preprocessing module. The data acquisition module acquires lightning observation data and corresponding reanalysis meteorological data within a preset time period and transmits the raw data to the data preprocessing module. The data preprocessing module performs time alignment, spatial interpolation, regional masking, and missing value filtering on the multi-source data to generate standardized input data with uniform spatial resolution and time scale. The preprocessed data is then fed into the feature construction module to construct a set of feature variables describing meteorological environmental conditions, forming feature samples for model training and prediction. These feature samples are further input into the machine learning modeling module, which constructs multiple machine learning prediction models and trains them based on lightning observation data. After training, the prediction results of each model are sent to both the accuracy evaluation module (to calculate model prediction performance indicators) and the key factor analysis module (to analyze the relative importance and physical mechanisms of different meteorological features in lightning density prediction). Finally, the model integration and data output module fuses the prediction results of multiple machine learning models using a ridge regression integration method and outputs a global lightning density dataset in a unified data format, thus forming a lightning density construction result with temporal continuity and spatial consistency.

[0067] Based on the system of this invention, this invention also proposes a method for constructing lightning density based on multi-source meteorological factors and machine learning, including the following steps: Step 1: The data acquisition module is used for collecting lightning observation data; Step 2: The data preprocessing module is used to preprocess lightning observation data and outputs standardized meteorological environmental variable data and lightning target data in a unified spatiotemporal grid. Step 3: The feature construction module is used to encode standardized meteorological environmental variable data and lightning target data into learnable inputs for the model, resulting in a standardized feature matrix and corresponding lightning density label data. Step 4: The machine learning modeling module is used to capture the evolution pattern of lightning density through the learning bias of different algorithms to obtain prediction results; Step 5: The accuracy assessment module is used to generate grid-level assessment fields and statistical tables based on the prediction results, and output assessment reports and log records. Step 6: The key factor analysis module is used to output a list of key factors, contribution measurement results, and explanatory analysis graphs based on the prediction results and the output of the accuracy assessment module. Step 7: The model integration and data output module is used to save the multi-model prediction results and the final output data.

[0068] Furthermore, step 4 includes: The multi-algorithm parallel modeling steps for heterogeneous model grouping are as follows: construct multiple machine learning models with different structural types to perform lightning density prediction modeling in parallel, and classify them into different model categories according to the structural characteristics of the models, including ensemble learning models based on gradient boosting decision trees, random forest models based on bagging, and deep learning models based on multi-layer neural network structures. The automated parameter optimization steps based on validation feedback are as follows: During the training of each model, the key parameter set is determined for different model structures, and automated parameter optimization is introduced within the preset parameter space. The model parameters are iteratively adjusted based on the prediction performance feedback on the validation data. The model selection and overfitting control steps based on stability assessment are as follows: During the model training phase, the changes in the model's predictive performance under different sample partitioning conditions are evaluated through cross-validation; the generalization stability of the model is assessed based on the performance fluctuations under multiple training or different validation conditions. Model solidification and call preparation steps: After completing model training, parameter optimization and stability determination, the trained model is solidified by serializing and storing its model structure information and corresponding parameters, and establishing model index relationships.

[0069] Furthermore, step 4 includes: Step 6 includes: Model selection and interpretation path selection steps: After completing model training and ensemble prediction, models with relatively low prediction performance are first excluded; based on the evaluation results, the corresponding feature contribution analysis path is determined for different models according to the structural type of the machine learning model; among them, for models based on gradient boosting decision tree structure, the feature contribution decomposition method is selected; for random forest models based on bagging method, the feature importance measurement method based on splitting or permutation mechanism is selected. Preliminary evaluation steps for feature contributions based on model structure: For different models, calculate the overall contribution of their input features to the prediction results: For gradient boosting decision tree models, use a feature contribution decomposition method based on game theory to quantify the feature contribution; for random forest models, use a feature importance measurement method based on tree structure statistics or feature permutation to evaluate the relative importance of each feature in the model prediction. The steps for determining the cross-model key feature set are as follows: Based on the feature contribution ranking results obtained from different models, the features with high contribution in each model are summarized and cross-compared, and the features that show high contribution in multiple models are selected to construct a cross-model consistent key feature set; Key feature response relationship analysis steps: After determining the model with the best prediction performance and its corresponding set of key features, select the key features with the highest feature contribution. Based on the sample-level feature contribution decomposition results, analyze the changes in the contribution of key features to the prediction results in different value ranges. Ultimately, the dependencies and response trends between key features and prediction results are obtained, which are used to identify the impact characteristics of key meteorological factors on lightning density changes under different conditions.

[0070] The above embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Any modifications, alterations, alterations, or substitutions made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.

Claims

1. A lightning density construction system based on multi-source meteorological factors and machine learning, characterized in that, include: The module comprises a data acquisition module, a data preprocessing module, a feature construction module, a machine learning modeling module, an accuracy evaluation module, a key factor analysis module, and a model integration and data output module; among which: The data acquisition module is used for collecting lightning observation data; The data preprocessing module is used to preprocess lightning observation data and outputs standardized meteorological environmental variable data and lightning target data in a unified spatiotemporal grid. The feature construction module is used to encode standardized meteorological environmental variable data and lightning target data into learnable inputs for the model, resulting in a standardized feature matrix and corresponding lightning density label data; The machine learning modeling module is used to capture the evolution of lightning density by learning biases of different algorithms, and obtain prediction results; The accuracy assessment module is used to generate grid-level assessment fields and statistical tables based on the prediction results, and output assessment reports and log records. The key factor analysis module is used to output a list of key factors, contribution measurement results, and explanatory analysis graphs based on the prediction results and the output of the accuracy assessment module. The model integration and data output module is used to save the prediction results of multiple models and the final output data.

2. A method for constructing lightning density based on multi-source meteorological factors and machine learning, characterized in that, Includes the following steps: Step 1: The data acquisition module is used for collecting lightning observation data; Step 2: The data preprocessing module is used to preprocess lightning observation data and outputs standardized meteorological environmental variable data and lightning target data in a unified spatiotemporal grid. Step 3: The feature construction module is used to encode standardized meteorological environmental variable data and lightning target data into learnable inputs for the model, resulting in a standardized feature matrix and corresponding lightning density label data. Step 4: The machine learning modeling module is used to capture the evolution pattern of lightning density through the learning bias of different algorithms to obtain prediction results; Step 5: The accuracy assessment module is used to generate grid-level assessment fields and statistical tables based on the prediction results, and output assessment reports and log records. Step 6: The key factor analysis module is used to output a list of key factors, contribution measurement results, and explanatory analysis graphs based on the prediction results and the output of the accuracy assessment module. Step 7: The model integration and data output module is used to save the multi-model prediction results and the final output data.

3. The lightning density construction method based on multi-source meteorological factors and machine learning according to claim 2, characterized in that, Step 4 includes: The multi-algorithm parallel modeling steps for heterogeneous model grouping are as follows: construct multiple machine learning models with different structural types to perform lightning density prediction modeling in parallel, and classify them into different model categories according to the structural characteristics of the models, including ensemble learning models based on gradient boosting decision trees, random forest models based on bagging, and deep learning models based on multi-layer neural network structures. The automated parameter optimization steps based on validation feedback are as follows: During the training of each model, the key parameter set is determined for different model structures, and automated parameter optimization is introduced within the preset parameter space. The model parameters are iteratively adjusted based on the prediction performance feedback on the validation data. The model selection and overfitting control steps based on stability assessment are as follows: During the model training phase, the changes in the model's predictive performance under different sample partitioning conditions are evaluated through cross-validation; the generalization stability of the model is assessed based on the performance fluctuations under multiple training or different validation conditions. Model solidification and call preparation steps: After completing model training, parameter optimization and stability determination, the trained model is solidified by serializing and storing its model structure information and corresponding parameters, and establishing model index relationships.

4. The lightning density construction method based on multi-source meteorological factors and machine learning according to claim 2, characterized in that, Step 6 includes: Model selection and interpretation path selection steps: After completing model training and ensemble prediction, models with relatively low prediction performance are first excluded; based on the evaluation results, the corresponding feature contribution analysis path is determined for different models according to the structural type of the machine learning model; among them, for models based on gradient boosting decision tree structure, the feature contribution decomposition method is selected; for random forest models based on bagging method, the feature importance measurement method based on splitting or permutation mechanism is selected. Preliminary evaluation steps for feature contributions based on model structure: For different models, calculate the overall contribution of their input features to the prediction results: For gradient boosting decision tree models, use a feature contribution decomposition method based on game theory to quantify the feature contribution; for random forest models, use a feature importance measurement method based on tree structure statistics or feature permutation to evaluate the relative importance of each feature in the model prediction. The steps for determining the cross-model key feature set are as follows: Based on the feature contribution ranking results obtained from different models, the features with high contribution in each model are summarized and cross-compared, and the features that show high contribution in multiple models are selected to construct a cross-model consistent key feature set; Key feature response relationship analysis steps: After determining the model with the best prediction performance and its corresponding set of key features, select the key features with the highest feature contribution. Based on the sample-level feature contribution decomposition results, analyze the changes in the contribution of key features to the prediction results in different value ranges. Ultimately, the dependencies and response trends between key features and prediction results are obtained, which are used to identify the impact characteristics of key meteorological factors on lightning density changes under different conditions.