Fixed-point Lightning Warning Method Based on Multi-source Data

By performing feature extraction and fusion analysis on multi-source data, a lightning warning model was established, and the operability problem of point-shaped target lightning warning was solved, and accurate warnings for flammable and explosive places and scenic spots were achieved, reducing the losses of lightning disasters.

CN120178384BActive Publication Date: 2025-07-25CHONGQING LIGHTNING PROTECTION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510671704.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-07-25
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing technology is difficult to provide accurate and timely lightning warnings on point-shaped targets such as flammable and explosive places and scenic spots, and lacks operability.

Method used

By extracting and fusion analysis of multi-source monitoring data such as atmospheric electric field data, new generation weather radar data, lightning positioning data, etc., a lightning warning model is established using principal component analysis and binary logistic regression method to conduct fixed-point lightning warning.

Benefits of technology

Provide operational lightning warning methods for lightning disaster-sensitive places such as oil depots, schools, scenic spots, and industrial parks to reduce disaster losses and ensure safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120178384B_ABST
    Figure CN120178384B_ABST
Patent Text Reader

Abstract

The present invention discloses a fixed-point lightning warning method based on multi-source data, comprising the steps of: data collection and preprocessing; sorting the multi-source data set to obtain a sample set; performing sample factor correlation analysis on the sample set to obtain a sample matrix; performing Z-Score standardization on the sample matrix to obtain a standardized matrix; performing linear transformation on the standardized matrix by using the principal component analysis method to extract the principal components; based on the extracted principal components, establishing a lightning warning model by using the binary logistic regression method, and inputting a training data set to train the lightning warning model; and performing fixed-point lightning warning by using the trained lightning warning model. The remarkable effect is that it provides an operable and implementable lightning warning method for lightning disaster-sensitive places such as oil depots, schools, scenic spots, industrial parks, mines, etc., and provides a scientific basis for managers to take lightning prevention measures in advance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lightning warning, and particularly relates to a fixed-point lightning warning method based on multi-source data. Background Technique

[0002] Lightning is one of the most serious natural disasters, which greatly endangers people's lives and property. The harm of lightning strikes to some flammable and explosive places and lightning disaster-sensitive places such as oil depots, gas filling stations, fireworks and firecracker warehouses, schools, scenic spots, etc. is particularly serious. Once lightning strike explosions or casualty accidents occur, it will cause irreparable losses and great social impacts. Doing a good job in lightning monitoring and warning services is an important part of lightning disaster prevention work. However, lightning has the characteristics of instantaneity, randomness, and suddenness in space and time. There are still certain difficulties in carrying out accurate and timely lightning forecasting and warning services, especially for point-like and small-scale targets such as oil depots and scenic spots. Therefore, improving the level of fixed-point lightning forecasting and warning is a difficult point in current scientific research work.

[0003] Lightning is a product of severe convective weather processes. The generation, dissipation, aggregation, and variation laws of severe convective cloud clusters and the lightning generation mechanism of thunderclouds are very complex, resulting in the characteristics of instantaneity, randomness, and suddenness in the spatial and temporal distribution of lightning. Early lightning nowcasting mainly used the extrapolation of weather radar and cloud images. Later, with the development of numerical prediction models and data assimilation technologies, the combination of extrapolation and high-resolution numerical prediction was proposed. More research tended to establish prediction equations or models by obtaining lightning observation data in real time, combining weather radar, meteorological satellites, sounding data, ground electric field data, lightning location data, etc., and using various mathematical methods to find the relationship between the observation data and thunderstorm phenomena.

[0004] In recent years, with the development of machine learning and deep learning technologies, various machine learning methods have been widely applied to lightning warning research. For example, a lightning nowcasting model is constructed based on neural network algorithms, or feature factors are extracted from multi-source data, and various machine learning methods are used to establish lightning forecasting models for comparative research. In addition, a deep semantic segmentation model is also used to fuse multi-source observation data to extract multi-source data forecasting factors to achieve lightning nowcasting. However, the research of domestic and foreign scholars mostly focuses on the lightning potential forecasting and lightning nowcasting on the surface or in a grid format, and the operability for point-like targets and small-scale targets is not strong enough.

[0005] In practical work, there is a relatively urgent need and actual application scenario for short-term fixed-point lightning warnings in lightning disaster-sensitive sites such as flammable and explosive areas, scenic spots, mines, and industrial parks. Therefore, if we extract features and conduct fusion analysis on multi-source monitoring data such as atmospheric electric field data, new-generation weather radar data, and lightning location data, explore fixed-point lightning warning methods based on multi-source data, and establish corresponding warning models, it will provide operable and implementable lightning warning methods for lightning disaster-sensitive sites such as oil depots, schools, scenic spots, industrial parks, and mines, and provide a scientific basis for managers to take lightning prevention measures in advance, which will be of great significance in reducing lightning disaster losses, ensuring enterprise production safety, and protecting people's lives and safety. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a fixed-point lightning warning method based on multi-source data. By extracting features and conducting fusion analysis on multi-source monitoring data such as atmospheric electric field data, new-generation weather radar data, and lightning location data, a lightning warning model is constructed for fixed-point lightning warning, which can solve the technical defect that the operability for point targets and small-range targets in the background technology is not strong enough.

[0007] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0008] The present invention proposes a fixed-point lightning warning method based on multi-source data, which is characterized by including the following steps:

[0009] Step 1: Data collection and preprocessing to obtain a multi-source data set including atmospheric electric field data, lightning location data, and weather radar data;

[0010] Step 2: Sort the multi-source data set to obtain a sample set;

[0011] Step 3: Conduct sample factor correlation analysis on the sample set to obtain a sample matrix;

[0012] Step 4: Perform Z-Score standardization on the sample matrix to obtain a standardized matrix;

[0013] Step 5: Use the principal component analysis method to perform linear transformation on the standardized matrix, and extract five principal components: atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor;

[0014] Step 6: Based on the five principal components of the extracted atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor, use the binary logistic regression method to establish a lightning warning model, and input the training data set to train the lightning warning model;

[0015] Step 7: Use the trained lightning warning model for fixed-point lightning warning;

[0016] The expression of the trained lightning warning model is:

[0017] ;

[0018] where ~ are the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor respectively, p is the probability of lightning occurrence.

[0019] Further, the data preprocessing in Step 1 includes the following processes:

[0020] Step 1.1: According to the lightning warning location and the defined coverage area range, divide the collected multi-source data into data coverage areas to achieve preliminary data screening;

[0021] Step 1.2: For the multi-source data after preliminary data screening, use different methods for data processing according to the data type to obtain a multi-source data set, where: the atmospheric electric field data is processed using the mean imputation method; the weather radar data is processed using the bilinear interpolation algorithm.

[0022] Further, the processing process of the lightning location data in Step 1.2 is as follows:

[0023] Determine the distance between the target point and the lightning location point according to the longitude and latitude of the target point, screen out all lightning location data with a distance less than the preset distance threshold from the target point, and form a lightning location data set for the target area after sorting by time;

[0024] Eliminate the lightning location data with a lightning current value less than the preset lightning current amplitude in the lightning location data set for the target area;

[0025] According to the positioning principle of the lightning locator, eliminate the lightning location data with a positioning method lower than three stations in the lightning location data set.

[0026] Further, the sample set obtained by sample sorting in Step 2 specifically includes:

[0027] Determine the sample elements;

[0028] Determine the sample recording time and the time interval between each group of samples;

[0029] Determine the variables included in a single sample and their meanings;

[0030] Sort out the data in the multi-source dataset according to the sample elements, sample recording time, time interval between each group of samples, variables included in a single sample and their meanings, and obtain the sample set.

[0031] Further, the variables included in a single sample include sample time, mean value, absolute value mean, absolute value maximum, mean value of difference score, maximum value of difference score, variance, standard deviation, number of reversals, whether reversed, echo mean, proportion of 15 km, proportion of 30 km.

[0032] Further, the formula for performing sample factor correlation analysis on the sample set in step 3 is:

[0033] ;

[0034] Among them, is the correlation coefficient between samples and , n is the sample size, and and are the standard scores, sample mean values and sample standard deviations of samples respectively, and and are the standard scores, sample mean values and sample standard deviations of samples respectively.

[0035] Further, the specific steps for obtaining the standardized matrix by performing Z-Score standardization on the sample matrix in step 4 include:

[0036] Calculate the mean value and standard deviation SD of each sample in the sample matrix;

[0037] Perform standardization according to the formula ; among them, is the value of a certain sample in the sample matrix.

[0038] Further, in step 5, the principal component analysis method is used to perform linear transformation on the standardized matrix, and the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor are extracted. Specifically, it includes:

[0039] Step 5.1: Establish the covariance matrix R of the standardized matrix;

[0040] Step 5.2: Solve the eigenvalues, principal component contribution rates and cumulative variance contribution rates of the covariance matrix R, and determine the number of principal components;

[0041] Step 5.3: According to the principle of selecting the number of principal components, count the eigenvalues whose eigenvalues are greater than 1 and the cumulative contribution rate reaches more than 80% λ 1 , λ 2 , …, λ m and the corresponding m;

[0042] Step 5.4: After performing a linear transformation with m unit eigenvectors as coefficients, obtain the eigenvector matrix a T , and calculate the first m sample principal components according to the formula to obtain the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor;

[0043] Step 5.5: Establish an initial factor loading matrix to interpret the principal components.

[0044] Furthermore, the mathematical expressions of the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor are:

[0045] ;

[0046] ;

[0047] ;

[0048] ;

[0049] ;

[0050] where ~ are the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor respectively, ~ are the variables included in a single sample respectively.

[0051] Furthermore, in Step 6, based on the five principal components extracted, a lightning warning model is established using the binary logistic regression method, and the training data set is input to train the lightning warning model, specifically including:

[0052] Step 6.1: Using "whether there is lightning" as the dependent variable and the five principal components of the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor as covariates, introduce the sigmond function to establish a lightning warning model;

[0053] Step 6.2: Input the training data set, estimate the parameters of the lightning warning model using the maximum likelihood estimation method, and train the lightning warning model using the gradient descent method and quasi-Newton method.

[0054] The remarkable effects of the present invention are as follows:

[0055] 1. By performing feature extraction and fusion analysis on multi-source lightning monitoring data such as atmospheric electric field data, new generation weather radar data, and lightning location data, and conducting data quality control such as interpolation and cleaning, a thunderstorm sample set and a non-thunderstorm sample set are formed. Then, the relationship between multiple warning factors and lightning occurrence results is analyzed using statistical methods, and correlation analysis, Z-Score standardization, principal component analysis (PCA), binary Logistic regression, etc. are performed on the sample data. The method and standard for extracting the principal component factors are established, and the first to fifth principal components are extracted as the modeling factors of the Logistic regression model to establish a fixed-point lightning warning model based on multi-source data, providing an operable and implementable lightning warning method for lightning disaster-sensitive places such as oil depots, schools, scenic spots, industrial parks, and mines, providing a scientific basis for managers to take lightning prevention measures in advance, and having important significance in reducing lightning disaster losses, ensuring enterprise production safety, and protecting people's lives.

[0056] 2. Thirteen sample factors that are all positively correlated with "whether there is lightning" are sorted out, and considering factors such as eigenvalue, single-item variance percentage, scree plot test, and cumulative variance percentage, a total of 5 principal components are extracted by fusing the multi-source data of the 13 factors. Certain meanings are given to and these 5 principal components are named according to the factor loading coefficients. Among them, the first principal component is the atmospheric electric field discrete factor, the second principal component is the radar echo intensity factor, the third principal component is the atmospheric electric field reversal factor, the fourth principal component is the atmospheric electric field intensity factor, and the fifth principal component is the radar echo extreme value factor. Thus, a binary Logistic regression model is established. From various test indicators such as the classification result, performance indicators, and evaluation indicators of the model, the model is effective and has a good fitting degree. Description of the Drawings

[0057] Figure 1 is the flowchart of the method described in the present invention;

[0058] Figure 2 is the schematic diagram of area division;

[0059] Figure 3 is the trend chart of the correlation coefficient between each sample factor and "whether there is lightning" in the present invention changing with the lead time;

[0060] Figure 4It is a heat map of the correlation between each factor and "whether there is lightning";

[0061] Figure 5 It is a scree plot of the eigenvalues of the principal components;

[0062] Figure 6 It is a schematic diagram of the establishment process of the lightning warning model;

[0063] Figure 7 It is a trend chart of the performance parameters of the binary Logistic regression model changing with the lead time t when the cut-off value is 0.5;

[0064] Figure 8 It is a trend chart of the performance parameters of the binary Logistic regression model changing with the lead time t when the cut-off value is 0.3;

[0065] Figure 9 It is the F1 value changing with the lead time t when the cut-off values of the model are 0.2 to 0.7 respectively;

[0066] Figure 10 It is a schematic block diagram of the system described in the present invention;

[0067] Figure 11 It is a schematic block diagram of the computer device described in the present invention. Specific embodiments

[0068] The following further elaborates in detail on the specific embodiments and working principles of the present invention in conjunction with the accompanying drawings.

[0069] Embodiment

[0070] As Figure 1 shown, this embodiment provides a fixed-point lightning warning method based on multi-source data, and the specific steps are as follows:

[0071] Step 1: Data collection and preprocessing to obtain a multi-source data set including atmospheric electric field data, lightning location data, and weather radar data;

[0072] In this example, the atmospheric electric field data is from the data continuously collected by the Pre-storm2.0 type field mill atmospheric electric field instrument installed in the field of the Jinfo Mountain lightning special experiment from 2020 to 2022. The basic working principle of this electric field instrument is: the motor drives the shielding metal sheet to rotate, making the induction metal sheet alternately exposed to the electric field or shielded, thereby generating induced charges proportional to the external electric field intensity. The induction sheet is connected to the amplification processing circuit and the waveform adjustment circuit, and outputs a voltage signal. After calibration, this voltage can characterize the intensity and polarity change of the atmospheric electric field. Its main technical parameters are shown in Table 1.

[0073]

[0074] In this example, the lightning location data is from the location data of the ADTD lightning location system in Chongqing from 2020 to 2022. The system consists of a second-generation ADTD lightning locator, a central data processing station, a user data service network, and a graphic display terminal. The observation network of the system consists of 5 monitoring stations in Chongqing and 10 monitoring stations in neighboring provinces. The system realizes continuous automatic monitoring of elements such as the time, location (longitude, latitude), peak lightning current, and polarity of cloud-to-ground lightning. The cloud-to-ground lightning detection efficiency is above 85%, the detection and location accuracy within the network is less than 300 m, and the processing time of lightning return strokes is about 1 ms.

[0075] The weather radar data uses the mosaic product of the combined reflectivity factor of a new generation of Doppler weather radar, which is from the Chongqing Meteorological Observatory. The time resolution is 6 min, and the spatial resolution is 0.01°×0.01° (about 1 km×1 km). The data format is a png image file, and the size of the map is 1400×1200 pixels, that is, the image coverage area is approximately (100°E, 23°N)~(114°E, 35°N). The echo intensity of the pixel points in the image is the gray level of the points. Since the file is an 8-bit gray-scale image, there are 256 gray levels in total.

[0076] It should be noted that in addition to the above data types, other types of lightning observation data such as trial lightning data, meteorological sounding data, and meteorological satellite data can also be included.

[0077] In some specific embodiments, the data preprocessing in step 1 includes the following processes:

[0078] Step 1.1: According to the lightning warning location and the defined coverage area range, divide the collected multi-source data into data coverage areas to achieve preliminary data screening.

[0079] By observing the obtained lightning observation data of different data types, it can be seen that except for the atmospheric electric field data which is point data, the rest of the data are surface data. In order to conduct preliminary screening of the data range, it is necessary to first divide the data coverage area. Referring to the division of the target area, surrounding area, monitoring area, and coverage area in "Lightning Protection Thunderstorm Warning System" (GB / T38121—2023), the following areas are defined in this embodiment, and the data is preliminarily screened accordingly:

[0080] Target area: The actual geographical range where lightning warning needs to be carried out, which can be a single location such as an oil depot, a school, a building, etc., or can be extended to a larger area such as a scenic area, a chemical industrial park, etc. In this embodiment, it refers to the outer field of the special lightning experiment on Jinfo Mountain (latitude: 28.9451N, longitude: 107.1462E).

[0081] Monitoring area: The area where an alarm needs to be triggered when lightning is approaching or occurring. In this embodiment, it refers to the area within a circle with a radius of 5 km centered on the outdoor field of the special lightning experiment on Jinfo Mountain.

[0082] Coverage area: Taking the outdoor field of the special lightning experiment on Jinfo Mountain as the center, the set coverage area radius is 15 km. The main considerations are as follows: (1) For local detectors / atmospheric electric field meters, the electric field of thunderstorm clouds within 15 km has a greater impact on the electric field at the location of the atmospheric electric field meter in the target area; (2) When the high-altitude wind speed is 2 km / min, the time for a thunderstorm cloud at a distance of 15 km to move above the monitoring area is about 5 - 6 minutes, which is equivalent to the volume scan cycle of the new generation weather radar.

[0083] The schematic diagram of area division is as Figure 2 shown. According to the area division results, the multi-source data collected is divided into data coverage areas to achieve preliminary data screening.

[0084] Step 1.2: For the multi-source data that has undergone preliminary data screening, different methods are used for data processing according to the data type to obtain a multi-source data set, where:

[0085] 1) Atmospheric electric field data processing:

[0086] For atmospheric electric field data, network interruption and delay may both cause packet loss. The processing of missing values is carried out according to the following method: When data loss is detected, statistical analysis is performed on the data for 1 minute, which is 30 seconds before and 30 seconds after the missing bit, to obtain the average value, and the missing bit is filled with the average value, that is, the mean interpolation method. For the situation of long-term data loss caused by equipment failure, etc., no interpolation processing is performed, and such time periods are not used as the source time for sample extraction.

[0087] 2) Lightning location data processing:

[0088] (1) Distance screening

[0089] Set the installation location of the atmospheric electric field meter as the target point. Then, lightning occurring near this point may cause lightning strike damage to the protected objects in the target area. Assume that lightning occurring within a circle with a radius of r centered on this point is regarded as the object to be pre-warned. Then, there are the following considerations: If r takes a too large value, the possibility of lightning far from the target area harming the protected objects is very small, and at the same time, there are more pre-warning objects, resulting in too many alarm numbers, which is not conducive to the development of lightning protection work; if r takes a too small value, not only will there be missed reports, but also the positioning data is less, resulting in a reduction in the screened thunderstorm samples, which is not conducive to the construction of the model in subsequent work. After repeated screening and comparison, in this embodiment,r Set to 5 km.

[0090] The distance calculation formula between the target point determined by latitude and longitude and the lightning location point is:

[0091] (1)

[0092] Among them, lat 1 , lon 1 , lat 2 , lon 2 are the latitude, longitude, latitude, and longitude of the target point of the lightning location point respectively. Use the above formula to screen out all lightning location data with a distance less than 5 km from the target point and sort them by time to form a lightning location data set for the target area.

[0093] (2) Noise and interference data processing

[0094] Since the ADTD lightning location system mainly uses the time difference of the electromagnetic waves radiated by each ground flash return stroke reaching each station for location, in the current complex electromagnetic environment, noise will inevitably be introduced. At the same time, due to various factors such as cloud flash interference, false ground flash location data will be generated. Its characteristic is that the current value in the location data is small, and such data should be excluded. The processing of location data with too small lightning current values in the lightning location data set follows the following principle: Exclude lightning location data with a lightning current amplitude of 0 - 2 kA. The "Technical Guide for Lightning Disaster Risk Zoning" (QX / T 405—2017) also points out that data above 200 kA should be excluded. However, in this embodiment, considering that the target area is above 2000 meters above sea level and the possibility of high-intensity ground flashes is relatively large, and in addition, it is also found in the statistical process that the overall proportion of data with a lightning current amplitude above 200 kA is very small and has a relatively minor impact on the research results, so it is not excluded.

[0095] (3) Deletion of unreliable data

[0096] According to the location principle of the ADTD lightning location system, its two-station location uses the magnetic orientation intersection method, and the location accuracy is poor due to terrain interference. Considering the topographic and geomorphic characteristics of Chongqing mainly being mountains and hills, all records in the lightning location data set with a location method lower than three stations are deleted.

[0097] 3) Weather radar data processing

[0098] Regarding the missing frame phenomenon in weather radar data, considering that the echo generally changes continuously and the change between adjacent two frames is small, in order to improve the utilization rate of data, the missing data is filled by bilinear interpolation. The bilinear interpolation algorithm is a commonly used one in image interpolation algorithms, that is, using the radar echoes at T-1 and T+1 moments, the radar echo at T moment is calculated by the bilinear interpolation method. Its basic principle is to calculate through the weighted sum of the 4 nearest neighbor pixel values of the pixel to be obtained in the source image.

[0099] Use Matlab software for image recognition and process the jigsaw product into element data such as the mean echo intensity and maximum value of the monitoring area, and the proportion of echoes exceeding 35 dBz within the coverage area. In addition, for reference and comparison, the proportion of echoes exceeding 35 dBz within 30 km is also counted as one of the attributes of the sample.

[0100] 4) Other data processing measures

[0101] Only use those objects with all attribute values. Since various data sources are different, it is impossible to ensure the consistency of data recording time. For example, due to equipment and network failures, the atmospheric electric field instrument has data missing for a relatively long period, and this kind of missing cannot be filled by interpolation, while the radar data may be relatively complete during this corresponding period. Therefore, for the samples sorted out, there must be some missing attribute values. For such situations, the corresponding samples are all deleted, and only the samples with all attribute values are retained.

[0102] Through the above data preprocessing and quality control, useless data can be effectively reduced, which helps to improve the accuracy of fixed-point lightning warning.

[0103] Step 2: Sort out the samples from the multi-source dataset to obtain a sample set;

[0104] In some embodiments, the specific implementation process of the sample data is as follows:

[0105] Step 2.1: Determine the basic principles for sorting out thunderstorm weather samples;

[0106] To avoid frequent switching of the warning status, the thunderstorm warning system can use the alarm retention time to set the minimum time for maintaining the alarm. If the retention time is too short, multiple alarms may be issued during a single thunderstorm process or the alarm status may switch frequently. If the retention time is too long, the alarm may not be lifted in a timely manner. In this embodiment, the retention time is set to 30 min, and based on this, the lightning location data is grouped: if the time difference between the recording time of a lightning location data record and the recording time of the adjacent data record behind it is less than 30 min, these two pieces of data are regarded as being in the same group. A group formed in this way is recorded as a thunderstorm process, and the time of the first lightning occurrence in each process is used as the reference time. According to the above principles, the lightning location data set in the target area from 2020 to 2022 is processed, and a total of 105 thunderstorm processes are sorted out.

[0107] Step 2.2: Determine the sample elements, sample recording time, time interval between each group of samples, the variables included in a single sample and their representative meanings;

[0108] For fixed-point lightning warning, the performance of various lightning-related monitoring data within a period of time before the first lightning occurrence in each thunderstorm process is the object of study in this embodiment. Assume that the time of the first lightning occurrence is T l , and the warning lead time is t lead , then the sample recording time is T s = T l - t lead . Since the volume scan period of the new generation weather radar is 6 min, the time interval between each group of samples should be 6 min. A single sample includes the following variables:

[0109] Sample recording time T s ; Atmospheric electric field strength value and its partial statistics; Echo intensity of the new generation weather radar and its partial statistics; Data from other sources. The variables included in a single sample and their representative meanings are shown in Table 2:

[0110]

[0111] Step 2.3: Sort out the data in the multi-source data set according to the sample elements, sample recording time, time interval between each group of samples, the variables included in a single sample and their representative meanings to obtain a sample set. For non-thunderstorm weather processes, this embodiment conducts statistics according to the following principles: ① No lightning occurs in the target area and the monitoring area during the statistical period; ② Precipitation occurs in the target area during the statistical period and the precipitation is greater than or equal to 5 mm.

[0112] Step 3: Conduct a sample factor correlation analysis on the sample set to obtain a sample matrix;

[0113] In some embodiments, the specific implementation process of conducting a sample factor correlation analysis on the sample set is as follows:

[0114] Correlation analysis refers to the analysis of two or more correlated variables to measure the degree of correlation between them. There must be a certain connection or probability between variables for correlation analysis to be carried out.

[0115] Conducting a multivariate correlation analysis on the sorted sample factors can initially understand whether there is a relationship between factors, between factors and the "whether lightning" result, the strength and direction of the relationship, and whether this relationship is statistically significant, etc. In addition, since principal component analysis requires a certain correlation between factors, it is necessary to conduct a correlation analysis on the sample factors before the modeling work.

[0116] The Pearson correlation coefficient is widely used to measure the degree of correlation between two variables, and its value ranges from -1 to 1. The Pearson correlation coefficient between two variables is defined as the quotient of the covariance and standard deviation between the two variables, so this correlation coefficient is also called the "Pearson product-moment correlation coefficient":

[0117] (2)

[0118] For samples, the equivalent expression is as follows:

[0119] (3)

[0120] Where is the correlation coefficient between samples , n is the sample size, , , are the standardized scores, sample means, and sample standard deviations of samples respectively, , , are the standardized scores, sample means, and sample standard deviations of samples respectively.

[0121] The trend of the correlation coefficient between each factor and "whether lightning" with the lead time is as Figure 3 , and the correlation heat map (t lead = 18) is as Figure 4 . It can be seen that:

[0122] (1)The correlation coefficient between each factor of the atmospheric electric field and "whether there is lightning" is significantly greater than that between each factor of the radar and "whether there is lightning", indicating that the correlation between each factor of the atmospheric electric field and "whether there is lightning" is stronger.

[0123] (2)Except for the "mean value", the correlation coefficients between the remaining factors of the atmospheric electric field and "whether there is lightning" show a downward trend with the increase of t lead The correlation coefficients between each factor of the radar and "whether there is lightning" basically remain unchanged with the increase of t lead This indicates that as time approaches, the relationship between each factor of the atmospheric electric field and "whether there is lightning" becomes stronger. However, because the atmospheric electric field has positive and negative values and the mean value changes greatly within one hour, the correlation between the mean value and "whether there is lightning" does not show an obvious increasing or decreasing trend over time, but shows a large degree of randomness. Within 1 hour, the relationship between each factor of the radar and "whether there is lightning" will not change significantly over time;

[0124] (3)The several factors with the strongest correlations are "whether it reverses", "absolute value mean", "standard deviation", "absolute value maximum", "maximum difference value", and "mean difference value".

[0125] In summary, there are correlations of varying degrees among the factors, and the vast majority of factors are positively correlated (confidence level 0.95). Among them, the correlations among the factors of the atmospheric electric field and among the factors of the radar are relatively high, but the correlation between the factors of the atmospheric electric field and the factors of the radar is weak, and some factors are negatively correlated.

[0126] Step 4: Perform Z-Score standardization on the sample matrix to obtain a standardized matrix;

[0127] In the specific implementation process, since the magnitudes and units of the factors in the sample will cause certain difficulties to the analysis work or seriously affect the PCA output, thereby affecting the accuracy of the subsequent modeling, in order to eliminate this influence, data standardization is required. This embodiment uses the Z-Score standardization method, and the specific calculation method is as follows:

[0128] Calculate the mean value of the factor and the standard deviation SD , and then subtract the mean value from each value of the factor and divide by the standard deviation, that is:

[0129] (4)

[0130] After Z-Score standardization, the data is closer to the standard normal distribution, with the mean of the variable being 0 and the standard deviation being 1. Through transformation, multiple sets of data are converted into unitless Z-Score values, standardizing the data and improving data comparability.

[0131] Step 5: Use the principal component analysis method to perform a linear transformation on the standardized matrix, and extract five principal components: the discrete factor of atmospheric electric field, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor.

[0132] Analysis of the sample factors shows that the number of sample factors (i.e., variables) is relatively large, and from the perspective of data structure, it belongs to a high-dimensional space. In this way, the difficulty of analyzing and modeling numerous factors increases sharply with the increase in the number of factors. At the same time, the correlation between sample factors is relatively high (the problem of multicollinearity), and the information redundancy is relatively high. In addition, when a serious collinearity problem occurs, it will lead to unstable analysis results, and the sign of the regression coefficient is completely opposite to the actual situation. At this time, it involves data dimensionality reduction processing and elimination of the influence of multicollinearity. Data dimensionality reduction refers to the process of mapping high-dimensional data to a low-dimensional space, and its purpose is to reduce the dimensionality of data while trying to retain the internal structure of the data. The general idea is to find the internal structure of the data and compress the original high-dimensional data into a low-dimensional space through techniques such as projection, coding, or reconstruction, for the convenience of data processing, analysis, and visualization.

[0133] Principal Component Analysis (PCA) is an attribute aggregation technique proposed by Karl Pearson in 1901. The basic idea is to map high-dimensional data to the principal component subspace through orthogonal transformation, retain the main components, and remove the secondary components. Usually, several new variables that are fewer than the number of original variables and can explain most of the differences in the data are selected, that is, the so-called principal components. For the lightning warning sample data, the main functions of principal component analysis are: one is to reduce the dimensionality of the data and simplify the complexity of the subsequent modeling process; the other is that the newly constructed variables are not correlated with each other, which is convenient for the next step of binary logistic regression modeling. The specific calculation steps are as follows:

[0134] Step 5.1: Use the data in the standardized matrix to establish a covariance matrix R;

[0135] Covariance is a statistical indicator reflecting the degree of closeness of the correlation between standardized factors. The larger the value, the more necessary it is to perform principal component analysis on the data. Among them, R ij (i, j = 1, 2,..., p) is the correlation coefficient between factor X i and X j . According to the calculation method of the correlation coefficient, R is a real symmetric matrix (i.e., Rij =R ji ), it is only necessary to calculate its upper triangular elements or lower triangular elements, and its calculation formula is:

[0136] (5)

[0137] Step 5.2, Solve the eigenvalues, principal component contribution rates and cumulative variance contribution rates of the covariance matrix R, and determine the number of principal components;

[0138] Solve the characteristic equation to find the eigenvalues λ i (i = 1, 2,..., p) and the corresponding eigenvectors a i . Since R is a positive definite matrix, its eigenvalues λ i are all positive numbers. Arrange them in descending order, that is λ 1 ≥ λ 2 ≥... ≥ λ i ≥ 0. The eigenvalues are the variances of the respective principal components, and their magnitudes reflect the influence of each principal component. Then the contribution rate of the i th principal component:

[0139] (6)

[0140] The cumulative contribution rate is:

[0141] (7)

[0142] Step 5.3, According to the principle of selecting the number of principal components, count the eigenvalues whose eigenvalues are greater than 1 and the cumulative contribution rate reaches more than 80% λ 1 , λ 2 , …, λ m and the corresponding 1, 2, …, m (m ≤ p), and m is the number of principal components;

[0143] Step 5.4, After performing a linear transformation with m unit eigenvectors as coefficients, obtain the eigenvector matrix a T , and find the first m sample principal components to obtain the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor;

[0144] (8)

[0145] Step 5.5. Establish the initial factor loading matrix and interpret the principal components:

[0146] (9)

[0147] Taking the sample with t lead = 18 as an example, the steps of principal component analysis using SPSS are as follows:

[0148] 1) Set "variables", and set all standardized Z-Score variables as "variables";

[0149] 2) Factor analysis: Description, select "KMO and Bartlett's test of sphericity";

[0150] 3) Factor analysis: Extraction, select "Principal components", extract eigenvalues greater than 1, and display the "Scree plot";

[0151] 4) Factor analysis: Rotation, select "Varimax method" for the method, and display the "Rotated solution";

[0152] 5) Check whether the cumulative rate C of the "Sum of squared loadings" in the "Total Variance Explained" table is greater than or equal to 80%. If C < 80%, go back to step 3) and make corresponding adjustments to the size of the extracted eigenvalues;

[0153] 6) Based on the results of factor analysis, define the same number of new variables, and input the factor loadings in the "Component Matrix" into the new variables defined in the new data file respectively;

[0154] 7) Use "Transform - Compute Variable" to calculate the characteristic variables and further obtain the eigenvector matrix;

[0155] 8) Further calculate the principal components from the eigenvector matrix.

[0156] Perform the KMO (Kaiser - Meyer - Olkin) and Bartlett's test of sphericity on the obtained principal components. The results are shown in Table 3:

[0157]

[0158] The null hypothesis of Bartlett's test of sphericity is that the correlation coefficient matrix is an identity matrix. Its approximate chi - square statistic is 4969.168, and the p - value 0.000 < 0.001. Therefore, the null hypothesis is rejected, indicating that there is a correlation relationship among the variables and it is suitable for factor analysis. In addition, the KMO test statistic is a sampling adequacy test proposed by Kaiser, Meyer, and Olkin, which is an index for comparing the simple correlation coefficients and partial correlation coefficients among variables, and its value ranges between 0 and 1.

[0159] Generally speaking, the closer the KMO value is to 1, the stronger the correlation between variables, and the more suitable the original variables are for factor analysis. It can be seen that after the original data is standardized by Z-Score, the KMO test statistic is 0.728, the data structure is reasonable, and it is suitable for factor analysis.

[0160] The communalities are shown in Table 4. It can be seen that the minimum value of the extracted communality is 0.774, indicating that after the Z-Score standardized data undergoes a principal component transformation, more information of the original variables can be extracted, and the common factors have a good expression for the Z-Score variables.

[0161]

[0162] The total variance explained is shown in Table 5. Among them, the left part is the initial eigenvalues, the middle is the result of extracting the main factors, and the right is the result of the rotated factor analysis. It can be seen that after the transformation of 13 factors, there are 13 main factors in total. Among them, the eigenvalue of the first main factor is 5.826, accounting for 45.095% of the total variance percentage. Similarly, the eigenvalues of the second to fifth main factors are 2.649, 1.192, 0.947, and 0.84 in turn. Generally speaking, the cumulative variance contribution rate reflects the proportion of information retained by the main factors. In Table 5, the cumulative variance percentage of the first five main factors is 88.393%, which can reflect most of the information of the original data.

[0163]

[0164] The number of main factors in the factor analysis results determines the number of principal components in the principal component analysis. The number of main factors to be extracted can be determined according to one or several of the following methods:

[0165] (1) The eigenvalue is greater than 1;

[0166] (2) The percentage of single variance is greater than 5%;

[0167] (3) Using the scree plot test;

[0168] (4) The cumulative variance percentage is greater than 80%.

[0169] Although the eigenvalues of the fourth and fifth main factors are slightly less than 1, the percentages of single variance are 7.286% and 6.461% respectively, both of which are greater than 5%. The cumulative variance percentage of the first to fifth main factors reaches 88.393%. Considering the aforementioned extraction principles comprehensively, a total of the first five main factors are extracted. At this time, less information of the original data is lost, the effect of the principal component analysis is relatively ideal, and it has research significance. Figure 5The scree plot with eigenvalues shows a steep slope for the larger factors and a gentle tail for the remaining factors. Generally, the main factors are selected on the steep slope, while the factors on the gentle slope have very little explanatory power for the variation. It can be seen that the curve becomes gentle after the fifth main factor, indicating that the first five principal components should be extracted.

[0170] Analyzing the rotated component matrix shown in Table 6, in the first principal component, the load coefficient values of 5 factors, namely the maximum absolute value, the mean difference value, the maximum difference value, the variance, and the standard deviation, are relatively large, indicating that the first principal component mainly explains these 5 factors; in the second principal component, the load coefficient values of 3 factors, namely the echo mean, the proportion at 15 km, and the proportion at 30 km, are relatively large; in the third principal component, the load coefficient values of 2 factors, namely the number of reversals and whether to reverse, are relatively large; in the fourth principal component, the load coefficient value of the mean is relatively large; in the fifth principal component, the load coefficient value of the maximum echo is relatively large. In the table, the data with factor load values exceeding 0.7 are all shown in bold. According to the factor load coefficients, certain meanings are assigned to these 5 principal components and named. For example, the first principal component is the atmospheric electric field discrete factor, the second principal component is the radar echo intensity factor, the third principal component is the atmospheric electric field reversal factor, the fourth principal component is the atmospheric electric field intensity factor, and the fifth principal component is the radar echo extreme value factor.

[0171]

[0172] The calculation of the eigenvector is carried out according to the following formula:

[0173] (10)

[0174] where is the i th column vector in the component matrix, and is the eigenvalue corresponding to the i th principal component.

[0175] After "Transform - Compute", the eigenvector matrix a T is as follows:

[0176]

[0177] Then the principal components can be calculated, that is, the eigenvector matrix is multiplied by the standardized sample matrix. The mathematical expressions of the atmospheric electric field discrete factor, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor are as follows:

[0178] ;

[0179] ;

[0180] ;

[0181] ;

[0182] ;

[0183] Step 6. Based on the five principal components of the extracted atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor, establish a lightning warning model using the binary logistic regression method, and input the training data set to train the lightning warning model;

[0184] Through principal component analysis, the original 13 factors (variables) are expressed by 5 principal components, which are: atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor. Principal component analysis realizes the dimensionality reduction of data, and the obtained component factors are not correlated with each other. The next step is to use binary Logistic regression to model the fixed-point lightning warning. The specific process is as follows:

[0185] Step 6.1. Take "whether there is lightning" as the dependent variable, and take the five principal components of the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor as covariates, and introduce the sigmond function to establish a lightning warning model;

[0186] The value "1" indicates "lightning occurs", and the value "0" indicates "no lightning occurs". If represents the probability of lightning occurrence, when its value is greater than or equal to 0.5, it is predicted that lightning occurs, and when it is less than 0.5, it is predicted that no lightning occurs, then there is:

[0187] (11)

[0188] Introduce the sigmond function as follows:

[0189] (12)

[0190] Then the mathematical model of the lightning warning model can be constructed as follows:

[0191] (13)

[0192] In formula (13), obeys the binary Logistic distribution, p represents the probability of lightning occurrence, (1 - p ) represents the probability of no lightning occurrence, β is the parameter to be estimated, yis the input variable. In this embodiment, y refers to the principal component extracted after standardization and linear transformation. It can be seen from the above formula that the Logistic model establishes the relationship between the probability of lightning occurrence and the principal component (explanatory variable). The model is shown as Figure 6 shown.

[0193] Step 6.2: Input the training data set , and use the maximum likelihood estimation method to estimate the parameters of the lightning warning model β , so as to obtain the Logistic regression model of the lightning warning model. Its log-likelihood function is:

[0194] (14)

[0195] Find the maximum value of to obtain the estimated value of β . In this way, the problem becomes an optimization problem with the log-likelihood function as the objective function. The methods commonly used in Logistic regression learning are the gradient descent method and the quasi-Newton method.

[0196] After obtaining the lightning warning model, this embodiment also conducts performance evaluation on it, which is specifically as follows:

[0197] The performance of the lightning warning model is mainly evaluated using three indicators: the effective alarm rate, the missed alarm rate, and the false alarm rate. The higher the effective alarm rate and the lower the missed alarm rate and the false alarm rate, the better the performance of the model. For a classification model, these three indicators are all related to the classification accuracy, but only when the class distribution in the classification is equal, the accuracy is a useful indicator. This means that under the condition of data imbalance, that is, when the data of a certain classification in the sample is significantly more than that of another classification, the accuracy cannot fully reflect the model performance, resulting in the fact that these three indicators of the effective alarm rate, the missed alarm rate, and the false alarm rate cannot very accurately reflect the advantages and disadvantages of the model.

[0198] One of the methods to solve the imbalance problem is to use better indicators. Here, the F1 score is introduced. It not only considers the number of errors predicted by the model but also considers the type of errors. As a performance measure of the classifier, the F1 score is defined as the harmonic mean of the precision and the recall rate. Its calculation formula is as follows:

[0199] (15)

[0200] The value of the F1 score ranges from 0 to 1. Generally speaking, the larger the F1 value, the better the performance of the classifier, and only when the FAR and the FTWR are both close to 0 at the same time, the F1 score is close to the maximum value of 1.

[0201] Taking tTaking 18:00 as an example, the results of establishing a binary Logistic regression model using SPSS are as follows:

[0202]

[0203]

[0204] From the Omnibus test of the model coefficients, it can be seen that the corresponding p-value is less than 0.001, which is less than the significance level of 0.05, indicating that the model is effective and has a good fitting degree. In addition, the ranges of Cox & Snell R Square and Nagelkerke R Square in Table 8 are from 0 to 1, and the larger the value, the better the model fitting degree. Here, they are 0.407 and 0.674 respectively, indicating that the model has a good fitting degree.

[0205] To obtain the correct classification rate of an object through a classification model, it must be better than the correct rate obtained by placing all objects in the class with the largest number of objects (the majority class, that is, no thunderstorm occurred). From the classification result a shown in Table 9, the classification correct rate for no thunderstorm occurred is 99.3%, the classification correct rate for thunderstorm occurred is 71.2%, and the overall percentage is 94.4%, which is better than the overall percentage of 82.5% (the classification result b shown in Table 10) by classifying all objects with thunderstorm occurred as no thunderstorm occurred. The model has improved the classification correct rate by nearly 12%, indicating that the model is effective.

[0206]

[0207]

[0208] The Wald chi-square value and the significance level p are the hypothesis tests for the regression coefficient B value. In Table 11, the significance p-values are all less than 0.05, indicating that the effects of all principal component factors on the results are statistically significant.

[0209]

[0210] Therefore, the final expression of the lightning warning model is:

[0211] (16)

[0212] Where, ~ are the discrete factor of atmospheric electric field, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor respectively, p is the probability of lightning occurrence.

[0213] Through a series of data analysis and processing methods such as data standardization, principal component analysis, and binary Logistic regression, a functional expression for the probability p of thunderstorm occurrence was finally obtained. The dependent variable of the function is the first to fifth principal components. The principal component analysis - binary Logistic regression model is as shown in Figure 6 . In this model, the data standardization and principal component analysis processes are carried out for the entire sample. As the sample size and sample values change, the expression of the model will also change accordingly. For comparison, Figure 7 、 Figure 8 respectively show the changing trends of the performance parameters of the binary Logistic regression model with the lead time t when the cut-off values are 0.5 and 0.3, and the following conclusions are drawn:

[0214] 1) There is no significant difference in the effective alarm rates of the two models, and they are basically above 90%;

[0215] 2) When the cut-off value is 0.5, the false alarm rate of the model is relatively low and does not exceed 7.41% at most. However, as t increases, the miss rate increases rapidly. When t = 54, the miss rate is as high as 55.36%. When the cut-off value is 0.3, a relatively low false alarm rate can be maintained. For example, when t ≤ 48, it is basically below 10%;

[0216] 3) When t ≤ 18, both the miss rate and false alarm rate of the model with a cut-off value of 0.3 are less than those of the model with a cut-off value of 0.5. At this time, the performance of the model with a cut-off value of 0.3 is better;

[0217] 4) When t ≥ 24, the two models have their own focuses. The miss rate of the model with a cut-off value of 0.3 is about 10% lower than that of the model with a cut-off value of 0.5, but the false alarm rate is significantly higher than that of the model with a cut-off value of 0.5.

[0218] Due to the imbalance of sample classification data, the non-thunderstorm samples are significantly more than the thunderstorm samples. It is necessary to further calculate the F1 score value to evaluate the quality of the model. Figure 9 is the curve of the F1 value changing with the lead time t when the cut-off values of the model are from 0.2 to 0.7. Through comparison, the following conclusions can be drawn:

[0219] 1) As the lead time t increases, the F1 score value shows a downward trend, indicating that the larger the lead time, the worse the classification effect and the worse the early warning effect;

[0220] 2) For the principal component analysis - binary Logistic regression model, when the cut-off value changes from 0.2 to 0.7, the F1 value first increases and then decreases. When the cut-off value is 0.3, the classification effect of the model is the best.

[0221] It can be seen that when the sample data is class-imbalanced, introducing the F1 score value can more objectively and accurately reflect the performance of the model. The lead time and threshold have a great impact on the model performance parameters. Specifically, the larger the lead time \(t\), the worse the performance. When the threshold changes from small to large, the F1 value first increases and then decreases. When the threshold is 0.3, the F1 value is the highest at 0.92 (\(t = 6\)) and the lowest at 0.68 (\(t = 54\)). And in most cases, the F1 score value is significantly higher than the F1 score values at other thresholds. It shows that the model has the best classification effect when the threshold is 0.3. Taking the lead time \(t = 18\) minutes as an example, the effective alarm rate, missed alarm rate, and false alarm rate of the early warning are 95.83%, 22.03%, and 4.17% respectively.

[0222] Step 7: Use the trained lightning early warning model to conduct fixed-point lightning early warning.

[0223] Embodiment

[0224] As Figure 10 shown, this embodiment provides a multi-source data-based fixed-point lightning early warning system for implementing the method described in Embodiment 1. The system includes:

[0225] A data collection and preprocessing module for collecting and preprocessing multi-source lightning monitoring data to obtain a multi-source data set including atmospheric electric field data, lightning location data, and weather radar data;

[0226] A sample sorting module for sorting the multi-source data set to obtain a sample set;

[0227] A correlation analysis module for performing sample factor correlation analysis on the sample set to obtain a sample matrix;

[0228] A normalization processing module for performing Z-Score normalization on the sample matrix to obtain a normalized matrix;

[0229] A principal component extraction module for performing linear transformation on the normalized matrix by using the principal component analysis method to extract five principal components: the atmospheric electric field discrete factor, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor;

[0230] A model establishment and training module for establishing a lightning early warning model by using the binary logistic regression method according to the five principal components of the atmospheric electric field discrete factor, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor extracted, and inputting a training data set to train the lightning early warning model;

[0231] An early warning module for using the trained lightning early warning model to conduct fixed-point lightning early warning.

[0232] Embodiment

[0233] As Figure 11 shown, this embodiment provides a computer device for implementing the fixed-point lightning warning method based on multi-source data described in Embodiment 1. The computer device includes:

[0234] One or more processors, one or more power supplies, one or more operating systems, one or more computer programs, one or more databases, a memory, one or more network interfaces, and one or more input / output interfaces.

[0235] The processor can execute the method steps described in the aforementioned Embodiment 1, which will not be elaborated here.

[0236] The power supply can meet the power requirements for the normal operation or overclocking of the computer device.

[0237] The operating system such as TM, MacOSXTM, UnixTM, LinuxTM, etc. Note that when selecting the operating system, pay attention to the version of the running code corresponding to the operating system.

[0238] The memory stores one or more application programs or data, which can be volatile storage or persistent storage. The programs stored in the memory include one or more modules, and each module can include a series of instruction operations on the computer device. The central processing unit can communicate with the memory and execute a series of instruction operations in the memory on the computer device.

[0239] Embodiment

[0240] This embodiment provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the method described in Embodiment 1.

[0241] It should be noted that the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, and the readable storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories, random access memories, or optical discs, including several instructions to enable a computer device, which can be a personal computer, a server, or a network device, etc., to execute all or part of the steps of the methods described in various embodiments of this application.

[0242] In summary, the present invention extracts features and performs fusion analysis on multi-source lightning monitoring data such as atmospheric electric field data, new generation weather radar data, and lightning location data, and conducts data quality control such as interpolation and cleaning to form a thunderstorm sample set and a non-thunderstorm sample set. Then, statistical methods are used to analyze the relationship between multiple warning factors and lightning occurrence results, and correlation analysis, Z-Score standardization, principal component analysis (PCA), binary Logistic regression, etc. are performed on the sample data. The method and standard for extracting principal component factors are established, and the first to fifth principal components are extracted as the modeling factors of the Logistic regression model to establish a fixed-point lightning warning model based on multi-source data, providing an operable and implementable lightning warning method for lightning disaster-sensitive sites such as oil depots, schools, scenic spots, industrial parks, and mines, providing a scientific basis for managers to take lightning prevention measures in advance, and having important significance in reducing lightning disaster losses, ensuring enterprise production safety, and protecting people's lives and safety.

[0243] The technical solution provided by the present invention has been introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A fixed-point lightning warning method based on multi-source data, characterized in that, Including the steps: Step 1, data collection and preprocessing, obtaining a multi-source data set including atmospheric electric field data, lightning location data, and weather radar data; Step 2, sorting out the samples of the multi-source data set to obtain a sample set; Step 3, performing sample factor correlation analysis on the sample set to obtain a sample matrix; Step 4, performing Z-Score standardization on the sample matrix to obtain a standardized matrix; Step 5, using the principal component analysis method to perform a linear transformation on the standardized matrix, and extracting five principal components: atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor; Step 6, based on the five principal components of the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor extracted, using the binary logistic regression method to establish a lightning warning model, and inputting the training data set to train the lightning warning model; Step 7, using the trained lightning warning model for fixed-point lightning warning; The expression of the trained lightning warning model is: ; Among them, ~ are the discrete factor of atmospheric electric field, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor respectively, p is the probability of lightning occurrence.

2. The method for fixed-point lightning warning based on multi-source data according to claim 1, wherein: The data preprocessing described in Step 1 includes the following processes: Step 1.1, according to the lightning warning location and the defined coverage area range, dividing the collected multi-source data into data coverage areas to achieve preliminary data screening; Step 1.2, for the multi-source data that has passed the preliminary data screening, using different methods for data processing according to the data type to obtain a multi-source data set, where: the atmospheric electric field data is processed using the mean imputation method; the weather radar data is processed using the bilinear interpolation algorithm.

3. The fixed-point lightning warning method based on multi-source data according to claim 2, characterized in that: The processing process of the lightning location data in Step 1.2 is as follows: Determine the distance between the target point and the lightning location point according to the longitude and latitude of the target point, screen out all lightning location data with a distance less than the preset distance threshold from the target point, and form a lightning location data set in the target area after sorting by time; Delete the lightning location data with a lightning current value less than the preset lightning current amplitude in the lightning location data set in the target area; According to the positioning principle of the lightning locator, delete the lightning location data with a positioning method lower than three stations in the lightning location data set.

4. The fixed-point lightning warning method based on multi-source data according to claim 1, wherein: The specific process of obtaining the sample set by sorting out the samples in Step 2 includes: Determine the sample elements; Determine the sample recording time and the time interval between each group of samples; Determine the variables included in a single sample and their meanings; Sort out the data in the multi-source data set according to the sample elements, sample recording time, time interval between each group of samples, variables included in a single sample and their meanings to obtain a sample set.

5. The fixed-point lightning warning method based on multi-source data according to claim 4, characterized in that: The variables included in a single sample include sample time, mean value, absolute value mean, absolute value maximum, difference value mean, difference value maximum, variance, standard deviation, number of reversals, whether to reverse, echo mean, 15km ratio, 30km ratio.

6. The method for fixed-point lightning warning based on multi-source data according to claim 1, characterized in that: The formula for performing sample factor correlation analysis on the sample set in Step 3 is: ; wherein, is the sample , is the correlation coefficient between, n is the sample size, , , are respectively the standard score, sample mean and sample standard deviation of the sample , , , are respectively the standard score, sample mean and sample standard deviation of the sample .

7. The method for fixed-point lightning warning based on multi-source data according to claim 1, wherein: The specific process of performing Z-Score standardization on the sample matrix in Step 4 to obtain a standardized matrix includes: Calculate the mean and standard deviation of each sample in the sample matrix and standard deviation SD ; Normalize according to the formula ; where is the value of a certain sample in the sample matrix.

8. The fixed-point lightning warning method based on multi-source data according to claim 1, wherein: In step 5, the principal component analysis method is used to perform a linear transformation on the standardized matrix, and the discrete factor of atmospheric electric field, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor are extracted, specifically including: Step 5.1: Establish the covariance matrix R of the standardized matrix; Step 5.2: Solve for the eigenvalues, main component contribution rates and cumulative variance contribution rates of the covariance matrix R, and determine the number of main components; Step 5.3: According to the principle of selecting the number of principal components, count the eigenvalues whose eigenvalues are greater than 1 and the cumulative contribution rate reaches more than 80% λ 1 , λ 2 , …, λ m and the corresponding m; Step 5.4: After performing a linear transformation with m unit eigenvectors as coefficients, an eigenvector matrix is obtained a T , and according to the formula calculate the first m sample principal components to obtain the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor; Step 5.5: Establish the initial factor loading matrix to interpret the principal components.

9. The method for fixed-point lightning warning based on multi-source data according to claim 8, wherein: The mathematical expressions of the discrete factor of atmospheric electric field, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor are: ; ; ; ; ; Among them, ~ are the discrete factor of atmospheric electric field, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor, respectively. ~ are the variables contained in a single sample, respectively.

10. The fixed-point lightning warning method based on multi-source data according to claim 1, characterized in that: In step 6, based on the five extracted principal components, a lightning warning model is established using the binary logistic regression method, and the training data set is input to train the lightning warning model, specifically including: Step 6.1: Using "whether there is lightning" as the dependent variable and the five principal components of the discrete factor of atmospheric electric field, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor as the covariates, introduce the sigmond function to establish the lightning warning model; Step 6.2: Input the training data set, estimate the parameters of the lightning warning model using the maximum likelihood estimation method, and train the lightning warning model using the gradient descent method and the quasi-Newton method.

Citation Information

Patent Citations

  • Thunder and lightning early warning data correction method based on real-time thunder and lighting positioning data correlation analysis

    CN105004932A

  • Multi-parameter and multi-algorithm integrated lightning refined monitoring and early warning algorithm

    CN114705922A