Fixed-point thunder and lightning early warning method based on multi-source data

By performing feature extraction and fusion analysis on multi-source lightning monitoring data, a lightning warning model was established, which solved the problem that the existing technology was difficult to achieve effective lightning warnings for point-like goals and small-scale goals, and achieved operational lightning warnings for lightning disaster-sensitive places, reducing losses and ensuring safety.

CN120178384AActive Publication Date: 2025-06-20CHONGQING LIGHTNING PROTECTION CENT
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510671704.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-20
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

It is difficult for the existing technology to achieve effective lightning warnings for point-based goals and small-scale goals, especially in places with lightning disaster-sensitive areas such as oil depots and scenic spots.

Method used

By performing feature extraction and fusion analysis on multi-source monitoring data such as atmospheric electric field data, new generation weather radar data, lightning positioning data, etc., a lightning warning model is established using principal component analysis and binary logistic regression method to achieve fixed-point lightning warning.

Benefits of technology

It provides operational and implementable lightning warning methods for lightning disaster-sensitive places such as oil depots, schools, scenic spots, industrial parks, and mines, reducing lightning disaster losses and ensuring production safety and life safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120178384A_ABST
    Figure CN120178384A_ABST
Patent Text Reader

Abstract

The invention discloses a fixed-point thunder and lightning early warning method based on multi-source data. The method comprises the following steps: collecting and preprocessing data; performing sample sorting on the multi-source data set to obtain a sample set; performing sample factor correlation analysis on the sample set to obtain a sample matrix; performing Z-Score standardization on the sample matrix to obtain a standardized matrix; carrying out linear transformation on the standardized matrix by adopting a principal component analysis method, and extracting principal components; based on the extracted principal component, establishing a thunder and lightning early warning model by using a binary logistic regression method, and inputting a training data set to train the thunder and lightning early warning model; and performing fixed-point lightning early warning by using the trained lightning early warning model. The method has the remarkable effects that an operable and practicable lightning early warning method is provided for lightning disaster sensitive places such as oil depots, schools, scenic spots, industrial parks and mines, and a scientific basis is provided for managers to take lightning prevention measures in advance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of lightning warning, and particularly relates to a fixed-point lightning warning method based on multi-source data. Background Art

[0002] Lightning is one of the most serious natural disasters, which greatly endangers people's lives and property. The harm of lightning strikes to some flammable and explosive places and lightning disaster-sensitive places such as oil depots, gas filling stations, fireworks and firecracker warehouses, schools, scenic spots, etc. is particularly serious. Once lightning strike explosions or casualty accidents occur, it will cause irreparable losses and great social impacts. Doing a good job in lightning monitoring and warning services is an important part of lightning disaster prevention work. However, lightning has the characteristics of instantaneousness, randomness, and suddenness in space and time. There are still certain difficulties in carrying out accurate and timely lightning forecasting and warning services, especially for point-like and small-scale targets such as oil depots and scenic spots. Therefore, improving the level of fixed-point lightning forecasting and warning is a difficult point in current scientific research work.

[0003] Lightning is a product of severe convective weather processes. The generation, dissipation, aggregation, and change laws of severe convective cloud clusters and the lightning generation mechanism of thunderclouds are very complex, resulting in the characteristics of instantaneousness, randomness, and suddenness in the spatial and temporal distribution of lightning. Early lightning nowcasting mainly used the extrapolation of weather radar and cloud images. Later, with the development of numerical prediction models and data assimilation technologies, the combination of extrapolation and high-resolution numerical prediction was proposed. More research tended to establish prediction equations or models by obtaining lightning observation data in real time, combining weather radar, meteorological satellites, sounding data, ground electric field data, lightning location data, etc., and using various mathematical methods to find the relationship between the observation data and thunderstorm phenomena.

[0004] In recent years, with the development of machine learning and deep learning technologies, various machine learning methods have been widely used in lightning warning research. For example, a lightning nowcasting model is constructed based on a neural network algorithm, or characteristic factors are extracted from multi-source data, and various machine learning methods are used to establish lightning forecasting models for comparative research. In addition, a deep semantic segmentation model is also used to fuse multi-source observation data to extract multi-source data forecasting factors to achieve lightning nowcasting. However, the research of domestic and foreign scholars mostly focuses on the lightning potential forecasting and lightning nowcasting on the surface or in a grid format, and the operability for point-like targets and small-scale targets is not strong enough.

[0005] In practical work, there is a relatively urgent need and actual application scenario for short-term fixed-point lightning warnings in lightning disaster-sensitive sites such as flammable and explosive places, scenic spots, mines, and industrial parks. Therefore, if feature extraction and fusion analysis are carried out on multi-source monitoring data such as atmospheric electric field data, new-generation weather radar data, and lightning location data, and a fixed-point lightning warning method based on multi-source data is explored, and a corresponding warning model is established, so as to provide an operable and implementable lightning warning method for lightning disaster-sensitive sites such as oil depots, schools, scenic spots, industrial parks, and mines, and provide a scientific basis for managers to take lightning prevention measures in advance, it will be of great significance in reducing lightning disaster losses, ensuring enterprise production safety, and protecting people's lives. Summary of the Invention

[0006] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a fixed-point lightning warning method based on multi-source data. By performing feature extraction and fusion analysis on multi-source monitoring data such as atmospheric electric field data, new-generation weather radar data, and lightning location data, a lightning warning model is constructed for fixed-point lightning warning, which can solve the technical defect that the operability for point targets and small-range targets in the background technology is not strong enough.

[0007] To achieve the above object, the technical solution adopted by the present invention is as follows: The present invention proposes a fixed-point lightning warning method based on multi-source data, which is characterized in that it includes the steps: Step 1, data collection and preprocessing to obtain a multi-source data set including atmospheric electric field data, lightning location data, and weather radar data; Step 2, sorting the samples of the multi-source data set to obtain a sample set; Step 3, performing sample factor correlation analysis on the sample set to obtain a sample matrix; Step 4, performing Z-Score standardization on the sample matrix to obtain a standardized matrix; Step 5, using the principal component analysis method to perform a linear transformation on the standardized matrix, and extracting five principal components: the atmospheric electric field discrete factor, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor; Step 6, based on the five principal components of the atmospheric electric field discrete factor, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor extracted, using the binary logistic regression method to establish a lightning warning model, and inputting the training data set to train the lightning warning model; Step 7, using the trained lightning warning model for fixed-point lightning warning; The expression of the trained lightning warning model is: ; Among them, ~ are respectively the atmospheric electric field discrete factor, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor, p is the probability of lightning occurrence.

[0008] Furthermore, the data preprocessing described in step 1 includes the following processes: Step 1.1: According to the lightning warning position and the defined coverage area range, divide the collected multi-source data into data coverage areas to achieve preliminary data screening; Step 1.2: For the multi-source data that has undergone preliminary data screening, use different methods for data processing according to the data type to obtain a multi-source data set. Among them: the atmospheric electric field data is processed using the mean imputation method; the weather radar data is processed using the bilinear interpolation algorithm.

[0009] Furthermore, the processing process of the lightning location data in step 1.2 is as follows: Determine the distance between the target point and the lightning location points according to the longitude and latitude of the target point, screen out all lightning location data with a distance less than the preset distance threshold from the target point, and form a lightning location data set for the target area after sorting by time; Eliminate the lightning location data with a lightning current value less than the preset lightning current amplitude in the lightning location data set for the target area; According to the positioning principle of the lightning locator, eliminate the lightning location data with a positioning method lower than three stations in the lightning location data set.

[0010] Furthermore, the sample set obtained by sample sorting in step 2 specifically includes: Determine the sample elements; Determine the sample recording time and the time interval between each group of samples; Determine the variables included in a single sample and their representative meanings; Sort out the data in the multi-source data set according to the sample elements, sample recording time, time interval between each group of samples, variables included in a single sample and their representative meanings to obtain a sample set.

[0011] Furthermore, the variables included in a single sample include sample time, mean value, absolute value mean, absolute value maximum, difference value mean, difference value maximum, variance, standard deviation, number of reversals, whether reversed, echo mean, 15km ratio, 30km ratio.

[0012] Furthermore, the formula for sample factor correlation analysis of the sample set in step 3 is: ; Among them, is the sample , , where n is the sample size, , , are respectively the standard score, sample mean and sample standard deviation of the sample ; , , are respectively the standard score, sample mean and sample standard deviation of the sample .

[0013] Furthermore, the specific steps of obtaining the standardized matrix by performing Z-Score standardization on the sample matrix in step 4 include: Calculating the mean value of each sample in the sample matrix and the standard deviation SD ; Performing standardization according to the formula ; where is the value of a certain sample in the sample matrix.

[0014] Furthermore, in step 5, the principal component analysis method is used to perform a linear transformation on the standardized matrix, and the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor are extracted. Specifically, it includes: Step 5.1: Establish the covariance matrix R of the standardized matrix; Step 5.2: Solve the eigenvalues, principal component contribution rates and cumulative variance contribution rates of the covariance matrix R to determine the number of principal components; Step 5.3: According to the principle of selecting the number of principal components, count the eigenvalues whose eigenvalues are greater than 1 and the cumulative contribution rate reaches more than 80% λ 1 , λ 2 , …, λ m and the corresponding m; Step 5.4: Perform a linear transformation with m unit eigenvectors as coefficients to obtain the eigenvector matrix a T , and calculate the first m sample principal components according to the formula to obtain the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor; Step 5.5: Establish the initial factor loading matrix and interpret the principal components.

[0015] Further, the mathematical expressions of the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor are as follows: ; ; ; ; ; Among them, ~ are the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor respectively, ~ are the variables included in a single sample respectively.

[0016] Further, in step 6, based on the five principal components extracted, a lightning warning model is established using the binary logistic regression method, and the training data set is input to train the lightning warning model, which specifically includes: Step 6.1: Taking "whether there is lightning" as the dependent variable and the five principal components of the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor as the covariates, introducing the sigmond function to establish the lightning warning model; Step 6.2: Input the training data set, estimate the parameters of the lightning warning model using the maximum likelihood estimation method, and train the lightning warning model using the gradient descent method and quasi-Newton method.

[0017] The remarkable effects of the present invention are as follows: 1. Through feature extraction and fusion analysis of multi-source lightning monitoring data such as atmospheric electric field data, new generation weather radar data, and lightning location data, and performing data quality control such as interpolation and cleaning, a thunderstorm sample set and a non-thunderstorm sample set are formed. Then, using statistical methods to analyze the relationship between multiple warning factors and lightning occurrence results, correlation analysis, Z-Score standardization, principal component analysis (PCA), binary Logistic regression, etc. are performed on the sample data, establishing the method and standard for extracting principal component factors, and extracting the first to fifth principal components as the modeling factors of the Logistic regression model to establish a fixed-point lightning warning model based on multi-source data, providing an operable and implementable lightning warning method for lightning disaster-sensitive places such as oil depots, schools, scenic spots, industrial parks, and mines, providing a scientific basis for managers to take lightning prevention measures in advance, and having important significance in reducing lightning disaster losses, ensuring enterprise production safety, and protecting people's lives.

[0018] 2. Thirteen sample factors that are all positively correlated with "whether lightning" are sorted out. Considering factors such as eigenvalue, single-item variance percentage, scree plot test, and cumulative variance percentage, multi-source data of the thirteen factors are fused to extract a total of five principal components. Certain meanings are assigned to these five principal components according to the factor loading coefficients and named. Among them, the first principal component is the atmospheric electric field discrete factor, the second principal component is the radar echo intensity factor, the third principal component is the atmospheric electric field reversal factor, the fourth principal component is the atmospheric electric field intensity factor, and the fifth principal component is the radar echo extreme value factor. Thus, a binary Logistic regression model is established. From various test indexes such as the classification result, performance indexes, and evaluation indexes of the model, the model is effective and has a good fitting degree. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is the flow chart of the method described in the present invention; Figure 2 is the schematic diagram of regional division; Figure 3 is the trend chart of the correlation coefficient between each sample factor and "whether lightning" in the present invention changing with the lead time; Figure 4 is the heat map of the correlation between each factor and "whether lightning"; Figure 5 is the scree plot of the eigenvalues of the principal components; Figure 6 is the schematic diagram of the establishment process of the lightning warning model; Figure 7 is the trend chart of the performance parameters of the binary Logistic regression model changing with the lead time t when the cut-off value is 0.5; Figure 8 is the trend chart of the performance parameters of the binary Logistic regression model changing with the lead time t when the cut-off value is 0.3; Figure 9 is the F1 value changing with the lead time t when the cut-off values of the model are 0.2 to 0.7; Figure 10 is the principle block diagram of the system described in the present invention; Figure 11 is the principle block diagram of the computer device described in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The following further elaborates in detail the specific embodiments and working principles of the present invention with reference to the drawings. Embodiment

[0021] As Figure 1As shown in the figure, this embodiment provides a fixed-point lightning warning method based on multi-source data. The specific steps are as follows: Step 1: Data collection and preprocessing to obtain a multi-source data set including atmospheric electric field data, lightning location data, and weather radar data; In this example, the atmospheric electric field data is from the data continuously collected by the Pre-storm2.0 field mill type atmospheric electric field instrument installed in the outdoor field of the Jinfo Mountain Lightning Special Experiment from 2020 to 2022. The basic working principle of this electric field instrument is that the motor drives the shielding metal sheet to rotate, making the induction metal sheet alternately exposed to the electric field or shielded, thereby generating induced charges proportional to the external electric field intensity. The induction sheet is connected to an amplification and processing circuit and a waveform adjustment circuit, and outputs a voltage signal. After calibration, this voltage can represent the intensity and polarity changes of the atmospheric electric field. Its main technical parameters are shown in Table 1.

[0022]

[0023] In this example, the lightning location data is from the location data of the Chongqing ADTD Lightning Location System from 2020 to 2022. This system consists of a second-generation ADTD lightning locator, a central data processing station, a user data service network, and a graphic display terminal. The observation network of the system consists of 5 monitoring stations in Chongqing and 10 monitoring stations in neighboring provinces. The system realizes continuous automatic monitoring of elements such as the time, location (longitude, latitude), peak lightning current, and polarity of cloud-to-ground lightning. The cloud-to-ground lightning detection efficiency is above 85%, the detection and positioning accuracy within the network is less than 300 m, and the processing time of lightning strikes is about 1 ms.

[0024] The weather radar data uses the composite reflectivity factor mosaic product of a new generation of Doppler weather radar, which is from the Chongqing Meteorological Observatory. The time resolution is 6 min, and the spatial resolution is 0.01°×0.01° (about 1 km×1 km). The data format is a png image file, and the size of the image is 1400×1200 pixels, that is, the image coverage area is approximately (100°E, 23°N)~(114°E, 35°N). The echo intensity of the pixel point in the image is the gray level of this point. Since the file is an 8-bit gray-scale image, there are a total of 256 gray levels.

[0025] It should be noted that in addition to the above data types, it may also include other types of lightning observation data such as trial lightning data, meteorological sounding data, and meteorological satellite data.

[0026] In some specific embodiments, the data preprocessing in Step 1 includes the following processes: Step 1.1: According to the lightning warning location and the defined coverage area range, divide the collected multi-source data into data coverage areas to achieve preliminary data screening; Observing the lightning observation data of different data types obtained, it can be seen that except for the atmospheric electric field data which is point data, the rest of the data are surface data. In order to preliminarily screen the data range, it is necessary to divide the data coverage area first. Referring to the division of the target area, surrounding area, monitoring area, and coverage area in "Lightning Protection Thunderstorm Warning System" (GB / T 38121—2023), the following areas are defined in this embodiment, and the data is preliminarily screened accordingly: Target area: The actual geographical range where lightning warning needs to be carried out, which can be a single location such as an oil depot, school, building, etc., or can be extended to a larger area such as a scenic area, chemical industrial park, etc. In this embodiment, it refers to the outdoor field of the special lightning experiment on Jinfo Mountain (latitude: 28.9451N, longitude: 107.1462E).

[0027] Monitoring area: The area where an alarm needs to be triggered when lightning is approaching or occurring. In this embodiment, it refers to the area included within a circle with a radius of 5 km centered on the outdoor field of the special lightning experiment on Jinfo Mountain.

[0028] Coverage area: With the outdoor field of the special lightning experiment on Jinfo Mountain as the center, the radius of the coverage area is set to 15 km. The main considerations are as follows: (1) For local detectors / atmospheric electric field meters, the electric field of thunderstorm clouds within 15 km has a greater impact on the electric field at the location of the atmospheric electric field meter in the target area; (2) When the high-altitude wind speed is 2 km / min, the time for a thunderstorm cloud at a distance of 15 km to move above the monitoring area is about 5 - 6 minutes, which is equivalent to the volume scan period of the new generation weather radar.

[0029] The schematic diagram of the area division is as Figure 2 shown. According to the area division results, the collected multi-source data is divided according to the data coverage area to achieve preliminary data screening.

[0030] Step 1.2: For the multi-source data that has undergone preliminary data screening, different methods are used for data processing according to the data type to obtain a multi-source data set, where: 1) Atmospheric electric field data processing:

[0031] For atmospheric electric field data, network interruption and delay may both cause the loss of data packets. The processing of missing values is carried out according to the following method: when data loss is detected, statistical analysis is performed on the data of 1 minute, which is 30 seconds before and 30 seconds after the missing bit, to obtain the average value, and the missing bit is filled with the average value, that is, the mean interpolation method. For the situation of long-term data loss caused by equipment failure, etc., interpolation processing is no longer carried out, and such time periods are not used as the source time for sample extraction. 2) Lightning location data processing:

[0032] (1) Distance screening Set the installation location of the atmospheric electric field instrument as the target point. Then, lightning occurring near this point may cause lightning strike damage to the protected objects within the target area. Assume that lightning occurring within a circle with this point as the center and a radius of r is regarded as the object to be pre-warned. Then, the following considerations are taken: If r takes a too large value, the possibility of lightning far from the target area causing harm to the protected objects is very small. At the same time, too many pre-warning objects lead to an excessive number of alarms, which is not conducive to the development of lightning protection work; if r takes a too small value, not only will there be missed reports, but also fewer positioning data will result in a reduction in the selected thunderstorm samples, which is not conducive to the construction of the model in subsequent work. After repeated screening and comparison, in this embodiment, r is set to 5 km.

[0033] The distance calculation formula between the target point determined according to longitude and latitude and the lightning location point is: (1) Among them, lat 1 、 lon 1 、 lat 2 、 lon 2 are the latitude, longitude, latitude, and longitude of the lightning location point and the target point respectively. Use the above formula to screen out all lightning location data with a distance less than 5 km from the target point and sort them by time to form the lightning location data set of the target area.

[0034] (2)Noise and interference data processing Since the ADTD lightning location system mainly uses the time difference of the electromagnetic waves radiated by each ground flash return stroke reaching each station for positioning, in the current complex electromagnetic environment, it is inevitable to introduce noise. At the same time, due to various factors such as cloud flash interference, false ground flash location data will be generated, and its characteristic is that the current value in the location data is too small. Such data should be excluded. The processing of location data with too small lightning current values in the lightning location data set follows the following principle: Exclude lightning location data with a lightning current amplitude of 0 - 2 kA. The "Technical Guide for Lightning Disaster Risk Zoning" (QX / T 405—2017) also points out that data above 200 kA should be excluded. However, in this embodiment, considering that the target area is above 2000 meters above sea level and the possibility of high-intensity ground flashes is relatively large. In addition, it is also found in the statistics that the overall proportion of data with a lightning current amplitude above 200 kA is very small and has a relatively minor impact on the research results, so it is not excluded.

[0035] (3)Deletion of unreliable data According to the positioning principle of the ADTD lightning location system, the two-station positioning uses the magnetic orientation intersection method, which is affected by terrain interference and results in poor positioning accuracy. Considering the topographical and geomorphic characteristics of Chongqing, which are mainly mountainous and hilly, all records in the lightning location dataset with a positioning method lower than three stations are deleted. 3) Weather radar data processing

[0036] Regarding the missing frame phenomenon in weather radar data, considering that the echo generally changes continuously and the change between adjacent two frames is small, in order to improve the utilization rate of data, the missing data is filled by bilinear interpolation. The bilinear interpolation algorithm is a commonly used one in image interpolation algorithms, that is, using the radar echoes at T-1 and T+1 moments, the radar echo at T moment is calculated by the bilinear interpolation method. Its basic principle is calculated through the weighted sum of the 4 nearest neighbor pixel values of the pixel to be solved in the source image.

[0037] Use Matlab software for image recognition and process the mosaic products into element data such as the mean echo intensity and maximum value of the monitoring area, and the proportion of echoes exceeding 35 dBz within the coverage area. In addition, for reference and comparison, the proportion of echoes exceeding 35 dBz within 30 km is also statistically counted as one of the attributes of the sample. 4) Other data processing measures

[0038] Only use those objects with all attribute values. Since the various data sources are different, it is impossible to ensure the consistency of data in the recording time. For example, due to equipment and network failures, the atmospheric electric field instrument has data missing for a relatively long period, and this kind of missing cannot be filled by interpolation, while the radar data may be relatively complete during this corresponding period. Therefore, for the samples sorted out, there must be some missing attribute values. For such situations, the corresponding samples are all deleted, and only the samples with all attribute values are retained.

[0039] Through the above data preprocessing and quality control, useless data can be effectively reduced, which helps to improve the accuracy of fixed-point lightning warning.

[0040] Step 2: Sort out the multi-source dataset to obtain a sample set; In some embodiments, the specific implementation process of the sample data is as follows: Step 2.1: Determine the basic principles for sorting out thunderstorm weather samples; To avoid frequent switching of the warning status, the thunderstorm warning system can use the alarm retention time to set the minimum time for maintaining the alarm. If the retention time is too short, multiple alarms may be issued during a single thunderstorm process or the alarm status may switch frequently. If the retention time is too long, the alarm may not be lifted in a timely manner. In this embodiment, the retention time is set to 30 min, and based on this, the lightning location data is grouped: If the time difference between the recording time of a lightning location data record and the recording time of the adjacent data record behind it is less than 30 min, then these two pieces of data are regarded as being in the same group. A group formed in this way is recorded as a thunderstorm process, and the time of the first lightning occurrence in each process is used as the reference time. According to the above principles, the lightning location data set in the target area from 2020 to 2022 is processed, and a total of 105 thunderstorm processes are sorted out.

[0041] Step 2.2: Determine the sample elements, sample recording time, time interval between each group of samples, the variables included in a single sample, and their representative meanings; For fixed-point lightning warning, the performance of various lightning-related monitoring data within a period of time before the first lightning occurrence in each thunderstorm process is the object of study in this embodiment. Assume that the time of the first lightning occurrence is T l , and the warning lead time is t lead , then the sample recording time is T s = T l - t lead . Since the volume scan period of the new generation weather radar is 6 min, the time interval between each group of samples should be 6 min. A single sample includes the following variables: Sample recording time T s ; Atmospheric electric field strength value and its partial statistics; New generation weather radar echo intensity and its partial statistics; Data from other sources. The variables included in a single sample and their representative meanings are shown in Table 2:

[0042] Step 2.3: Sort out the data in the multi-source data set according to the sample elements, sample recording time, time interval between each group of samples, the variables included in a single sample, and their representative meanings to obtain a sample set. For non-thunderstorm weather processes, this embodiment conducts statistics according to the following principles: ① There is no lightning occurrence in both the target area and the monitoring area during the statistical period; ② There is precipitation in the target area during the statistical period and the precipitation is greater than or equal to 5 mm.

[0043] Step 3: Conduct a correlation analysis of sample factors on the sample set to obtain a sample matrix; In some embodiments, the specific implementation process of performing sample factor correlation analysis on a sample set is as follows: Correlation analysis refers to analyzing two or more correlated variables to measure the degree of correlation between the two variables. There needs to be a certain connection or probability between the variables for correlation analysis to be performed.

[0044] Performing multivariate correlation analysis on the sorted sample factors can initially understand whether there is a relationship between the factors, between the factors and the "whether lightning" result, the strength and direction of the relationship, and whether this relationship is statistically significant, etc. In addition, since principal component analysis requires a certain correlation between the factors, it is necessary to perform correlation analysis on the sample factors before the modeling work.

[0045] The Pearson correlation coefficient is widely used to measure the degree of correlation between two variables, and its value ranges from -1 to 1. The Pearson correlation coefficient between two variables is defined as the quotient of the covariance and standard deviation between the two variables, so this correlation coefficient is also called the "Pearson product-moment correlation coefficient": (2) For the sample, the equivalent expression is as follows: (3) Among them, is the correlation coefficient between samples , , n is the sample size, , , are the standardized scores, sample mean, and sample standard deviation of sample respectively, , , are the standardized scores, sample mean, and sample standard deviation of sample respectively.

[0046] The trend of the correlation coefficient between each factor and "whether lightning" with the lead time is as Figure 3 , and the correlation heat map (t lead = 18) is as Figure 4 . It can be seen that: (1) The correlation coefficient between each factor of the atmospheric electric field and "whether lightning" is significantly greater than the correlation coefficient between each factor of the radar and "whether lightning", indicating that the correlation between each factor of the atmospheric electric field and "whether lightning" is stronger.

[0047] (2) Except for the "mean value", the correlation coefficient between the other factors of the atmospheric electric field and "whether lightning" changes with t leadshows a downward trend as it increases, and the correlation coefficient between each factor of the radar and "whether there is lightning" varies with t lead remains basically unchanged as it increases, indicating that as time approaches, the relationship between each factor of the atmospheric electric field and "whether there is lightning" becomes stronger. However, since the atmospheric electric field has positive and negative values and the mean value changes greatly within one hour, the correlation between the mean value and "whether there is lightning" does not show an obvious increasing or decreasing trend over time, but rather shows a large degree of randomness. Within 1 hour, the relationship between each factor of the radar and "whether there is lightning" will not change significantly over time; (3)The several factors with the strongest correlations are "whether it reverses", "absolute value mean", "standard deviation", "absolute value maximum", "maximum difference value", and "mean difference value".

[0048] In summary, there are correlations of varying degrees among the factors, and the vast majority of factors are positively correlated (confidence level 0.95). Among them, the correlations among the factors of the atmospheric electric field and among the factors of the radar are relatively high, but the correlation between the factors of the atmospheric electric field and the factors of the radar is weak, and there is a negative correlation between individual factors.

[0049] Step 4: Perform Z-Score standardization on the sample matrix to obtain a standardized matrix; In the specific implementation process, since the magnitudes and units of the factors of the sample will cause certain difficulties to the analysis work or seriously affect the PCA output, thus affecting the accuracy of the later modeling. To eliminate this influence, data standardization is required. This embodiment adopts the Z-Score standardization method, and the specific calculation method is as follows: Calculate the mean value of the factor and the standard deviation SD , and then subtract the mean value from each value of the factor and divide by the standard deviation, that is: (4) After Z-Score standardization, the data is closer to the standard normal distribution, the mean value of the variable is 0, and the standard deviation is 1. Through the transformation, multiple groups of data are converted into unitless Z-Score values, making the data standardized and improving the comparability of the data.

[0050] Step 5: Use the principal component analysis method to perform a linear transformation on the standardized matrix, and extract five principal components: the discrete factor of the atmospheric electric field, the echo intensity factor of the radar, the reversal factor of the atmospheric electric field, the intensity factor of the atmospheric electric field, and the extreme value factor of the radar echo; From the analysis results of the sample factors, it can be seen that the number of sample factors (i.e., variables) is relatively large, belonging to a high-dimensional space in terms of data structure. In this way, the difficulty of analyzing and modeling numerous factors increases sharply with the increase in the number of factors. At the same time, the correlation between sample factors is relatively high (the problem of multicollinearity), and the information redundancy is relatively high. In addition, when a serious collinearity problem occurs, it will lead to unstable analysis results, and the sign of the regression coefficient is completely opposite to the actual situation. At this time, it involves data dimensionality reduction and elimination of the influence of multicollinearity. Data dimensionality reduction refers to the process of mapping high-dimensional data to a low-dimensional space, and its purpose is to reduce the dimensionality of data while trying to retain the internal structure of the data. The general idea is to find the internal structure of the data and compress the original high-dimensional data into a low-dimensional space through techniques such as projection, coding, or reconstruction, so as to facilitate data processing, analysis, and visualization.

[0051] Principal Component Analysis (PCA) is an attribute aggregation technique proposed by Karl Pearson in 1901. The basic idea is to map high-dimensional data to the principal component subspace through orthogonal transformation, retain the main components, and remove the secondary components. Usually, several new variables that are fewer than the number of original variables and can explain most of the differences in the data are selected, that is, the so-called principal components. For the lightning warning sample data, the main functions of principal component analysis are: First, reduce the dimensionality of the data and simplify the complexity of the subsequent modeling process; Second, the newly constructed variables are not correlated with each other, which is convenient for the next step of binary logistic regression modeling. The specific calculation steps are as follows: Step 5.1: Use the data in the standardization matrix to establish the covariance matrix R; Covariance is a statistical index that reflects the degree of closeness of the correlation between standardized factors. The larger the value, the more necessary it is to perform principal component analysis on the data. Among them, R ij (i, j = 1, 2,..., p) is the correlation coefficient between factor X i and X j . According to the calculation method of the correlation coefficient, R is a real symmetric matrix (i.e., R ij = R ji ), and only the upper triangular elements or lower triangular elements need to be calculated. Its calculation formula is: (5) Step 5.2: Solve the eigenvalues, principal component contribution rates and cumulative variance contribution rates of the covariance matrix R to determine the number of principal components; Solve the characteristic equation to find the eigenvalues λ i (i = 1, 2,..., p) and the corresponding eigenvectors ai Since R is a positive definite matrix, its eigenvalues λ i are all positive numbers. Arrange them in descending order, that is λ 1 ≥ λ 2 ≥ … ≥ λ i ≥ 0. The eigenvalues are the variances of the principal components, and their magnitudes reflect the influence of each principal component. Then the contribution rate of the i th principal component is: (6) The cumulative contribution rate is: (7) Step 5.3: According to the principle of selecting the number of principal components, count the eigenvalues whose eigenvalues are greater than 1 and the cumulative contribution rate reaches more than 80% λ 1 , λ 2 , …, λ m and the corresponding 1, 2, …, m (m ≤ p), where m is the number of principal components; Step 5.4: After performing a linear transformation with m unit eigenvectors as coefficients, obtain the eigenvector matrix a T , and calculate the first m sample principal components to obtain the discrete factor of the atmospheric electric field, the radar echo intensity factor, the reverse factor of the atmospheric electric field, the atmospheric electric field intensity factor, and the radar echo extreme value factor; (8) Step 5.5: Establish the initial factor loading matrix and interpret the principal components: (9) Next, take the sample with t lead = 18 as an example, and use SPSS to perform the principal component analysis steps: 1) Set "variables", and set all standardized Z - Score variables as "variables"; 2) Factor analysis: Description, select "KMO and Bartlett's sphericity test"; 3) Factor analysis: Extraction, select "Principal components", extract eigenvalues greater than 1, and display the "scree plot"; 4) Factor analysis: Rotation, select "Varimax method" for the method, and display the "rotated solution"; 5) Check whether the cumulative rate C of the "extracted sum of squared loadings" in the "Total Variance Explained" table is greater than or equal to 80%. If C < 80%, go back to step 3) and make corresponding adjustments to the size of the extracted eigenvalues. 6) Based on the factor analysis results, define the same number of new variables, and input the factor loadings in the "Component Matrix" into the new variables defined in the new data file respectively. 7) Use "Transform - Variable" to calculate the characteristic variables and further obtain the eigenvector matrix. 8) Further calculate the principal components from the eigenvector matrix.

[0052] Perform the KMO (Kaiser - Meyer - Olkin) and Bartlett's test of sphericity on the obtained principal components. The results are shown in Table 3:

[0053] The null hypothesis of Bartlett's test of sphericity is that the correlation coefficient matrix is an identity matrix. Its approximate chi - square statistic is 4969.168, and the p - value 0.000 < 0.001. Therefore, the null hypothesis is rejected, indicating that there is a correlation relationship among the variables and it is suitable for factor analysis. In addition, the KMO test statistic is a sampling adequacy test proposed by Kaiser, Meyer, and Olkin, which is an index for comparing the simple correlation coefficients and partial correlation coefficients among variables, and its value ranges between 0 and 1.

[0054] Generally speaking, the closer the KMO value is to 1, the stronger the correlation among the variables, and the more suitable the original variables are for factor analysis. It can be seen that after the original data is standardized by Z - Score, the KMO test statistic is 0.728, the data structure is reasonable, and it is suitable for factor analysis.

[0055] The communalities are shown in Table 4. It can be seen that the minimum value of the extracted communalities is 0.774, indicating that after the Z - Score standardized data undergoes a principal component transformation, more information of the original variables can be extracted, and the common factors have a good expression for the Z - Score variables.

[0056]

[0057] The total variance explanation is shown in Table 5. Among them, the left part is the initial eigenvalue, the middle is the result of extracting the main factors, and the right is the result of factor analysis after rotation. It can be seen that there are 13 main factors after the transformation of 13 factors. Among them, the eigenvalue of the first main factor is 5.826, accounting for 45.095% of the total variance. Similarly, the eigenvalues of the second to fifth main factors are 2.649, 1.192, 0.947, and 0.84 in turn. Generally speaking, the cumulative variance contribution rate reflects the proportion of information retained by the main factors. The cumulative variance percentage of the first five main factors in Table 5 is 88.393%, which can reflect most of the information of the original data.

[0058]

[0059] The number of main factors in the factor analysis result determines the number of main components in the principal component analysis. The number of main factors to be extracted can be determined according to one or several of the following methods: (1)The eigenvalue is greater than 1; (2)The single-variance percentage is greater than 5%; (3)Using the scree plot test; (4)The cumulative variance percentage is greater than 80%.

[0060] Although the eigenvalues of the fourth and fifth main factors are slightly less than 1, the single-variance percentages are 7.286% and 6.461% respectively, both of which are greater than 5%. The cumulative variance percentage of the first to fifth main factors reaches 88.393%. Considering the above extraction principles comprehensively, a total of the first five main factors are extracted. At this time, less information of the original data is lost, and the effect of the principal component analysis is relatively ideal and has research significance. Figure 5 It is the scree plot of the eigenvalues. This figure shows the steep slope of the larger factors and the gentle tails of the remaining factors. Generally, the main factors are selected on the steep slope, while the factors on the gentle slope have very little explanation for the variation. It can be seen that the curve tends to be gentle after the fifth main factor, which also shows that the first five main components should be extracted.

[0061] Analyzing the rotated component matrix shown in Table 6, in the first principal component, the load factor values of five factors, namely the maximum absolute value, the mean of the difference values, the maximum of the difference values, the variance, and the standard deviation, are relatively large, indicating that the first principal component mainly explains these five factors; in the second principal component, the load factor values of three factors, namely the mean of echo, the proportion at 15 km, and the proportion at 30 km, are relatively large; in the third principal component, the load factor values of two factors, namely the number of reversals and whether there is a reversal, are relatively large; in the fourth principal component, the load factor value of the mean is relatively large; in the fifth principal component, the load factor value of the maximum of echo is relatively large. In the table, the data with factor load values exceeding 0.7 are all shown in bold. According to the factor load coefficients, certain meanings are assigned to these five principal components and named. For example, the first principal component is the atmospheric electric field discrete factor, the second principal component is the radar echo intensity factor, the third principal component is the atmospheric electric field reversal factor, the fourth principal component is the atmospheric electric field intensity factor, and the fifth principal component is the radar echo extreme value factor.

[0062]

[0063] The calculation of the eigenvector is carried out according to the following formula: (10) where is the i th column vector in the component matrix, and is the eigenvalue corresponding to the i th principal component.

[0064] After "Transformation - Calculation", the eigenvector matrix a T is as follows:

[0065] Then, each principal component can be calculated, that is, the eigenvector matrix is multiplied by the standardized sample matrix. The mathematical expressions of the atmospheric electric field discrete factor, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor are as follows: ; ; ; ; ; Step 6: Based on the five principal components of the extracted atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor, use the binary logistic regression method to establish a lightning warning model, and input the training data set to train the lightning warning model; Through principal component analysis, the original 13 factors (variables) were expressed by 5 principal components, which are: atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor. Principal component analysis realizes the dimensionality reduction of data, and the obtained component factors are not related to each other. The next step is to use binary Logistic regression to build a model for fixed-point lightning warning. The specific process is as follows: Step 6.1: Take "whether there is lightning" as the dependent variable, and take the five principal components of atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor as covariates, and introduce the sigmond function to establish a lightning warning model; The value "1" represents "lightning occurs", and the value "0" represents "no lightning occurs". If represents the probability of lightning occurrence, when its value is greater than or equal to 0.5, it is predicted that lightning will occur, and when it is less than 0.5, it is predicted that no lightning will occur, then there is: (11) Introduce the sigmond function as follows: (12) Then the mathematical model of the lightning warning model can be constructed as follows: (13) In formula (13), obeys the binary Logistic distribution, p represents the probability of lightning occurrence, (1 - p ) represents the probability of no lightning occurrence, β is the parameter to be estimated, y is the input variable. In this embodiment, y refers to the principal component extracted after standardization and linear transformation. It can be seen from the above formula that the Logistic model establishes the relationship between the probability of lightning occurrence and the principal component (explanatory variable). The model is shown as Figure 6 shown.

[0066] Step 6.2: Input the training data set , and use the maximum likelihood estimation method to estimate the parameters β of the lightning warning model, so as to obtain the Logistic regression model of the lightning warning model. Its log-likelihood function is: (14) Find the maximum value of to obtain βThe estimated value. In this way, the problem becomes an optimization problem with the log-likelihood function as the objective function. The methods commonly used in Logistic regression learning are the gradient descent method and the quasi-Newton method.

[0067] After obtaining the lightning warning model, this embodiment also evaluated its performance, specifically as follows: The evaluation of the performance of the lightning warning model mainly uses three indicators: the effective alarm rate, the missed alarm rate, and the false alarm rate. The higher the effective alarm rate and the lower the missed alarm rate and false alarm rate, the better the performance of the model. For a classification model, these three indicators are all related to the classification accuracy. However, only when the class distribution in the classification is equal, the accuracy is a useful indicator. This means that under the condition of data imbalance, that is, when the data of a certain classification in the sample is significantly more than that of another classification, the accuracy cannot fully reflect the performance of the model, resulting in the fact that these three indicators of the effective alarm rate, the missed alarm rate, and the false alarm rate cannot very accurately reflect the advantages and disadvantages of the model.

[0068] One of the methods to solve the imbalance problem is to use better indicators. Here, the F1 score is introduced. It not only considers the number of errors predicted by the model but also considers the type of errors. As a performance measure of a classifier, the F1 score is defined as the harmonic mean of precision and recall, and its calculation formula is as follows: (15) The value of the F1 score ranges from 0 to 1. Generally speaking, the larger the F1 value, the better the performance of the classifier. And only when both the FAR and FTWR are close to 0 at the same time, the F1 score approaches the maximum value of 1.

[0069] Take t = 18 as an example. The results of establishing a binary Logistic regression model using SPSS are as follows:

[0070]

[0071] It can be seen from the Omnibus test of the model coefficients that the corresponding p-value is less than 0.001, which is less than the significance level of 0.05, indicating that the model is effective and has a good fitting degree. In addition, the values of Cox & Snell R Square and Nagelkerke R Square in Table 8 range from 0 to 1, and the larger the value, the better the fitting degree of the model. Here, they are 0.407 and 0.674 respectively, indicating that the model has a good fitting degree.

[0072] To obtain the correct classification rate of an object through a classification model, it must be better than the correct rate obtained by placing all objects in the class with the largest number of objects (the majority class, i.e., no thunderstorm occurred). From the classification result a shown in Table 9, the classification accuracy rate for no thunderstorm occurred is 99.3%, the classification accuracy rate for thunderstorm occurred is 71.2%, and the overall percentage is 94.4%, which is better than the overall percentage of 82.5% when classifying all objects with thunderstorm occurred as no thunderstorm occurred (classification result b shown in Table 10). The model has improved the classification accuracy rate by nearly 12%, indicating that the model is effective.

[0073]

[0074]

[0075] The Wald chi-square value and the significance level p are the hypothesis tests for the regression coefficient B value. In Table 11, the significance p values are all less than 0.05, indicating that the influence of all principal component factors on the result is statistically significant.

[0076]

[0077] Therefore, the final expression of the lightning warning model is: (16) Wherein, ~ are the discrete factor of atmospheric electric field, the radar echo intensity factor, the reverse factor of atmospheric electric field, the atmospheric electric field intensity factor, and the extreme value factor of radar echo respectively, p is the probability of lightning occurrence.

[0078] Through a series of data analysis and processing such as data standardization, principal component analysis, and binary Logistic regression, a function expression for the probability p of thunderstorm occurrence is finally obtained. The dependent variable of the function is the first to fifth principal components. The principal component analysis - binary Logistic regression model is as Figure 6 . In this model, the data standardization and principal component analysis processes are carried out for the entire sample. With the change of the sample size and sample values, the expression of the model will also change accordingly. As a comparison, Figure 7 、 Figure 8 respectively give the change trends of the performance parameters of the binary Logistic regression model with the lead time t when the cut-off values are 0.5 and 0.3, and the following conclusions are obtained: 1) There is no significant difference in the effective alarm rates of the two models, and basically both are above 90%; 2) When the threshold value is 0.5, the false alarm rate of the model is relatively low and does not exceed 7.41% at most. However, as t increases, the miss rate increases rapidly. When t = 54, the miss rate is as high as 55.36%. When the threshold value is 0.3, a relatively low false alarm rate can be maintained. For example, when t ≤ 48, it is basically below 10%. 3) When t ≤ 18, both the miss rate and the false alarm rate of the model with a threshold value of 0.3 are less than those of the model with a threshold value of 0.5. At this time, the performance of the model with a threshold value of 0.3 is better. 4) When t ≥ 24, the two models have their own focuses. The miss rate of the model with a threshold value of 0.3 is about 10% lower than that of the model with a threshold value of 0.5, but the false alarm rate is significantly higher than that of the model with a threshold value of 0.5.

[0079] Due to the imbalance of sample classification data, the non - thunderstorm samples are significantly more than the thunderstorm samples. It is necessary to further calculate the F1 - score value to evaluate the pros and cons of the model. Figure 9 It is the change curve of the F1 value with the lead time when the threshold values of the model are 0.2 - 0.7 respectively. t The following conclusions can be drawn through comparison: 1) As the lead time t increases, the F1 - score value shows a downward trend, indicating that the larger the lead time, the worse the classification effect and the worse the early warning effect. 2) For the principal component analysis - binary Logistic regression model, when the threshold value changes from 0.2 to 0.7, the F1 value first increases and then decreases. When the threshold value is 0.3, the classification effect of the model is the best.

[0080] It can be seen that when the sample data classification is unbalanced, introducing the F1 - score value can more objectively and accurately reflect the pros and cons of the model performance. The lead time and the threshold value have a greater impact on the model performance parameters. Specifically, the larger the lead time t, the worse the performance. When the threshold value changes from small to large, the F1 value first increases and then decreases. When the threshold value is 0.3, the highest F1 value is 0.92 (t = 6), and the lowest is 0.68 (t = 54). And in most cases, the F1 - score value is significantly higher than the F1 - score values at other threshold values, indicating that when the threshold value is 0.3, the classification effect of the model is the best. Taking the lead time t = 18 minutes as an example, the effective alarm rate, miss rate, and false alarm rate of the early warning are 95.83%, 22.03%, and 4.17% respectively.

[0081] Step 7: Use the trained lightning early warning model for fixed - point lightning early warning. Embodiment

[0082] As Figure 10 shown, this embodiment provides a multi - source data - based fixed - point lightning early warning system for implementing the method described in Embodiment 1. The system includes: A data collection and preprocessing module for collecting and preprocessing multi-source lightning monitoring data to obtain a multi-source data set including atmospheric electric field data, lightning location data, and weather radar data; A sample sorting module for sorting the multi-source data set to obtain a sample set; A correlation analysis module for performing sample factor correlation analysis on the sample set to obtain a sample matrix; A normalization processing module for performing Z-Score normalization on the sample matrix to obtain a normalized matrix; A principal component extraction module for performing linear transformation on the normalized matrix using the principal component analysis method to extract five principal components: atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor; A model establishment and training module for establishing a lightning warning model using the binary logistic regression method based on the five principal components of the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor, and inputting a training data set to train the lightning warning model; A warning module for performing fixed-point lightning warning using the trained lightning warning model. Embodiment

[0083] As Figure 11 shown, this embodiment provides a computer device for implementing the fixed-point lightning warning method based on multi-source data described in Embodiment 1. The computer device includes: One or more processors, one or more power supplies, one or more operating systems, one or more computer programs, one or more databases, a memory, one or more network interfaces, and one or more input / output interfaces.

[0084] The processor can execute the method steps described in the aforementioned Embodiment 1, which will not be elaborated here.

[0085] The power supply can meet the power requirements for the normal operation or overclocking of the computer device.

[0086] The operating system is, for example TM, MacOSXTM, UnixTM, LinuxTM, etc. Pay attention to the version of the operating system corresponding to the running code when selecting the operating system.

[0087] The memory stores one or more application programs or data, which can be volatile storage or persistent storage. The programs stored in the memory include one or more modules, and each module can include a series of instruction operations for a computer device. The central processing unit can communicate with the memory and execute a series of instruction operations in the memory on the computer device. Embodiment

[0088] This embodiment provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in Embodiment 1 are implemented.

[0089] It should be noted that the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, and the readable storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories, random access memories, or optical discs, etc., and includes several instructions to enable a computer device, which can be a personal computer, a server, or a network device, etc., to execute all or part of the steps of the methods described in various embodiments of this application.

[0090] In summary, the present invention extracts and analyzes the features of multi-source lightning monitoring data such as atmospheric electric field data, new-generation weather radar data, and lightning location data, and performs data quality control such as interpolation and cleaning to form a thunderstorm sample set and a non-thunderstorm sample set. Then, statistical methods are used to analyze the relationship between multiple warning factors and lightning occurrence results, and correlation analysis, Z-Score standardization, principal component analysis (PCA), binary Logistic regression, etc. are performed on the sample data to establish a method and standard for extracting principal component factors, and the first to fifth principal components are extracted as modeling factors for the Logistic regression model to establish a fixed-point lightning warning model based on multi-source data, providing an operable and implementable lightning warning method for lightning disaster-sensitive places such as oil depots, schools, scenic spots, industrial parks, and mines, providing a scientific basis for managers to take lightning prevention measures in advance, and having important significance in reducing lightning disaster losses, ensuring enterprise production safety, and protecting people's lives and safety.

[0091] The technical solution provided by the present invention has been introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A fixed-point lightning warning method based on multi-source data, characterized in that, Including the steps: Step 1, data collection and preprocessing, to obtain a multi-source data set including atmospheric electric field data, lightning location data, and weather radar data; Step 2, perform sample sorting on the multi-source data set to obtain a sample set; Step 3, perform sample factor correlation analysis on the sample set to obtain a sample matrix; Step 4, perform Z-Score standardization on the sample matrix to obtain a standardized matrix; Step 5, use the principal component analysis method to perform a linear transformation on the standardized matrix, and extract five principal components: atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor; Step 6, based on the five principal components of the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor extracted, use the binary logistic regression method to establish a lightning warning model, and input the training data set to train the lightning warning model; Step 7, use the trained lightning warning model for fixed-point lightning warning; The expression of the trained lightning warning model is: ; Among them, ~ are respectively the atmospheric electric field discrete factor, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor, p is the probability of lightning occurrence.

2. The fixed-point lightning warning method based on multi-source data according to claim 1, characterized in that: The data preprocessing in Step 1 includes the following processes: Step 1.1, according to the lightning warning location and the defined coverage area range, perform data coverage area division on the collected multi-source data to achieve preliminary data screening; Step 1.2, for the multi-source data that has passed the preliminary data screening, use different methods for data processing according to the data type to obtain a multi-source data set, where: the atmospheric electric field data is processed using the mean imputation method; the weather radar data is processed using the bilinear interpolation algorithm.

3. The fixed-point lightning warning method based on multi-source data according to claim 2, characterized in that: The processing process of the lightning location data in Step 1.2 is as follows: Determine the distance between the target point and the lightning location point according to the longitude and latitude of the target point, screen out all lightning location data with a distance less than the preset distance threshold from the target point, and form a lightning location data set in the target area after sorting by time; Eliminate the lightning location data with a lightning current value less than the preset lightning current amplitude in the lightning location data set in the target area; According to the positioning principle of the lightning locator, eliminate the lightning location data with a positioning method lower than three stations in the lightning location data set.

4. The fixed-point lightning warning method based on multi-source data according to claim 1, characterized in that: The sample sorting to obtain the sample set in Step 2 specifically includes: Determine the sample elements; Determine the sample recording time and the time interval between each group of samples; Determine the variables included in a single sample and their meanings; Sort the data in the multi-source data set according to the sample elements, sample recording time, time interval between each group of samples, variables included in a single sample and their meanings to obtain a sample set.

5. The fixed-point lightning warning method based on multi-source data according to claim 4, characterized in that: The variables included in a single sample include sample time, mean value, absolute value mean, absolute value maximum, difference value mean, difference value maximum, variance, standard deviation, number of reversals, whether to reverse, echo mean, proportion within 15 km, proportion within 30 km.

6. The fixed-point lightning warning method based on multi-source data according to claim 1, characterized in that: The formula for performing sample factor correlation analysis on the sample set in Step 3 is: ; Among them, is the correlation coefficient between the samples , , n is the sample size, , , are respectively the standard scores, sample means and sample standard deviations of the sample , , , are respectively the standard scores, sample means and sample standard deviations of the sample .

7. The fixed-point lightning warning method based on multi-source data according to claim 1, characterized in that: The specific process of performing Z-Score standardization on the sample matrix in Step 4 to obtain a standardized matrix includes: Calculate the mean of each sample in the sample matrix and the standard deviation SD ; Normalize according to the formula ; where is the value of a certain sample in the sample matrix.

8. The fixed-point lightning warning method based on multi-source data according to claim 1, wherein: In step 5, the principal component analysis method is used to perform a linear transformation on the standardized matrix, and the discrete factor of atmospheric electric field, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor are extracted, specifically including: Step 5.1: Establish the covariance matrix R of the standardized matrix; Step 5.2: Solve the eigenvalues, principal component contribution rates and cumulative variance contribution rates of the covariance matrix R to determine the number of principal components; Step 5.3: According to the principle of selecting the number of principal components, count the eigenvalues whose eigenvalues are greater than 1 and the cumulative contribution rate reaches more than 80% λ 1 , λ 2 ,…, λ m and the corresponding m; Step 5.4: After performing a linear transformation with m unit eigenvectors as coefficients, an eigenvector matrix is obtained a T , and according to the formula the first m sample principal components are calculated to obtain the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor; Step 5.5: Establish the initial factor loading matrix to interpret the principal components.

9. The fixed-point lightning warning method based on multi-source data according to claim 8, wherein: The mathematical expressions of the discrete factor of atmospheric electric field, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor are: ; ; ; ; ; Among them, ~ are the atmospheric electric field discrete factor, radar echo intensity factor, atmospheric electric field reversal factor, atmospheric electric field intensity factor, and radar echo extreme value factor respectively, ~ are the variables included in a single sample respectively.

10. The fixed-point lightning warning method based on multi-source data according to claim 1, wherein: In step 6, based on the five extracted principal components, a lightning warning model is established using the binary logistic regression method, and the training data set is input to train the lightning warning model, specifically including: Step 6.1: Take "whether there is lightning" as the dependent variable, and take the five principal components of the discrete factor of atmospheric electric field, the radar echo intensity factor, the atmospheric electric field reversal factor, the atmospheric electric field intensity factor, and the radar echo extreme value factor as covariates, and introduce the sigmond function to establish the lightning warning model; Step 6.2: Input the training data set, estimate the parameters of the lightning warning model using the maximum likelihood estimation method, and train the lightning warning model using the gradient descent method and the quasi-Newton method.

Citation Information

Patent Citations

  • Thunder and lightning early warning data correction method based on real-time thunder and lighting positioning data correlation analysis

    CN105004932A

  • Multi-parameter and multi-algorithm integrated lightning refined monitoring and early warning algorithm

    CN114705922A

  • Lightning monitoring and early warning method and device

    CN114994801A

  • Lightning early warning system

    CN118962261A