A seismic numerical prediction method according to nuclear density and machine learning

By standardizing magnitude and performing kernel density analysis, combined with earthquake frequency and energy distribution across multiple time scales, and utilizing a random forest model for earthquake prediction, this approach addresses the shortcomings of insufficient feature extraction and bandwidth optimization in existing technologies, thereby achieving precise earthquake prediction and disaster prevention and mitigation support.

CN121741838BActive Publication Date: 2026-07-07INST OF EARTHQUAKE SCI CHINA EARTHQUAKE ADMINISTATION
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Existing earthquake prediction technologies suffer from limited feature extraction dimensions, lack of targeted optimization of bandwidth parameters, absence of long-term benchmark comparisons, and coarse prediction granularity, resulting in insufficient prediction accuracy and spatial applicability.

Method used

By standardizing earthquake magnitude and screening earthquake catalog data for completeness, combined with kernel density bandwidth sensitivity testing and multi-timescale analysis, anomaly distributions of earthquake frequency and energy kernel density are constructed, and refined predictions are made using a random forest model.

Benefits of technology

It enables multi-dimensional characterization of seismic activity, improves the accuracy of earthquake prediction and the rationality of spatial distribution modeling, meets the needs of precise disaster prevention and mitigation, and provides more targeted prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121741838B_ABST
    Figure CN121741838B_ABST
Patent Text Reader

Abstract

The application discloses a kind of according to nuclear density and machine learning Seismic Numerical Prediction Method, belong to the technical field of earthquake prediction.The method includes: the original earthquake catalogue data is standardized and energy correlation is handled, and standard earthquake catalogue data is formed, and the minimum complete magnitude is determined by completeness processing to standard earthquake catalogue data, and the complete earthquake catalogue is obtained by screening;The first and second bandwidth are determined by nuclear density bandwidth sensitivity test to the complete earthquake catalogue, the time and magnitude sliding window are combined, the seismic frequency and energy nuclear density distribution of target area are calculated, and the nuclear density anomaly distribution data is obtained by difference with long-term average nuclear density distribution;The nuclear density anomaly distribution data of target magnitude range under multiple time scales is input into pre-training random forest model, and the occurrence prediction result of target magnitude earthquake in target area is output.The method improves the spatial accuracy and timeliness of earthquake prediction, and provides technical support for earthquake disaster risk prevention and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of earthquake prediction technology, and more specifically to a numerical earthquake prediction method based on kernel density and machine learning. Background Technology

[0002] Numerical earthquake prediction methods are a core technology in the field of earthquake disaster risk prevention and control. By processing earthquake monitoring data, extracting features, and analyzing models, they can predict the probability of earthquakes occurring and their spatial distribution in target areas, providing a scientific basis for earthquake disaster early warning, disaster prevention and mitigation planning, and emergency response deployment. They typically have basic earthquake data processing and statistical analysis capabilities, aiming to improve the accuracy and timeliness of earthquake prediction and reduce casualties and property losses caused by earthquake disasters.

[0003] For example, patent publication number "CN113610147A" discloses a short-term earthquake prediction method based on LSTM multi-potential subspace information fusion. This method collects multi-source electromagnetic monitoring data, divides the sliding time window to make samples, and uses LSTM network and logistic regression model to fuse short-term earthquake prediction. It breaks through the limitation of insufficient prediction accuracy of traditional single-dimensional data and solves the fundamental problem of the difficulty in effectively fusing multi-source earthquake precursor data. However, existing earthquake prediction technologies still have significant shortcomings: First, the feature extraction dimensions are limited, focusing primarily on precursory physical quantities such as electromagnetic and geological data, without fully exploring the frequency and energy spatial distribution characteristics in earthquake catalog data, making it difficult to accurately characterize the spatiotemporal clustering patterns of regional seismic activity; second, bandwidth parameters lack targeted optimization, with fixed bandwidth often used in kernel density analysis, failing to differentiate bandwidth selection based on the characteristics of frequency and energy samples, easily leading to distorted density distribution modeling; third, long-term benchmark comparisons are not introduced, density modeling is only performed on single-period data, without combining long-term average distribution data to calculate outliers, making it impossible to effectively identify significant anomalies in seismic activity; fourth, the prediction granularity is relatively coarse, mostly regional-level overall predictions, failing to achieve grid-level fine spatial division, making it difficult to meet the spatial deployment requirements for precise disaster prevention and mitigation. These shortcomings restrict the prediction accuracy and spatial applicability of existing technologies and urgently require improvement. Summary of the Invention

[0004] The purpose of this invention is to provide a numerical earthquake prediction method based on kernel density and machine learning to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A numerical earthquake prediction method based on kernel density and machine learning includes the following steps:

[0007] S1. After standardizing the magnitudes of each earthquake event in the original earthquake catalog data of the target area, calculate the released energy of each earthquake event and correlate them to form standard earthquake catalog data of the target area.

[0008] S2. Completeness processing is performed on the standard earthquake catalog data of the target area to obtain the minimum complete magnitude, and the standard earthquake catalog data is filtered based on the minimum complete magnitude to obtain a complete earthquake catalog.

[0009] S2. Perform kernel density bandwidth sensitivity testing on the complete earthquake catalog to determine the first and second bandwidths;

[0010] S4. Based on the first bandwidth and the second bandwidth, the spatial distribution of earthquakes in the target area is modeled using time sliding windows and magnitude sliding windows to obtain earthquake frequency kernel density anomaly distribution data and earthquake energy kernel density anomaly distribution data in the target area.

[0011] S5. Input the earthquake frequency kernel density anomaly distribution data and earthquake energy kernel density anomaly distribution data of the target magnitude range at multiple time scales in the target area into the pre-trained random forest model for prediction, and obtain the prediction results of the occurrence of earthquakes of the target magnitude in the target area.

[0012] Preferably, the original earthquake catalog data includes the earthquake occurrence time, epicenter coordinates, earthquake depth, and earthquake magnitude for all earthquake events within the target area, and the earthquake magnitude type includes local magnitude. Surface wave magnitude Moment and Magnitude .

[0013] Preferably, the magnitude normalization process involves adjusting the local magnitude. Surface wave magnitude Convert to moment magnitude using the magnitude conversion formula The energy released By moment magnitude The result was obtained by substituting into the magnitude energy conversion formula;

[0014] The magnitude conversion formula is:

[0015] ;

[0016] ;

[0017] ;

[0018] ;

[0019] The formula for magnitude energy conversion is: log 10E=1.5M W +4.8.

[0020] Preferably, the method for the completeness processing is as follows:

[0021] S21. Set the initial minimum complete magnitude. Based on the initial minimum complete magnitude Select earthquake events with magnitudes greater than or equal to the initial minimum complete magnitude from the standard earthquake catalog data. This constitutes a subset of earthquake data, the initial minimum complete magnitude. The initial setting can be made by combining the distribution characteristics of earthquake magnitudes of historical earthquake events in the target area, and the value can be taken as the lowest earthquake magnitude of earthquake events in the target area plus an empirical step size parameter. , experimental step size parameter The preferred value range is [0.1, 0.5], that is... ,in, Earthquake events in the standard earthquake catalog data for the target area earthquake magnitude The earthquake's magnitude The type is moment magnitude;

[0022] S22. Calculate the value of b using the maximum likelihood formula for a subset of seismic data, and calculate the uncertainty of the value of b using the uncertainty estimation formula;

[0023] The maximum likelihood formula is:

[0024] ;

[0025] in, For the value of b, This represents the average earthquake magnitude of all earthquake events in the standard earthquake catalog data for the target area. The magnitude of the earthquake is the slip width, and Generally, 0.1 is taken, where e is the base of the natural logarithm;

[0026] The uncertainty estimation formula is as follows:

[0027] ;

[0028] in, denoted as b, where b is the uncertainty of the value, and N is the total number of earthquake events in the standard earthquake catalog data for the target area.

[0029] S23. The initial minimum complete magnitude. Subtract a preset magnitude step A new initial minimum complete magnitude is obtained. Among them, the preset magnitude step size The size should not exceed the empirical step size parameter. Repeating S21-S22 yields a new initial minimum complete magnitude. The corresponding new b value and new uncertainty ;

[0030] S24. If the new b value and b value The absolute value of the difference between them is less than the uncertainty, that is... Then the initial minimum complete magnitude If the minimum complete magnitude is reached, then repeat S23-S24.

[0031] Preferably, the complete earthquake catalog consists of all earthquake events in the standard earthquake catalog data whose magnitude is greater than the minimum complete magnitude, and each earthquake event in the complete earthquake catalog is associated with the corresponding earthquake occurrence time, earthquake epicenter coordinates, earthquake depth, earthquake magnitude and released energy.

[0032] Preferably, the method for nuclear density bandwidth sensitivity testing is as follows:

[0033] S31. Divide the target area into several grids according to a preset resolution. Each grid is a spatial calculation unit, and record the coordinates of its center point. j is the raster number, j=1,2,...,m, where m is the total number of raster cells. Simultaneously, frequency kernel density calculation sample sets and energy kernel density calculation sample sets are constructed based on the complete earthquake catalog. The frequency kernel density calculation sample set consists of the epicenter coordinates of each earthquake event in the complete earthquake catalog. The energy kernel density calculation sample set consists of the epicenter coordinates of each earthquake event in a complete earthquake catalog. It consists of energy release and preset candidate bandwidth ranges. In the candidate bandwidth range Inner step size Select multiple bandwidths , For bandwidth numbering, =1,2,...,p, where p is the total bandwidth;

[0034] S32. Based on each bandwidth, calculate the kernel density value grid-by-grid for both the frequency kernel density calculation sample set and the energy kernel density calculation sample set to obtain the frequency kernel density raster dataset for each bandwidth. and energy kernel density raster dataset The frequency kernel density calculation sample is calculated grid-by-grid using the earthquake frequency kernel density formula, and the energy kernel density calculation sample is calculated grid-by-grid using the earthquake energy kernel density formula. The frequency kernel density raster dataset consists of the frequency kernel density values ​​of all grids in the target area, and the energy kernel density raster dataset consists of the energy kernel density values ​​of all grids in the target area.

[0035] The formula for the earthquake frequency kernel density value is:

[0036] ;

[0037] in, Let K be the frequency kernel density value of grid j, K be the kernel function, and n be the total number of earthquake events in the frequency kernel density calculation sample set;

[0038] The formula for the earthquake energy kernel density value is:

[0039] ;

[0040] in, Let K be the energy kernel density value of grid j, K(.) be the kernel function, preferably a Gaussian kernel function, and n be the total number of seismic events in the energy kernel density calculation sample set. For earthquake events The release of energy The energy kernel density is calculated to represent the total energy released by all seismic events in the sample set.

[0041] S33. From the frequency kernel density raster dataset corresponding to each bandwidth and energy kernel density raster dataset Extract the highest frequency kernel density value corresponding to the bandwidth from each. and the highest value of nuclear density with energy ;

[0042] S34. Plot the first relationship curve L1, with bandwidth as the x-axis and the highest kernel density as the y-axis, to show the relationship between the highest kernel density and bandwidth: And the second relationship curve L1 showing the maximum energy nuclear density as a function of bandwidth: ;

[0043] S35. Determine the curvature abrupt change inflection points of the two curves using the curvature calculation method, and select the bandwidth corresponding to the curvature abrupt change inflection point of the first relationship curve L1 as the first bandwidth for calculating the adaptation frequency kernel density. The bandwidth corresponding to the curvature abrupt inflection point of the second relationship curve is selected as the second bandwidth for the adaptation energy kernel density calculation. .

[0044] Preferably, the time sliding window includes 1 month, 2 months, 3 months, 6 months and 1 year, and the magnitude sliding window includes ≥3, ≥4, ≥5 and ≥6.

[0045] Preferably, the method for modeling the spatial distribution of earthquakes is as follows:

[0046] S41. Set the time-based sliding window set ,in, For January, For 2 months, For 3 months, For 6 months, Set a sliding window set for magnitude over a year. ,in, Level 3 or higher Level 4 or higher Level 5 or higher For levels ≥6, a sliding window based on time is used. ( =1,2,...,5) and magnitude sliding window ( Window combination conditions =1,2,...,4) From a complete earthquake catalog, select earthquakes whose occurrence time belongs to The earthquake magnitude is consistent with All earthquake events form multiple modeling sub-sample sets. ,in, To model the epicenter coordinates of the i-th earthquake event in the subsample set, To release energy for it, The total number of earthquake events in the modeled subsample set under this window combination condition;

[0047] S42. Based on the first bandwidth For each modeling subset Calculate the frequency kernel density distribution grid by grid, grid center point frequency kernel density value The calculation formula is: Simultaneously based on the second bandwidth For each modeling subset Calculate the energy kernel density distribution grid by grid, grid center point Energy nuclear density value The calculation formula is: After traversing all grid cells, the frequency kernel density distribution under this window combination condition is obtained. and energy nuclear density distribution ;

[0048] S43. Based on the complete earthquake catalog of the target area, select a length of... For long-term statistical periods, conditions are combined within the same window. The long-term statistical period is divided into S equal-length sub-windows, and the corresponding frequency kernel density distribution is calculated for each sub-window. and energy nuclear density distribution ( =1,2,...,S), where S is the total number of sub-windows in the long-term statistical period; then, the long-term average frequency kernel density distribution corresponding to the frequency kernel density distribution under the same window combination condition is calculated respectively. The long-term average energy kernel density distribution corresponding to the energy kernel density distribution The calculation formulas are as follows:

[0049] ;

[0050] ;

[0051] S44. Frequency kernel density distribution under each window combination condition. Compare with the corresponding long-term average frequency kernel density distributions respectively Subtracting grid by grid yields the frequency kernel density anomaly distribution data. The energy kernel density distribution under various window combinations Corresponding to the long-term average energy kernel density distribution By subtracting grid by grid, the energy kernel density anomaly distribution data are obtained. .

[0052] Preferably, the prediction result of the target magnitude earthquake includes the binary classification result of each grid within the target area. The binary classification result is either an earthquake or no earthquake. The prediction result of the target magnitude earthquake needs to be obtained by inputting the earthquake frequency kernel density anomaly distribution data and earthquake energy kernel density anomaly distribution data of the target magnitude range at multiple time scales into a pre-trained random forest model. The earthquake frequency kernel density anomaly distribution data and earthquake energy kernel density anomaly distribution data of the target magnitude range at multiple time scales are: the earthquake occurrence time of the most recent earthquake event in the target area is used as the time reference node. Based on time reference nodes Using [the endpoint], we trace back 1 month, 2 months, 3 months, 6 months, and 1 year to form statistical periods with multiple time scales. The statistical intervals corresponding to these time scales are […]. -January, ]、[ -February, ]、[ -March, ]、[ June, ]and[ -1 year, Simultaneously, within each statistical interval, earthquake events whose magnitudes fall within the target magnitude range are selected, i.e., any magnitude interval of ≥3, ≥4, ≥5, or ≥6 corresponding to the magnitude sliding window. Then, frequency kernel density anomaly distribution data and energy kernel density anomaly distribution data under the corresponding time sliding window and magnitude sliding window, obtained based on earthquake spatial distribution modeling methods, are integrated to form a multi-timescale kernel density anomaly feature dataset for the target magnitude range, i.e., earthquake frequency kernel density anomaly distribution data and earthquake energy kernel density anomaly distribution data for the target magnitude range at multiple time scales. The data is distributed normally; the target earthquake magnitude is any one of the magnitude intervals of ≥3, ≥4, ≥5, and ≥6 corresponding to the magnitude sliding window. The target magnitude interval can be flexibly selected according to the actual magnitude concern needs of disaster prevention and mitigation; the multiple time scales are 1 month, 2 months, 3 months, 6 months, and 1 year corresponding to the time sliding window. Different time scales can correspond to short-term, medium-term, and medium-to-long-term earthquake occurrence prediction scenarios, respectively, to achieve a comprehensive prediction of earthquake occurrence risk in the target area within different time spans, providing data support for refined disaster prevention and mitigation deployment.

[0053] Preferably, the training process of the random forest model is as follows:

[0054] Historical raw earthquake catalog data of the target area is collected. Following the processing flow of S1-S4 described above, earthquake frequency kernel density anomaly distribution data and energy kernel density anomaly distribution data at multiple time scales (1 month, 2 months, 3 months, 6 months, 1 year) and multiple magnitude intervals (≥3, ≥4, ≥5, ≥6) are generated. These are used as the model input feature set X for the random forest model. At the same time, the actual earthquake event occurrence of each grid in the target area within one time unit after the corresponding statistical period of each time scale and magnitude interval is used as the label set Y. The value of Y is a binary discrete value. If an earthquake of the corresponding magnitude interval occurs in the grid, it is marked as 1 (earthquake) and if no earthquake occurs, it is marked as 0 (no earthquake). The model input feature set X and the label set Y are matched one-to-one at the grid time dimension to form the initial training sample set.

[0055] The initial training sample set was randomly divided into training, validation, and test sets in a 7:2:1 ratio. The initial ranges for the hyperparameters of the random forest model were set as follows: the initial range for the number of decision trees was [50, 200], the initial range for the maximum depth of decision trees was [5, 20], the initial range for the minimum number of split samples per node was [2, 10], and the feature sampling method was random sampling. Using the training set as input, hyperparameter optimization was performed on the validation set based on cross-validation. The F1 score was used as the core evaluation metric, and the combination of hyperparameters with the highest F1 score was selected as the optimal hyperparameters for the random forest model.

[0056] A random forest model is constructed based on optimal parameters. Multiple subsets of samples are extracted from the training set using bootstrap resampling technology, and corresponding decision tree learners are trained on each subset. Each decision tree learner independently performs binary classification prediction on its subset, thus training the random forest model. The test set is input into the trained random forest model, and model performance is validated using accuracy and F1 score. If both accuracy and F1 score reach preset thresholds (accuracy ≥ 85%, F1 score ≥ 0.8), the random forest model is considered successfully trained and used as a pre-trained random forest model. If accuracy and F1 score do not meet the preset thresholds, the hyperparameter range is readjusted and the model is trained again until the accuracy and F1 score requirements are met.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] 1. This invention expands the feature extraction dimensions for earthquake prediction, innovatively mining the spatial distribution characteristics of earthquake frequency and energy from earthquake catalog data, and constructing a multi-dimensional seismic activity characterization system by combining kernel density analysis. This breaks through the limitations of traditional technologies that rely solely on physical precursors and are difficult to accurately depict the spatiotemporal clustering patterns of regional earthquakes, making the basic data for prediction more consistent with the essential laws of seismic activity.

[0059] 2. This invention achieves differentiated and precise optimization of kernel density bandwidth. It designs a dedicated bandwidth sensitivity test process for the characteristics of two types of samples: frequency and energy. By determining the appropriate bandwidth through the curvature abrupt change inflection point, it solves the problem of density modeling distortion caused by fixed bandwidth and significantly improves the accuracy and rationality of seismic spatial distribution modeling.

[0060] 3. This invention introduces a long-term benchmark comparison mechanism, which obtains outliers by subtracting the kernel density distribution of a specific time period from the long-term average distribution. This can effectively identify significant anomalies in seismic activity, making up for the shortcomings of traditional techniques that only analyze data from a single time period and cannot distinguish between normal fluctuations and abnormal activities. This provides more targeted feature basis for prediction.

[0061] 4. This invention achieves grid-level refined prediction by dividing the target area into grids and outputting binary prediction results grid by grid, replacing the traditional regional-level overall prediction mode. This meets the spatial deployment requirements for precise disaster prevention and mitigation and enhances the supporting value of prediction results for actual emergency prevention and control work. Attached Figure Description

[0062] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0063] Figure 1 This is a flowchart of the method steps of the present invention. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be described in detail below. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other implementation methods obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0065] Examples, such as Figure 1 As shown, a numerical earthquake prediction method based on kernel density and machine learning includes the following steps:

[0066] S1. After standardizing the magnitudes of each earthquake event in the original earthquake catalog data of the target area, calculate the released energy of each earthquake event and correlate them to form standard earthquake catalog data of the target area.

[0067] S2. Completeness processing is performed on the standard earthquake catalog data of the target area to obtain the minimum complete magnitude, and the standard earthquake catalog data is filtered based on the minimum complete magnitude to obtain a complete earthquake catalog.

[0068] S2. Perform kernel density bandwidth sensitivity testing on the complete earthquake catalog to determine the first and second bandwidths;

[0069] S4. Based on the first bandwidth and the second bandwidth, the spatial distribution of earthquakes in the target area is modeled using time sliding windows and magnitude sliding windows to obtain earthquake frequency kernel density anomaly distribution data and earthquake energy kernel density anomaly distribution data in the target area.

[0070] S5. Input the earthquake frequency kernel density anomaly distribution data and earthquake energy kernel density anomaly distribution data of the target magnitude range at multiple time scales in the target area into the pre-trained random forest model for prediction, and obtain the prediction results of the occurrence of earthquakes of the target magnitude in the target area.

[0071] Furthermore, the working principle of the present invention will be illustrated below through embodiments:

[0072] This embodiment takes a seismically active area in southwestern my country as an example.

[0073] Raw earthquake catalog data for the target region from 2000 to 2024 was collected. This raw catalog data includes the earthquake occurrence time, epicenter coordinates, depth, and magnitude for each earthquake event. Earthquake magnitude types include local magnitude, surface wave magnitude, and moment magnitude. Non-moment magnitude data underwent magnitude standardization transformation. For example, if the local magnitude of an earthquake event on May 12, 2023, is 5.2, then the magnitude is standardized using the formula... The calculated moment magnitude is 5.06. Substituting this into the magnitude-energy conversion formula log... 10 E=1.5M W +4.8 calculates the released energy log 10 E is 12.39, that is, E = 2.45 × Erge, after standardizing the magnitude and correlating the energy of all earthquake events, formed a standard earthquake catalog data; subsequently, he performed completeness processing on the standard earthquake catalog data and set an initial minimum complete magnitude. The value is set to 3.0. Earthquake events with a magnitude ≥ 3.0 are selected to form a subset of earthquake data. The maximum likelihood formula yields a value of b of 1.02 and an uncertainty of 0.02. The value is 0.05. Subtract the magnitude step =0.1 to get the new initial value =2.9, and after repeated calculation, the new value of b is 1.03. At this time, |1.03-1.02|=0.01<0.05, and the minimum complete magnitude is determined to be 3.0. Earthquake events with magnitude >3.0 are selected from the standard earthquake catalog to form a complete earthquake catalog. The data of each earthquake event in the complete earthquake catalog are associated with the earthquake occurrence time, epicenter coordinates, earthquake depth, moment magnitude and released energy.

[0074] The target region was divided into several grids with a resolution of 0.2°×0.2°. Frequency kernel density and energy kernel density calculation sample sets were constructed. A candidate bandwidth interval of [0.3, 1.5] was preset, and 13 bandwidths were selected with a step size ∆h=0.1. Kernel density values ​​were calculated grid-by-grid for each bandwidth. The highest frequency kernel density and highest energy kernel density values ​​were extracted for each bandwidth. A first relationship curve and a second relationship curve were plotted. Curvature calculations showed a curvature abrupt change in the first relationship curve at bandwidth h=0.6, which was determined as the first bandwidth. The value is 0.6. The second relationship curve shows a curvature inflection point at bandwidth h=0.8, which is determined to be the second bandwidth. It is 0.8.

[0075] Define a time sliding window set T = {1 month, 2 months, 3 months, 6 months, 1 year} and a magnitude sliding window set M = {≥3, ≥4, ≥5, ≥6}, and combine the windows according to the conditions. Earthquake events from a comprehensive earthquake catalog are selected to form multiple modeling sub-sample sets. Taking window combination conditions (3 months, ≥4 magnitude) as an example, based on the first bandwidth... The frequency kernel density distribution of the modeling subset is calculated with a value of 0.6, based on the second bandwidth. The energy kernel density distribution of the modeling subsample set is calculated with a value of 0.8. The long-term statistical period is selected from 2000 to 2024. After dividing the sub-windows into equal-length sub-windows, the long-term average frequency kernel density distribution and the long-term average energy kernel density distribution under the same window combination conditions are calculated. The frequency kernel density distribution and energy kernel density distribution of the current period are subtracted from the long-term average frequency kernel density distribution and the long-term average energy kernel density distribution grid by grid to obtain the frequency kernel density anomaly distribution data and energy kernel density anomaly distribution data under the window combination conditions. The calculation is completed in the same way for other window combination conditions (3 months, ≥4 levels).

[0076] The frequency and energy kernel density anomaly distribution data of earthquakes with magnitude ≥4 in the target area during January, February, and March 2024 (a 3-month timescale) were input into a pre-trained random forest model. This pre-trained random forest model was trained using frequency and energy kernel density anomaly distribution data of the same timescale and magnitude range from 1980 to 2023 as the feature set and the actual earthquake occurrence data of the corresponding grid cells one month later as the label set. The accuracy on the test set reached 88%, and the F1 score reached 0.83. The pre-trained random forest model output the binary classification results of each grid cell in the target area. Among them, all grid cells near 102°E and 25°N were predicted as "earthquake". The area corresponding to the subsequent grid cell prediction result of "earthquake" experienced a magnitude 4.2 earthquake on April 5, 2024, which is consistent with the prediction result. Most of the remaining grid cells were predicted as "no earthquake", which is consistent with the actual monitoring situation, thus verifying the effectiveness of the method.

[0077] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; modifications to the technical solutions described in the foregoing embodiments, or equivalent substitutions of some of the technical features, do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A numerical earthquake prediction method based on kernel density and machine learning, characterized in that, Includes the following steps: S1. After standardizing the magnitudes of each earthquake event in the original earthquake catalog data of the target area, calculate the released energy of each earthquake event and correlate them to form standard earthquake catalog data of the target area. S2. Completeness processing is performed on the standard earthquake catalog data of the target area to obtain the minimum complete magnitude, and the standard earthquake catalog data is filtered based on the minimum complete magnitude to obtain a complete earthquake catalog. S3. Perform kernel density bandwidth sensitivity testing on the complete earthquake catalog to determine the first and second bandwidths; The method for the kernel density bandwidth sensitivity test is as follows: S31. Divide the target area into several grids according to a preset resolution. At the same time, construct frequency kernel density calculation sample sets and energy kernel density calculation sample sets according to a complete earthquake catalog. Preset candidate bandwidth ranges and select multiple bandwidths within the candidate bandwidth ranges. S32. Based on each bandwidth, calculate the kernel density value grid by grid for the frequency kernel density calculation sample set and the energy kernel density calculation sample set respectively, to obtain the frequency kernel density raster dataset and the energy kernel density raster dataset under each bandwidth; S33. Extract the highest value of frequency kernel density and the highest value of energy kernel density corresponding to each bandwidth from the frequency kernel density raster dataset and the energy kernel density raster dataset, respectively; S34. Plot the first curve showing the relationship between the highest frequency kernel density and bandwidth, and the second curve showing the relationship between the highest energy kernel density and bandwidth, respectively. S35. Select the bandwidth corresponding to the curvature change inflection point of the first relationship curve as the first bandwidth, and select the bandwidth corresponding to the curvature change inflection point of the second relationship curve as the second bandwidth. S4. Based on the first bandwidth and the second bandwidth, the spatial distribution of earthquakes in the target area is modeled using time sliding windows and magnitude sliding windows to obtain earthquake frequency kernel density anomaly distribution data and earthquake energy kernel density anomaly distribution data in the target area. S5. Input the earthquake frequency kernel density anomaly distribution data and earthquake energy kernel density anomaly distribution data of the target magnitude range at multiple time scales in the target area into the pre-trained random forest model for prediction, and obtain the prediction results of the occurrence of earthquakes of the target magnitude in the target area.

2. The earthquake numerical prediction method based on kernel density and machine learning according to claim 1, characterized in that, The original earthquake catalog data includes the earthquake occurrence time, epicenter coordinates, earthquake depth, and earthquake magnitude for all earthquake events within the target area. The earthquake magnitude types include local magnitude, surface wave magnitude, and moment magnitude.

3. The earthquake numerical prediction method based on kernel density and machine learning according to claim 2, characterized in that, The magnitude standardization process involves converting local magnitudes and surface wave magnitudes into moment magnitudes using a magnitude conversion formula. The released energy is calculated by substituting the moment magnitude into the magnitude energy conversion formula.

4. The earthquake numerical prediction method based on kernel density and machine learning according to claim 3, characterized in that, The method for completeness processing is as follows: S21. Set an initial minimum complete magnitude, and based on the initial minimum complete magnitude, select earthquake events with magnitudes greater than or equal to the initial minimum complete magnitude from the standard earthquake catalog data to form a subset of earthquake data; S22. Calculate the value of b using the maximum likelihood formula for a subset of seismic data, and calculate the uncertainty of the value of b using the uncertainty estimation formula; S23. Subtract a preset magnitude step size from the initial minimum complete magnitude to obtain a new initial minimum complete magnitude. Repeat S21-S22 to obtain the new b value and new uncertainty corresponding to the new initial minimum complete magnitude. S24. If the absolute value of the difference between the new b value and the b value is less than the uncertainty, then the initial minimum complete magnitude is the minimum complete magnitude; otherwise, repeat S23-S24.

5. The earthquake numerical prediction method based on kernel density and machine learning according to claim 4, characterized in that, The complete earthquake catalog consists of all earthquake events in the standard earthquake catalog data whose magnitude is greater than the minimum complete magnitude. Each earthquake event in the complete earthquake catalog is associated with the corresponding earthquake occurrence time, epicenter coordinates, earthquake depth, earthquake magnitude, and released energy.

6. The earthquake numerical prediction method based on kernel density and machine learning according to claim 1, characterized in that, The time sliding window includes 1 month, 2 months, 3 months, 6 months and 1 year, and the magnitude sliding window includes ≥3, ≥4, ≥5 and ≥6.

7. The earthquake numerical prediction method based on kernel density and machine learning according to claim 6, characterized in that, The method for modeling the spatial distribution of earthquakes is as follows: S41. Based on the window combination conditions of time sliding window and magnitude sliding window, select all earthquake events that meet the window combination conditions from the complete earthquake catalog to form multiple sets of modeling subsamples; S42. Calculate the frequency kernel density distribution grid-by-grid for each modeling sub-sample set based on the first bandwidth, and calculate the energy kernel density distribution grid-by-grid for each modeling sub-sample set based on the second bandwidth; S43. Based on the complete seismic catalog of the target area, calculate the long-term average frequency kernel density distribution and the long-term average energy kernel density distribution corresponding to the frequency kernel density distribution under the same window combination conditions. S44. Subtract the frequency kernel density distribution under each window combination condition from the corresponding long-term average frequency kernel density distribution grid by grid to obtain the frequency kernel density anomaly distribution data. Subtract the energy kernel density distribution under each window combination condition from the corresponding long-term average energy kernel density distribution grid by grid to obtain the energy kernel density anomaly distribution data.

8. The earthquake numerical prediction method based on kernel density and machine learning according to claim 7, characterized in that, The predicted occurrence result of the target magnitude earthquake includes the binary classification result of each grid in the target area, wherein the binary classification result is either an earthquake or no earthquake; the target magnitude earthquake is any one of the magnitude intervals corresponding to the magnitude sliding window; and the multiple time scales are 1 month, 2 months, 3 months, 6 months and 1 year corresponding to the time sliding window.

Citation Information

Patent Citations

  • Multi-potential subspace information fusion earthquake short-term and temporary prediction method based on LSTM (Long Short Term Memory)

    CN113610147A

  • Methods and systems for earthquake detection and prediction

    US20210318455A1

  • Model generation device, model generation method and program

    WO2023021690A1