Karst development area geological disaster identification method and system based on artificial intelligence

Through artificial intelligence technology, combined with multi-source data fusion and GIS, the geological disaster risks in karst development areas can be predicted in real time, solving the problems of complex and time-consuming information acquisition in traditional methods, and realizing efficient and accurate karst collapse risk assessment and early warning.

CN120689981APending Publication Date: 2025-09-23贵州省地质矿产勘查开发局114地质大队
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510872131.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Traditional karst collapse research methods are complex in information acquisition, rely on expert experience, are time-consuming and labor-intensive, and lack technical means, making it difficult to quickly and accurately identify and assess karst collapse risks.

Method used

Using an artificial intelligence-based method, through multi-source heterogeneous data fusion, Bayesian fusion algorithm, random forest model and GIS technology, the geological disaster risks in karst development areas are predicted in real time, risk level distribution maps are generated and early warning information is issued.

Benefits of technology

It improves the efficiency of karst collapse identification and prevention, reduces dependence on manpower and time, improves the accuracy and reliability of research, and realizes dynamic updating and early warning of high-risk areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689981A_ABST
    Figure CN120689981A_ABST
Patent Text Reader

Abstract

The invention discloses a karst development area geological disaster identification method and system based on artificial intelligence, relates to the field of geological disaster early warning and risk assessment, and solves the problems that it is difficult to integrate geology and related data, it is difficult to construct a model to predict disaster categories based on historical disaster cases and real-time monitoring data, and the risk assessment efficiency is improved. The technical problem of lack of association of a prediction result with a geographic space to generate a risk distribution map and dynamically update a high-risk area is solved; comprising the following steps: introducing multi-source heterogeneous data, and combining historical geological disaster cases; preprocessing the data; extracting underground water level change rate, soil humidity and vibration frequency characteristics; combining a plurality of features into a feature matrix, and projecting the feature matrix to a principal component space; after training the model by using a random forest algorithm, optimizing the constructed model by combining cross validation with a Bayesian algorithm; the model prediction result is converted into a risk level value through a mapping rule; and generating a risk level distribution map based on the GIS data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of geological disaster early warning and risk assessment, and specifically relates to a geological disaster identification method and system in karst development areas based on artificial intelligence. Background Art

[0002] Karst collapse is a common geological hazard in limestone areas, posing a serious threat to public life and property. Currently, the underlying mechanisms and potential risks of karst collapse on urban roads remain unclear. In densely populated urban environments, especially during the flood season, adverse factors such as heavy rainfall can exacerbate the probability and risk of karst collapse. This not only poses a significant threat to the safety of passing vehicles and pedestrians on these roads, but can also lead to traffic disruptions and economic losses. Therefore, there is an urgent need to strengthen research on karst collapse to clarify its causes, assess potential risks, and implement effective preventive measures to ensure the safety and smooth flow of urban roads.

[0003] Traditional karst collapse research methods usually include data collection, field surveys, geophysical exploration, engineering geological drilling, and experimental testing. Through these comprehensive analyses, the development characteristics of karst and the cause of collapse can be identified, and corresponding prevention and control measures can be proposed. However, these methods also have some obvious problems: 1. Information acquisition is huge and complex: a large amount of geological information and data needs to be collected, and the information is diverse, making processing and analysis difficult; 2. Reliance on expert experience: Many judgments and analyses rely on the extensive experience of geologists, which may lead to subjectivity and uncertainty, thus affecting the accuracy of the conclusions; 3. Time-consuming and labor-intensive: Traditional methods usually require a long time for field investigation and data collection, consuming a lot of human resources and being relatively inefficient; 4. Technical limitations: In some cases, the accuracy and applicability of traditional exploration technologies may be insufficient, making it difficult to fully reflect the true situation of karst. It is also difficult to build models to predict disaster categories based on historical disaster cases and real-time monitoring data. There is a lack of linking prediction results with geographic space to generate risk distribution maps and dynamically update high-risk areas.

[0004] Therefore, there is an urgent need to explore more efficient and intelligent research methods, such as those combining modern technologies such as artificial intelligence, big data analysis, and remote sensing technology, to improve the identification, analysis, and prevention of karst collapses. These new methods can process and analyze data more quickly, reduce reliance on manpower and time, and improve the accuracy and reliability of overall research. Summary of the Invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes a method and system for identifying geological hazards in karst development areas based on artificial intelligence.

[0006] To solve the above problems, the first aspect of the present invention provides a method for identifying geological hazards in karst development areas based on artificial intelligence, comprising the following steps: S1: Introducing multi-source heterogeneous data, combined with historical geological disaster cases, and using the Bayesian fusion algorithm to calculate the posterior probability using prior probability and conditional probability; S2: preprocess the data and normalize the data using data normalization methods; S3: Extract groundwater level change rate, soil moisture and vibration frequency characteristics; S4: Combine the extracted groundwater level change rate, soil moisture, and vibration frequency features into an original feature matrix. Use the principal component analysis method to construct a projection matrix, project the original feature matrix into the principal component space, and obtain the feature matrix after dimensionality reduction. S5: Use the random forest algorithm to train the model, and the final random forest model RF votes on the prediction results of all decision trees; S6: Use cross-validation combined with Bayesian optimization to build a random forest model; S7: After normalization and PCA dimension reduction, the real-time features are input into the random forest model, and the prediction results are converted into risk level values ​​through mapping rules; S8: Generate a risk level distribution map based on GIS data, set thresholds to extract the boundaries of high-risk areas, dynamically update risk levels and issue warning information.

[0007] Preferably, the step S1 includes the following steps: In view of the geological characteristics of karst development areas, multi-source heterogeneous data fusion technology is introduced to collect geological data and combine it with meteorological data, topographic data and vegetation cover data. At the same time, historical geological disaster cases are collected; Use data fusion algorithms to fuse data from different sources; The data fusion algorithm adopts the Bayesian fusion algorithm to calculate the posterior probability based on the prior probability and conditional probability, specifically: in, is the posterior probability, is the conditional probability, which means the probability of event B occurring under the condition that event A occurs. is the probability of event A, which means the estimation of the probability of event A before there is any observation data. is the probability of event B occurring.

[0008] Preferably, the step S2 comprises the following steps: The data is preprocessed, including removing noise and irrelevant data, formatting the data from different sources uniformly, and using a data normalization method to normalize the data to the [a, b] interval. The data normalization method is specifically as follows: in is the normalized sample data, is the sample data, is a data set, and For the normalized lower and upper limits of the interval, that is, the minimum and maximum values ​​of the target interval, respectively, is the minimum value in the sample set, is the maximum value in the sample set.

[0009] Preferably, the step S3 comprises the following steps: The groundwater level change rate is characterized by time series analysis; The soil moisture is collected by a soil moisture sensor to extract features; The vibration frequency is converted into a frequency domain signal by Fourier transform to extract the vibration frequency feature.

[0010] Preferably, the step S4 comprises the following steps: The extracted groundwater level change rate, soil moisture value and vibration frequency features are combined into the original feature matrix; Normalize the original feature matrix, calculate the covariance matrix of the normalized matrix, perform eigenvalue decomposition on the covariance, obtain the eigenvectors and eigenvalues, and sort them from large to small according to the eigenvalues; Select the first p principal components and construct the projection matrix ,in, For the feature vectors; Project the original feature matrix into the principal component space to obtain the reduced dimension feature matrix, specifically: in, is the feature matrix after dimensionality reduction, is the original feature matrix, is the projection matrix.

[0011] Preferably, the step S5 comprises the following steps: The reduced feature matrix and the corresponding historical disaster labels are combined into a training dataset D. The historical disaster labels are geological disaster categories, including no disaster, low-risk disaster, medium-risk disaster, and high-risk disaster; The model is trained using the random forest algorithm. The final random forest model RF votes on the prediction results of all decision trees, specifically: in, Input samples for random forest The final prediction result is For the A decision tree for samples The predicted value of is the number of decision trees in the random forest, is the mode, that is, the value that occurs most frequently.

[0012] Preferably, the step S6 comprises the following steps: Select K-fold cross validation, define the hyperparameter search space, and use Bayesian optimization to find the optimal hyperparameters in cross validation; Train the random forest model on each fold of the training set, evaluate the performance using the validation set, and record the performance indicators for each validation; Analyze the results of cross-validation, adjust hyperparameters, feature engineering, or model structure, and re-perform cross-validation; Retrain the model on the entire training set using the optimal hyperparameters.

[0013] Preferably, the step S7 includes the following steps: Align the real-time feature vector with the statistics of the training dataset and perform normalization; Use principal component analysis to reduce the dimension of the normalized real-time feature vector to the same dimension as the training model; The real-time feature vector after dimensionality reduction is input into the trained random forest model, and the predicted geological hazard category is output. The geological hazard category is converted into a risk level value according to the mapping rules; The mapping rules are formulated based on historical data and expert experience, specifically: If the model prediction result is no disaster, it means safety and the risk level value is 0; if it is a low-risk disaster, it means low risk and the risk level value is 1; if it is a medium-risk disaster, it means medium risk and the risk level value is 2; if it is a high-risk disaster, it means high risk and the risk level value is 3.

[0014] Preferably, the step S8 comprises the following steps: Collect GIS data related to geological hazards, associate the predicted risk level value of each monitoring point with its geographical location, and generate a risk level distribution map; Based on the risk level distribution map, identify high-risk areas. Specifically, set risk level thresholds and use spatial analysis tools to extract high-risk area boundaries. Update the risk level of monitoring points in real time, dynamically adjust the boundaries of high-risk areas, issue early warning information through GIS, and mark high-risk areas.

[0015] The second aspect of the present invention provides a geological disaster identification system for karst development areas based on artificial intelligence, comprising the following modules: a data acquisition module: introducing multi-source heterogeneous data, combining historical geological disaster cases, and calculating posterior probabilities using prior probabilities and conditional probabilities through a Bayesian fusion algorithm; Data preprocessing module: preprocess the data and normalize the data using data normalization method; Geological disaster feature extraction module: extracts groundwater level change rate, soil moisture and vibration frequency features; Principal component analysis dimensionality reduction module: The extracted groundwater level change rate, soil moisture, and vibration frequency features are combined into the original feature matrix. The principal component analysis method is used to construct a projection matrix. The original feature matrix is ​​projected into the principal component space to obtain the reduced dimension feature matrix. Random forest model building module: Use the random forest algorithm to train the model, and the final random forest model RF votes on the prediction results of all decision trees; Model optimization module: Use cross-validation combined with Bayesian algorithm to optimize the constructed random forest model; Risk level mapping module: Real-time features are normalized and PCA dimension reduced before being input into the random forest model. The prediction results are converted into risk level values ​​through mapping rules. GIS risk visualization and dynamic warning module: Generate risk level distribution maps based on GIS data, set thresholds to extract high-risk area boundaries, dynamically update risk levels and issue warning information.

[0016] Beneficial effects of the present invention: This paper introduces multi-source heterogeneous data fusion technology based on the geological characteristics of karst development areas, combines meteorological, topographic, vegetation and other data, and dynamically calculates the posterior probability through the Bayesian fusion algorithm, thereby improving the accuracy of data fusion. The present invention extracts the groundwater level change rate, soil moisture, and vibration frequency characteristics based on the geological disaster characteristics of karst development areas. It then optimizes the feature matrix through PAC dimensionality reduction. Then, through Bayesian optimization and K-fold cross-validation, it optimizes the random forest model based on the characteristics of karst disaster data, thereby improving model performance. The present invention maps the disaster categories predicted by the model into risk level values, updates the boundaries of high-risk areas in real time through GIS, and introduces an anomaly detection mechanism based on the random forest model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of the method flow of the present invention; Figure 2 It is a schematic diagram of the module flow of the present invention. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0019] See also Figure 1 As shown, the present invention is a method for identifying geological hazards in karst development areas based on artificial intelligence, comprising the following steps: S1: Introducing multi-source heterogeneous data, combined with historical geological disaster cases, and using the Bayesian fusion algorithm to calculate the posterior probability using prior probability and conditional probability; S2: preprocess the data and normalize the data using data normalization methods; S3: Extract groundwater level change rate, soil moisture and vibration frequency characteristics; S4: Combine the extracted groundwater level change rate, soil moisture, and vibration frequency features into an original feature matrix. Use the principal component analysis method to construct a projection matrix, project the original feature matrix into the principal component space, and obtain the feature matrix after dimensionality reduction. S5: Use the random forest algorithm to train the model, and the final random forest model RF votes on the prediction results of all decision trees; S6: Use cross-validation combined with Bayesian optimization to build a random forest model; S7: After normalization and PCA dimension reduction, the real-time features are input into the random forest model, and the prediction results are converted into risk level values ​​through mapping rules; S8: Generate a risk level distribution map based on GIS data, set thresholds to extract the boundaries of high-risk areas, dynamically update risk levels and issue warning information.

[0020] In one embodiment of the present invention, step S1 includes the following steps: In view of the geological characteristics of karst development areas, multi-source heterogeneous data fusion technology is introduced to collect geological data and combine it with meteorological data, topographic data and vegetation cover data. At the same time, historical geological disaster cases are collected; Use data fusion algorithms to fuse data from different sources; The data fusion algorithm adopts the Bayesian fusion algorithm to calculate the posterior probability based on the prior probability and conditional probability, specifically: in, is the posterior probability, is the conditional probability, which means the probability of event B occurring under the condition that event A occurs. is the probability of event A, which means the estimation of the probability of event A before there is any observation data. is the probability of event B occurring.

[0021] Specifically, data on geological structure, rock layer distribution, groundwater level, soil type, etc. in karst development areas are collected, as well as corresponding meteorological data, topographic data, and vegetation coverage data. The time, location, cause, and impact information of historical disaster events in karst areas (such as volcanic eruptions and karst collapses) are collected. Data are preprocessed, including denoising and cleaning, standardization, and time-space alignment. Event A is defined as a geological disaster occurring in a karst area, and event B is the observed data. The prior probability To estimate the initial probability of event A based on historical disaster cases, the marginal probability is the overall probability of observing event B, regardless of event A, and the conditional probability The probability of observing event B under the condition that event A occurs, determined through statistical analysis or expert experience.

[0022] In one embodiment of the present invention, step S2 includes the following steps: The data is preprocessed, including removing noise and irrelevant data, formatting the data from different sources uniformly, and using a data normalization method to normalize the data to the [a, b] interval. The data normalization method is specifically as follows: in is the normalized sample data, is the sample data, is a data set, and For the normalized lower and upper limits of the interval, that is, the minimum and maximum values ​​of the target interval, respectively, is the minimum value in the sample set, is the maximum value in the sample set.

[0023] In one embodiment of the present invention, step S3 includes the following steps: The groundwater level change rate is characterized by time series analysis; The soil moisture is collected by a soil moisture sensor to extract features; The vibration frequency is converted into a frequency domain signal by Fourier transform to extract the vibration frequency feature.

[0024] Specifically, the groundwater level change rate is characterized by time series analysis. For time The groundwater level at , the groundwater level change rate is calculated using the difference method, specifically: in, For time The rate of change of groundwater level at For time The groundwater level at For time The groundwater level at is the time interval; The soil moisture is collected by sensors, assuming Indicates time The soil moisture at the location can be smoothed using the moving average method, specifically: in, For time The smoothed soil moisture value is is the window size of the moving average, For time The original soil moisture value at ; The vibration frequency is converted into a frequency domain signal by Fourier transform. For time The vibration signal at the frequency domain It is obtained by discrete Fourier transform, specifically: in, In the frequency domain The complex representation of the frequency components, In the time domain The vibration signal value of each sampling point, is the total number of sampling points of the signal, is the frequency index, is the imaginary unit, is a complex exponential function used to map time domain signals to frequency domain.

[0025] In one embodiment of the present invention, step S4 includes the following steps: The extracted groundwater level change rate, soil moisture value and vibration frequency features are combined into the original feature matrix; Normalize the original feature matrix, calculate the covariance matrix of the normalized matrix, perform eigenvalue decomposition on the covariance, obtain the eigenvectors and eigenvalues, and sort them from large to small according to the eigenvalues; Select the first p principal components and construct the projection matrix ,in, For the feature vectors; Project the original feature matrix into the principal component space to obtain the reduced dimension feature matrix, which is: in, is the feature matrix after dimensionality reduction, is the original feature matrix, is the projection matrix.

[0026] Specifically, the input data includes the extracted groundwater level change rate, soil moisture value and vibration frequency characteristics. The characteristics of each monitoring point are combined into a row vector, and the data of all monitoring points constitute the original feature matrix. For example, if each monitoring point has 3 features (groundwater level, soil moisture and vibration frequency), and there are m monitoring points, then the original feature matrix is ​​an m*3 matrix; the features are standardized to eliminate the dimensional differences of different features so that each feature has the same scale; the covariance matrix is ​​calculated to measure the linear correlation between the features, and the eigenvalue decomposition is performed to extract the eigenvectors and eigenvalues ​​of the covariance matrix to determine the direction of the principal component; the original features are projected into the principal component space to construct a projection matrix, and the original feature matrix is ​​reduced to p-dimensional space to obtain the reduced feature matrix with a dimension of m*p.

[0027] In one embodiment of the present invention, step S5 includes the following steps: The reduced feature matrix and the corresponding historical disaster labels are combined into a training dataset D. The historical disaster labels are geological disaster categories, including no disaster, low-risk disaster, medium-risk disaster, and high-risk disaster; The model is trained using the random forest algorithm. The final random forest model RF votes on the prediction results of all decision trees, specifically: in, Input samples for random forest The final prediction result is For the A decision tree for samples The predicted value of is the number of decision trees in the random forest, is the mode, that is, the value that occurs most frequently.

[0028] Specifically, the feature matrix after dimensionality reduction is combined with the corresponding historical disaster labels to form a training data set D, where the historical disaster labels are geological disaster categories, including no disaster, low-risk disaster, medium-risk disaster, and high-risk disaster; The model is trained using the random forest algorithm, specifically: S401: Randomly extract n samples with replacement from the training data set D as the training set of the decision tree Ti; S402: When each node is bifurcated, m features are randomly selected from the feature set as candidates, and the optimal splitting attribute is selected based on the Gini impurity; S403: recursively construct a decision tree based on the selected features and split points until a stopping condition is reached; S404: Repeat steps S401-S403 to construct multiple decision trees T1, T2, ...Tk; The final random forest model RF votes on the prediction results of all decision trees, specifically: in, Input samples for random forest The final prediction result is For the A decision tree for samples The predicted value of is the number of decision trees in the random forest, is the mode, that is, the value that occurs most frequently.

[0029] In one embodiment of the present invention, step S6 includes the following steps: Select K-fold cross validation, define the hyperparameter search space, and use Bayesian optimization to find the optimal hyperparameters in cross validation; Train the random forest model on each fold of the training set, evaluate the performance using the validation set, and record the performance indicators for each validation; Analyze the results of cross-validation, adjust hyperparameters, feature engineering, or model structure, and re-perform cross-validation; Retrain the model on the entire training set using the optimal hyperparameters.

[0030] Specifically, the input data is the training data set, including the feature matrix and labels after dimensionality reduction. The hyperparameter search space is defined, including random forest hyperparameters and Bayesian optimization objectives. Among them, the random forest hyperparameters include the number of decision trees, the maximum depth of the decision tree, the minimum number of samples for node splitting, and the number of features randomly selected for each tree. The Bayesian optimization goal is to maximize the performance index of cross-validation, accuracy or F1 score; initialize K-fold cross-validation with K folds, divide the training data set D into K mutually exclusive subsets, and in each fold, K-1 subsets are used as training sets, and the remaining subset is used as validation set. The model is trained on the training set and the performance is evaluated on the validation set; use the BayesianOptimization library to define the objective function: Under each set of hyperparameters, K-fold cross validation is performed and the average performance index is returned. The Bayesian optimizer selects the next set of hyperparameters for evaluation based on historical results, repeats the iteration until the maximum number of iterations is reached or convergence is achieved, and returns the hyperparameter combination that optimizes the cross-validation performance; records the accuracy, confusion matrix, F1 score, etc. of each fold. If the performance of a fold is significantly lower than that of other folds, check whether the data distribution is unbalanced or whether the feature engineering is reasonable. If the overall performance is not ideal, adjust the hyperparameter search space or feature engineering; based on the Bayesian optimization results, update the hyperparameter search space, rerun K-fold cross validation and Bayesian optimization, select the best hyperparameter combination based on the Bayesian optimization results, use the entire training set D and the optimal hyperparameters, and retrain the random forest model.

[0031] In one embodiment of the present invention, step S7 includes the following steps: Align the real-time feature vector with the statistics of the training dataset and perform normalization; Use principal component analysis to reduce the dimension of the normalized real-time feature vector to the same dimension as the training model; The real-time feature vector after dimensionality reduction is input into the trained random forest model, and the predicted geological hazard category is output. The geological hazard category is converted into a risk level value according to the mapping rules; The mapping rules are formulated based on historical data and expert experience, specifically: If the model prediction result is no disaster, it means safety and the risk level value is 0; if it is a low-risk disaster, it means low risk and the risk level value is 1; if it is a medium-risk disaster, it means medium risk and the risk level value is 2; if it is a high-risk disaster, it means high risk and the risk level value is 3.

[0032] Specifically, data is collected in real time, and the groundwater level change rate, soil moisture value, and vibration frequency are calculated and combined into a real-time feature vector; statistical data is extracted from the training data set, including the mean and standard deviation of each feature, and the real-time feature vector is aligned with the statistical data of the training data set and standardized; the covariance matrix of the training data set is used to calculate the principal components, and the first kk principal components are selected, which are consistent with the dimensions of the training model after dimensionality reduction. For example, if the training model is reduced to 2 dimensions, the first 2 principal components are selected; the standardized real-time feature vector is projected into the principal component space; the saved random forest model RF is loaded and the real-time feature vector after dimensionality reduction is input into the model; mapping rules are defined based on historical data and expert experience, and the prediction results are converted into risk level values ​​according to the mapping rules; Among them, no disaster means there are no signs of geological disasters; low-risk disasters mean small-scale disasters with little impact; medium-risk disasters mean medium-scale disasters that may cause losses; high-risk disasters mean large-scale disasters with serious threats.

[0033] In one embodiment of the present invention, step S8 includes the following steps: Collect GIS data related to geological hazards, associate the predicted risk level value of each monitoring point with its geographical location, and generate a risk level distribution map; Based on the risk level distribution map, identify high-risk areas. Specifically, set risk level thresholds and use spatial analysis tools to extract high-risk area boundaries. Update the risk level of monitoring points in real time, dynamically adjust the boundaries of high-risk areas, issue early warning information through GIS, and mark high-risk areas. Specifically, the data of each monitoring point includes its geographic location and predicted risk level value. GIS software (such as ArcGIS, QGIS) or Python libraries (such as gdal, rasterio) are used to interpolate the risk level value of the monitoring point to the entire area, and a color gradient is used to represent the risk level, with 0 = green, 1 = yellow, 2 = orange, and 3 = red. A high-risk threshold is set according to actual needs. For example, if the risk level is ≥ 2, the "Extract Raster" or "Classify" tool is used in the GIS software to extract areas with a risk level ≥ 2 as polygons, and the boundaries of the high-risk areas are exported as vector files. A real-time sensor network is built and a calculation model is established to obtain real-time monitoring data, which is then standardized, reduced in dimension, and input into the model to predict the risk level and update the risk level value of the monitoring point. After each update of the monitoring point data, a risk level distribution map is regenerated, the boundaries of the high-risk areas are re-extracted according to the threshold, the old and new boundaries are compared, and the changed areas are identified. Map services are published through GIS platforms (such as ArcGIS Online and QGIS Server), and the annotation content includes the boundaries of the high-risk areas, the risk level values ​​of the monitoring points, and early warning information.

[0034] See also Figure 2 As shown, the present invention is a geological disaster identification system for karst development areas based on artificial intelligence, which includes the following modules: Data acquisition module: Introducing multi-source heterogeneous data, combined with historical geological disaster cases, using the Bayesian fusion algorithm, and using prior probability and conditional probability to calculate posterior probability; Data preprocessing module: preprocess the data and normalize the data using data normalization method; Geological disaster feature extraction module: extracts groundwater level change rate, soil moisture and vibration frequency features; Principal component analysis dimensionality reduction module: The extracted groundwater level change rate, soil moisture, and vibration frequency features are combined into the original feature matrix. The principal component analysis method is used to construct a projection matrix. The original feature matrix is ​​projected into the principal component space to obtain the reduced dimension feature matrix. Random forest model building module: Use the random forest algorithm to train the model, and the final random forest model RF votes on the prediction results of all decision trees; Model optimization module: Use cross-validation combined with Bayesian algorithm to optimize the constructed random forest model; Risk level mapping module: Real-time features are normalized and PCA dimension reduced before being input into the random forest model. The prediction results are converted into risk level values ​​through mapping rules. GIS risk visualization and dynamic warning module: Generate risk level distribution maps based on GIS data, set thresholds to extract high-risk area boundaries, dynamically update risk levels and issue warning information.

[0035] The above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A method for identifying geological hazards in karst development areas based on artificial intelligence, characterized by: The following steps are involved: S1: Introducing multi-source heterogeneous data, combined with historical geological disaster cases, and using the Bayesian fusion algorithm to calculate the posterior probability using prior probability and conditional probability; S2: preprocess the data and normalize the data using data normalization methods; S3: Extract groundwater level change rate, soil moisture and vibration frequency characteristics; S4: Combine the extracted groundwater level change rate, soil moisture, and vibration frequency features into an original feature matrix. Use the principal component analysis method to construct a projection matrix, project the original feature matrix into the principal component space, and obtain the feature matrix after dimensionality reduction. S5: Use the random forest algorithm to train the model, and the final random forest model RF votes on the prediction results of all decision trees; S6: Use cross-validation combined with Bayesian optimization to build a random forest model; S7: After normalization and PCA dimension reduction, the real-time features are input into the random forest model, and the prediction results are converted into risk level values ​​through mapping rules; S8: Generate a risk level distribution map based on GIS data, set thresholds to extract the boundaries of high-risk areas, dynamically update risk levels and issue warning information.

2. The method for identifying geological hazards in karst development areas based on artificial intelligence according to claim 1, characterized in that: The step S1 comprises the following steps: In view of the geological characteristics of karst development areas, multi-source heterogeneous data fusion technology is introduced to collect geological data and combine it with meteorological data, topographic data and vegetation cover data. At the same time, historical geological disaster cases are collected; Use data fusion algorithms to fuse data from different sources; The data fusion algorithm adopts the Bayesian fusion algorithm to calculate the posterior probability based on the prior probability and conditional probability, specifically: in, is the posterior probability, is the conditional probability, which means the probability of event B occurring under the condition that event A occurs. is the probability of event A, which means the estimation of the probability of event A before there is any observation data. is the probability of event B occurring.

3. The method for identifying geological hazards in karst development areas based on artificial intelligence according to claim 1, characterized in that: The step S2 comprises the following steps: The data is preprocessed, including removing noise and irrelevant data, formatting the data from different sources uniformly, and using a data normalization method to normalize the data to the [a, b] interval. The data normalization method is specifically as follows: in is the normalized sample data, is the sample data, is a data set, and For the normalized lower and upper limits of the interval, that is, the minimum and maximum values ​​of the target interval, respectively, is the minimum value in the sample set, is the maximum value in the sample set.

4. The method for identifying geological hazards in karst development areas based on artificial intelligence according to claim 1, characterized in that: The step S3 comprises the following steps: The groundwater level change rate is characterized by time series analysis; The soil moisture is collected by a soil moisture sensor to extract features; The vibration frequency is converted into a frequency domain signal by Fourier transform to extract the vibration frequency feature.

5. The method for identifying geological hazards in karst development areas based on artificial intelligence according to claim 1, characterized in that: The step S4 comprises the following steps: The extracted groundwater level change rate, soil moisture value and vibration frequency features are combined into the original feature matrix; Normalize the original feature matrix, calculate the covariance matrix of the normalized matrix, perform eigenvalue decomposition on the covariance, obtain the eigenvectors and eigenvalues, and sort them from large to small according to the eigenvalues; Select the first p principal components and construct the projection matrix ,in, For the feature vectors; Project the original feature matrix into the principal component space to obtain the reduced dimension feature matrix, which is: in, is the feature matrix after dimensionality reduction, is the original feature matrix, is the projection matrix.

6. The method for identifying geological hazards in karst development areas based on artificial intelligence according to claim 1, characterized in that: The step S5 comprises the following steps: The reduced feature matrix and the corresponding historical disaster labels are combined into a training dataset D. The historical disaster labels are geological disaster categories, including no disaster, low-risk disaster, medium-risk disaster, and high-risk disaster; The model is trained using the random forest algorithm. The final random forest model RF votes on the prediction results of all decision trees, specifically: in, Input samples for random forest The final prediction result is For the A decision tree for samples The predicted value of is the number of decision trees in the random forest, is the mode, that is, the value that occurs most frequently.

7. The method for identifying geological hazards in karst development areas based on artificial intelligence according to claim 1, characterized in that: The step S6 comprises the following steps: Select K-fold cross validation, define the hyperparameter search space, and use Bayesian optimization to find the optimal hyperparameters in cross validation; Train the random forest model on each fold of the training set, evaluate the performance using the validation set, and record the performance indicators for each validation; Analyze the results of cross-validation, adjust hyperparameters, feature engineering, or model structure, and re-perform cross-validation; Retrain the model on the entire training set using the optimal hyperparameters.

8. The method for identifying geological hazards in karst development areas based on artificial intelligence according to claim 1, characterized in that: The step S7 comprises the following steps: Align the real-time feature vector with the statistics of the training dataset and perform normalization; Use principal component analysis to reduce the dimension of the normalized real-time feature vector to the same dimension as the training model; The real-time feature vector after dimensionality reduction is input into the trained random forest model, and the predicted geological hazard category is output. The geological hazard category is converted into a risk level value according to the mapping rules; The mapping rules are formulated based on historical data and expert experience, specifically: If the model prediction result is no disaster, it means safety and the risk level value is 0; if it is a low-risk disaster, it means low risk and the risk level value is 1; if it is a medium-risk disaster, it means medium risk and the risk level value is 2; if it is a high-risk disaster, it means high risk and the risk level value is 3.

9. The method for identifying geological hazards in karst development areas based on artificial intelligence according to claim 1, characterized in that: The step S8 comprises the following steps: Collect GIS data related to geological hazards, associate the predicted risk level value of each monitoring point with its geographical location, and generate a risk level distribution map; Based on the risk level distribution map, identify high-risk areas. Specifically, set risk level thresholds and use spatial analysis tools to extract high-risk area boundaries. Update the risk level of monitoring points in real time, dynamically adjust the boundaries of high-risk areas, issue early warning information through GIS, and mark high-risk areas.

10. The artificial intelligence-based geological disaster identification system for karst development areas is characterized by: Includes the following modules: Data acquisition module: Introducing multi-source heterogeneous data, combined with historical geological disaster cases, using the Bayesian fusion algorithm, and using prior probability and conditional probability to calculate posterior probability; Data preprocessing module: preprocess the data and normalize the data using data normalization method; Geological disaster feature extraction module: extracts groundwater level change rate, soil moisture and vibration frequency features; Principal component analysis dimensionality reduction module: The extracted groundwater level change rate, soil moisture, and vibration frequency features are combined into the original feature matrix. The principal component analysis method is used to construct a projection matrix. The original feature matrix is ​​projected into the principal component space to obtain the reduced dimension feature matrix. Random forest model building module: Use the random forest algorithm to train the model, and the final random forest model RF votes on the prediction results of all decision trees; Model optimization module: Use cross-validation combined with Bayesian algorithm to optimize the constructed random forest model; Risk level mapping module: Real-time features are normalized and PCA dimension reduced before being input into the random forest model. The prediction results are converted into risk level values ​​through mapping rules. GIS risk visualization and dynamic warning module: Generate risk level distribution maps based on GIS data, set thresholds to extract high-risk area boundaries, dynamically update risk levels and issue warning information.

Citation Information

Cited By

  • Water area karst risk grading method, device and equipment and readable storage medium

    CN121165209A

  • Water area karst risk grading method, device and equipment and readable storage medium

    CN121165209B

  • Multi-source fusion prediction method, device and equipment for karst geological disaster and storage medium

    CN121234175A

  • Karst geological disaster multi-source fusion prediction method, device and equipment and storage medium

    CN121234175B