Hidden fault prediction method and related device based on multi-source detection data fusion
By fusing gravity and magnetic detection data, combined with data augmentation and machine learning algorithms, the problem that traditional geological survey methods are difficult to identify faults in complex terrain areas is solved, and efficient and accurate prediction of hidden faults is achieved.
Patent Information
- Application Number
- CN202411890597.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-12-20
AI Technical Summary
Traditional geological survey methods are difficult to effectively identify faults in complex terrain and thick cover areas, have low efficiency, and lack comprehensive multi-source data analysis methods, so they fail to fully explore geological characteristics and laws.
Using the fusion of gravity and magnetic detection data, the training sample data set is constructed through data augmentation and machine learning algorithms, and the cryptic fracture prediction is performed using random forest and artificial neural network models to generate a fracture distribution map of the cover area.
It improves the recognition ability of hidden fractures, reduces dependence on high-end equipment and professionals, simplifies operating steps, and significantly improves the efficiency and accuracy of fracture prediction.
Smart Images

Figure CN119830209B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a hidden fracture prediction method based on multi-source detection data fusion and related devices. Background Art
[0002] Traditional geological survey methods face numerous challenges in this region, including complex terrain, thick overburden, and surface vegetation cover, making fault identification difficult and inefficient. Advances in remote sensing technology and geophysical methods, particularly the widespread application of gravity and magnetic data in geological research, have provided new avenues for fault identification and the study of regional geological characteristics.
[0003] The study of regional geological characteristics and deep geological structures is a fundamental and highly applied geological survey task, of great significance to socioeconomic development. This research plays a key role in various areas, including lithologic identification and classification, tectonic evolution, mineral prediction and development, and deep geological exploration. Gravity and magnetic data reflect diverse aspects of a region's geological characteristics. Overlaying these data through integrated methods effectively highlights geological features and provides a foundation for comprehensive regional geological analysis. Previous interpretations of gravity and magnetic data have primarily relied on separate analyses of Bouguer gravity and aeromagnetic anomalies, combined with manual interpretation based on geological conditions, resulting in multiple interpretations. Alternatively, joint inversions using multiple geophysical data, including gravity and magnetic, have been performed. Correlative studies of gravity and magnetic anomalies using traditional analytical methods lack the ability to integrate gravity and magnetic interpretation with geological data, resulting in the inability to effectively uncover some geological features, patterns, and genetic mechanisms. The development of big data analysis capabilities, artificial intelligence technologies, and increased computing power have provided new technical means for intelligently mining geological information from gravity and magnetic data, offering new opportunities for intelligent interpretation.
[0004] Currently, machine learning technology allows computers to discern the inherent patterns in data from training examples (i.e., generate models) and identify or predict new input samples. It possesses excellent capabilities for handling complex and nonlinear problems and is an effective means for data classification and regression prediction. In the era of big data, machine learning offers a data-based approach that prioritizes correlation over causality, opening up new avenues for scientific research. Machine learning analyzes correlations across multiple sources and heterogeneous data, thereby uncovering hidden relationships within complex data. This is of great significance for unlocking the hidden value of geological data. In recent years, machine learning has been actively applied in many fields, and some research has also applied machine learning techniques to solve problems in the geosciences, such as earthquake prediction, lithology classification and mapping, remote sensing information extraction, quantitative prediction of mineral resources, and identification and extraction of geochemical anomalies.
[0005] At present, machine methods have become a research hotspot in the geological field. How to use intelligent technologies and methods to process integrated geophysical exploration data, analyze and predict regional geological characteristics and deep geological structures, and improve the efficiency and accuracy of gravity and magnetic anomaly interpretation methods is one of the current research hotspots. The key scientific issues are as follows: (1) In view of the weak geophysical exploration signals in the covered area, how to process gravity and magnetic data and perform data enhancement and diversity characterization on gravity and magnetic information; (2) How to design the hyperparameters of machine learning to achieve the best model effect; (3) For fracture interpretation, it is not that the more features the better. How to select the best features from the data; (4) How to perform refined expression of small fractures. Based on this, gravity and magnetic data are collected on the basis of previous work; after data enhancement, the data are integrated to establish a regional spatial geographic information database. The known fractures are grouped, and the existing data sets are labeled as classification labels. The training set and test set are divided according to a certain ratio. The random forest (RF) and artificial neural network (ANNs) algorithms are used for training, and the parameters of RF and ANNs are optimized for hyperparameters; the selected model is applied to the covered area to predict the unknown faults in the covered area, and the results of fracture interpretation in the covered area are obtained. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies of the prior art. The present invention provides a method and related device for predicting hidden fractures based on the fusion of multi-source detection data, which reduces the dependence on high-end equipment and professional personnel and greatly improves the ability to identify hidden fractures.
[0007] To solve the above technical problems, an embodiment of the present invention provides a method for predicting hidden fractures based on the fusion of multi-source detection data. The method includes:
[0008] Performing detection and processing in a specified area based on gravity instrument equipment and magnetic force instrument equipment to obtain gravity detection data and magnetic force detection data. The specified area is a known fracture area. The gravity detection data is Bouguer gravity anomaly data, and the magnetic force detection data is aeromagnetic anomaly data;
[0009] Fusing the gravity detection data and the magnetic force detection data to form multi-source fusion detection data;
[0010] Performing enhancement processing on the multi-source fusion detection data based on a preset data enhancement algorithm to obtain enhanced multi-source fusion detection data;
[0011] Performing data set construction processing on the enhanced multi-source fusion detection data and performing annotation processing on the constructed data set to form a training sample data set;
[0012] Input the training sample data set into a preset machine learning model for training to obtain a trained and converged machine learning model, where the machine learning model includes a random forest model and an artificial neural network model;
[0013] Apply the trained and converged machine learning model to an unknown fracture target area for buried fracture prediction processing, output the prediction result in the form of spatial coordinates, and generate a buried fracture prediction distribution map covering the unknown fracture target area.
[0014] Optionally, the fusion processing of the gravity detection data and the magnetic detection data to form multi-source fusion detection data includes:
[0015] Perform abnormal data filtering processing on the gravity detection data and the magnetic detection data to obtain filtered gravity detection data and filtered magnetic detection data;
[0016] Perform smoothing and compensation processing on the filtered gravity detection data and the filtered magnetic detection data in sequence to obtain compensated gravity detection data and compensated magnetic detection data;
[0017] Perform fusion processing on the compensated gravity detection data and the compensated magnetic detection data based on a data fusion algorithm to obtain multi-source fusion detection data.
[0018] Optionally, the enhancement processing of the multi-source fusion detection data based on a preset data enhancement algorithm to obtain enhanced multi-source fusion detection data includes:
[0019] Use multiple data enhancement algorithms to perform enhancement processing on the geological structure in the multi-source fusion detection data to obtain enhanced multi-source fusion detection data;
[0020] Where the multiple data enhancement algorithms include the total horizontal derivative method, the vertical derivative method, the vertical second derivative, the tilt angle method, the tilt angle horizontal derivative method, the analytical signal method, the Theat method, the total horizontal derivative tilt angle method, the horizontal tilt angle method, the enhanced analytical signal method, and the enhanced tilt angle method.
[0021] Optionally, the dataset construction processing of the enhanced multi-source fusion detection data and the annotation processing of the constructed dataset to form a training sample dataset include:
[0022] Perform sample synthesis processing on the enhanced multi-source fusion detection data based on the synthetic minority over-sampling technique and perform dataset construction processing based on the synthetic sample data to obtain a constructed dataset;
[0023] Obtain the existing fracture distribution map and geological data of the specified area, and label the constructed data set based on the existing fracture distribution map and geological data to form a training sample data set.
[0024] Optionally, the sample synthesis process for the enhanced multi-source fusion detection data based on the Synthetic Minority Over-sampling Technique (SMOTE) includes:
[0025] For each minority class sample data X in the enhanced multi-source fusion detection data, new samples will be synthesized by calculating the difference between the minority class sample X and its nearest neighbor sample data according to the required oversampling quantity, specifically expressed as: ;
[0026] ;
[0027] where, represents a random function; represents the synthesized sample data.
[0028] Optionally, inputting the training sample data set into a preset machine learning model for training to obtain a trained and converged machine learning model includes:
[0029] When inputting the training sample data set into a preset machine learning model for training, adjust the hyperparameters in the preset machine learning model through a grid search algorithm or a genetic optimization algorithm, select the best parameter combination for training, and obtain a trained and converged machine learning model;
[0030] When the preset machine learning model is a random forest model, the hyperparameters are the number of decision trees, the maximum depth of the decision tree, the minimum number of samples in the leaf node, and the minimum number of samples for node splitting in the random forest model; when the preset machine learning model is an artificial neural network model, the hyperparameters are the number of layers and the number of nodes of the artificial neural network model.
[0031] Optionally, outputting the prediction result in the form of spatial coordinates and generating a predicted map of hidden fractures covering the unknown fracture target area includes:
[0032] Output the prediction result in the form of spatial coordinates, and perform visual analysis on the output prediction result in the form of spatial coordinates in combination with geographic information system software, and generate a predicted map of hidden fractures covering the unknown fracture target area.
[0033] In addition, an embodiment of the present invention also provides a device for predicting hidden fractures based on multi-source detection data fusion, and the device includes:
[0034] Detection module: It is used to perform detection processing in a specified area based on gravity instrument equipment and magnetic instrument equipment to obtain gravity detection data and magnetic detection data. The specified area is a known fracture area. The gravity detection data is Bouguer gravity anomaly data, and the magnetic detection data is aeromagnetic anomaly data;
[0035] Data fusion module: It is used to fuse the gravity detection data and the magnetic detection data to form multi-source fusion detection data;
[0036] Data enhancement module: It is used to perform enhancement processing on the multi-source fusion detection data based on a preset data enhancement algorithm to obtain enhanced multi-source fusion detection data;
[0037] Data construction and annotation module: It is used to perform dataset construction processing on the enhanced multi-source fusion detection data and perform annotation processing on the constructed dataset to form a training sample dataset;
[0038] Model training module: It is used to input the training sample dataset into a preset machine learning model for training processing to obtain a trained and converged machine learning model. The machine learning model includes a random forest model and an artificial neural network model;
[0039] Prediction generation module: It is used to apply the trained and converged machine learning model to an unknown fracture target area for buried fracture prediction processing, output the prediction result in the form of spatial coordinates, and generate a buried fracture prediction distribution map covering the unknown fracture target area.
[0040] In addition, an embodiment of the present invention further provides an electronic device, including a processor and a memory. The processor runs a computer program or code stored in the memory to implement the buried fracture prediction method described in any one of the above.
[0041] In addition, an embodiment of the present invention further provides a computer-readable storage medium for storing a computer program or code. When the computer program or code is executed by a processor, the buried fracture prediction method described in any one of the above is implemented.
[0042] In the embodiment of the present invention, by fusing multi-source detection data such as gravity and magnetism, the underground geological structure can be more comprehensively and deeply reflected; the superposition and comprehensive analysis of multi-source data make the fracture characteristics more obvious, greatly improving the recognition ability of buried fractures; and through intelligent algorithms and automated analysis techniques, the operation steps are simplified, the dependence on high-end equipment and professional personnel is reduced, and the overall cost is significantly reduced; at the same time, by optimizing the analysis process, a large amount of data can be quickly processed, improving the working efficiency of fracture prediction. Description of the Drawings
[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0044] Figure 1 is a schematic flowchart of a method for predicting hidden faults based on multi-source detection data fusion in an embodiment of the present invention;
[0045] Figure 2 is a schematic flowchart of a method for predicting hidden faults based on multi-source detection data fusion in another embodiment of the present invention;
[0046] Figure 3 is a schematic diagram of the structural composition of a device for predicting hidden faults based on multi-source detection data fusion in an embodiment of the present invention;
[0047] Figure 4 is a schematic diagram of the structural composition of an electronic device in an embodiment of the present invention. Specific Embodiments
[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0049] Embodiment 1, please refer to Figure 1 , Figure 1 is a schematic flowchart of a method for predicting hidden faults based on multi-source detection data fusion in an embodiment of the present invention.
[0050] As Figure 1 shown, a method for predicting hidden faults based on multi-source detection data fusion, the method includes:
[0051] S101: Perform detection processing in a specified area based on gravity instrument equipment and magnetic instrument equipment to obtain gravity detection data and magnetic detection data. The specified area is a known fault area. The gravity detection data is Bouguer gravity anomaly data, and the magnetic detection data is aeromagnetic anomaly data;
[0052] In the specific implementation process of the present invention, a designated area needs to be specified, and this designated area is a known fracture area; then, gravity instrument equipment and magnetic instrument equipment are used to conduct detection processing within the designated area, so as to obtain gravity detection data and magnetic detection data; among them, the gravimeter measures the minute changes in the earth's surface gravity field to provide Bouguer gravity anomaly data; the magnetometer records the magnetic field anomalies on the earth's surface to generate aeromagnetic anomaly data; the data is transmitted to the storage device through the data acquisition workstation for subsequent processing.
[0053] S102: Perform fusion processing on the gravity detection data and the magnetic detection data to form multi-source fusion detection data;
[0054] In the specific implementation process of the present invention, the performing of fusion processing on the gravity detection data and the magnetic detection data to form multi-source fusion detection data includes: performing abnormal data filtering processing on the gravity detection data and the magnetic detection data to obtain filtered gravity detection data and filtered magnetic detection data; sequentially performing smoothing and compensation processing on the filtered gravity detection data and the filtered magnetic detection data to obtain compensated gravity detection data and compensated magnetic detection data; performing fusion processing on the compensated gravity detection data and the compensated magnetic detection data based on a data fusion algorithm to obtain multi-source fusion detection data.
[0055] Specifically, for the obtained gravity detection data and magnetic detection data, preprocessing is first required, and the preprocessing here includes data cleaning, filtering, smoothing, and compensation processing, etc.; that is, performing abnormal data filtering processing on the gravity detection data and the magnetic detection data to obtain filtered gravity detection data and filtered magnetic detection data; then sequentially performing smoothing and compensation processing on the filtered gravity detection data and the filtered magnetic detection data to obtain compensated gravity detection data and compensated magnetic detection data; finally, performing fusion processing on the compensated gravity detection data and the compensated magnetic detection data through a data fusion algorithm to obtain multi-source fusion detection data.
[0056] S103: Perform enhancement processing on the multi-source fusion detection data based on a preset data enhancement algorithm to obtain enhanced multi-source fusion detection data;
[0057] In the specific implementation process of the present invention, enhancing the multi-source fusion detection data based on a preset data enhancement algorithm to obtain enhanced multi-source fusion detection data includes: enhancing the geological structure in the multi-source fusion detection data by using a variety of data enhancement algorithms to obtain enhanced multi-source fusion detection data; wherein the variety of data enhancement algorithms include the total horizontal derivative method, the vertical derivative method, the vertical second derivative, the tilt angle method, the tilt angle horizontal derivative method, the analytical signal method, the Theat method, the total horizontal derivative tilt angle method, the horizontal tilt angle method, the enhanced analytical signal method, and the enhanced tilt angle method.
[0058] Specifically, after obtaining the multi-source fusion detection data, it is necessary to enhance the geological structure in the multi-source fusion detection data by using a variety of data enhancement algorithms to obtain enhanced multi-source fusion detection data; wherein the variety of data enhancement algorithms include at least two combinations of the total horizontal derivative method, the vertical derivative method, the vertical second derivative, the tilt angle method, the tilt angle horizontal derivative method, the analytical signal method, the Theat method, the total horizontal derivative tilt angle method, the horizontal tilt angle method, the enhanced analytical signal method, and the enhanced tilt angle method.
[0059] Among them, for the total horizontal derivative method, the horizontal gradient of the gravity anomaly is first used for boundary recognition. Later, this method has been widely used in gravity and magnetic anomalies. Its principle is to assume that the anomaly is the boundary of a vertical body, and the modulus of the horizontal gradient of the anomaly takes the maximum value at the boundary, thereby determining the position of the boundary. However, as the burial depth of the geological body increases and the dip angle changes, the offset of the maximum value position from the boundary will also be greater, but this offset can be ignored when studying the anomalies in the study area; the calculation formula of the total horizontal derivative method can be expressed as:
[0060] ;
[0061] The calculation formula of the vertical derivative method is:
[0062] ;
[0063] The vertical derivative method is different from the total horizontal derivative method. The vertical derivative method determines the boundary by finding the zero value position of the gravity and magnetic anomalies, and as the burial depth increases, its offset from the true position will also be more obvious. Usually, the implementation of the vertical derivative is carried out in the frequency domain. Therefore, for the stability of the results, a low-pass filtering method such as upward continuation and the Winer filtering method is generally used to suppress high-frequency interference. At the same time, the vertical second derivative method can also be used, but increasing the order of the derivative may cause the identified boundary to be closer to the center position; as follows:
[0064] ;
[0065] The tilt angle method is also called the oblique derivative method, and its calculation formula is:
[0066] ;
[0067] Among them, and are obtained by the vertical derivative method and the total horizontal derivative method respectively; the tilt angle method is a boundary enhancement method with better balance; the ratio of the abnormal vertical derivative to the total horizontal derivative is used, which well balances the anomalies with high amplitude and low amplitude, and appropriately stretches the geological bodies with different burial depths. The first two methods perform poorly in identifying deep-source anomalies, while the tilt angle method is insensitive to the depth of the field source due to its balance and can well identify the boundary; however, the tilt angle method also has defects and may have the phenomenon of "analytical singularity" affecting the results; further obtaining the total horizontal derivative of the tilt angle not only eliminates the "analytical singularity", but also determines the boundary using the maximum position to improve the lateral resolution of identification; the improved calculation formula can be expressed as:
[0068] ;
[0069] The analytical signal method is also called the total gradient modulus method. Initially, the analytical signal was only used for two-dimensional space processing, and then the three-dimensional analytical signal method was developed; the analytical signal method has its own advantages when processing magnetic anomalies. It is not affected or has little influence by the magnetization direction and magnetic anomaly components. However, at the same time, its lateral resolution is very low compared with other methods. Therefore, improving its lateral resolution is also the focus of future research; the calculation formula of the analytical signal method is:
[0070] ;
[0071] Among them, and are obtained from the above formula respectively.
[0072] The calculation formula of the Theta map method is:
[0073] ;
[0074] Among them, THDR and ASM are obtained from the above respectively.
[0075] The Theta map identifies the boundary by taking the maximum position, and also well balances the anomaly amplitude values from different depths, enhancing both shallow-source anomalies and deep-source anomalies; it is considered that this method is affected by the magnetic anomaly components and magnetization direction to the same extent as the vertical derivative method and the tilt angle method; however, the Theta map also has the same "analytical singularity" problem as the tilt angle method, which can be improved by introducing a regularization factor.
[0076] The calculation formula of the horizontal tilt angle method is:
[0077] ;
[0078] Among them, THDR and VDR are given above; contrary to the amplitude of the normalized vertical derivative by the tilt angle method, the amplitude of the normalized horizontal derivative by the horizontal tilt angle method shows equally good effects on both deep and shallow anomalies in boundary recognition and is a very useful method for enhancing features in a given direction.
[0079] The enhanced analytical signal tilt angle method is an equalization filter formed based on the tilt angle method and the analytical signal method, aiming to be able to handle the situation where both positive and negative anomalies coexist. Its calculation formula is:
[0080] ;
[0081] Among them, ASM is given by the above formula.
[0082] At the same time, a method to increase the resolution is also proposed. A filter is defined using the vertical first derivative of the analytical signal amplitude:
[0083] ;
[0084] Among them, VAS is the vertical first derivative of ASM, that is:
[0085] ;
[0086] Generally, it is considered that the TAAS and TAVAS methods identify boundaries more accurately than traditional methods, and when both positive and negative anomalies are included, no redundant and incorrect boundary information will be generated. However, due to the calculation of high-order derivatives, TAAS and TAVAS are easily affected by external noise. Therefore, in order to obtain accurate fracture information, a filtering operation needs to be performed before preprocessing.
[0087] The calculation formula of the enhanced tilt angle method is:
[0088] ;
[0089] The enhanced tilt angle method is similar to the tilt angle method. The zero value provides the potential field boundary and also realizes equalization through the normalization of anomalies with different field source depths; its greatest advantage lies in being able to improve the accuracy of identifying small-scale linear structures from superimposed anomalies and can highlight the linear features of weak changes in practical applications; similarly, TAVDR is also affected by noise, so filtering should also be performed before enhancement.
[0090] In this embodiment, according to the abnormal discontinuity exhibited by gravity and magnetism at geological boundaries, the boundary is reasonably extracted by utilizing the large change rate of gravity and magnetic anomalies at the boundary, providing a basis for the subsequent identification of faults; the common methods for gravity and magnetic boundary extraction can generally be divided into mathematical statistics methods (such as small sub-domain filtering method, normalized standard deviation method, etc.), numerical calculation methods (methods of various order derivatives of potential field data), and other methods (such as directional filtering method, etc.); the most commonly used numerical calculation method for boundary detection is used, including a total of 12 methods including the first-order derivative and higher-order derivatives of the gravity and magnetic fields, and the enhanced result will be used as the characteristic data for fault identification.
[0091] S104: Process the enhanced multi-source fusion detection data to construct a data set, and perform annotation processing on the constructed data set to form a training sample data set;
[0092] In the specific implementation process of the present invention, the processing of constructing a data set from the enhanced multi-source fusion detection data and performing annotation processing on the constructed data set to form a training sample data set includes: performing sample synthesis processing on the enhanced multi-source fusion detection data based on the synthetic minority over-sampling algorithm, and performing data set construction processing based on the synthetic sample data to obtain the constructed data set; obtaining the existing fracture distribution map and geological data of the specified area, and performing annotation processing on the constructed data set based on the existing fracture distribution map and geological data to form a training sample data set.
[0093] Further, the performing sample synthesis processing on the enhanced multi-source fusion detection data based on the synthetic minority over-sampling algorithm includes:
[0094] For each minority class sample data X in the enhanced multi-source fusion detection data, new samples will be synthesized by calculating the difference between the minority class sample X and its nearest neighbor sample data according to the required oversampling quantity, which is specifically expressed as: ;
[0095] ;
[0096] where, represents a random function; represents the synthesized sample data.
[0097] Specifically, for the enhanced multi-source fusion detection data, this embodiment uses the Synthetic Minority Over-Sampling Technique (SMOTE) to address the problem of sample imbalance, that is, to increase the data volume of the minority class samples, so as to achieve the purpose of sample balance. The principle of the SMOTE method is to create synthetic samples instead of randomly replacing the oversampled ones; in this embodiment, for each minority class sample X, based on the required oversampling quantity, using the breakpoints of each group as the sampling benchmark, new samples are synthesized by calculating the difference between the minority samples and their nearest neighbor samples' data The specific expression is:
[0098] ;
[0099] Wherein, represents a random function; represents the synthesized sample data; the SMOTE method can not only fully learn the characteristics of the minority samples and improve the accuracy of the classifier, but also avoid the overfitting phenomenon caused by the random oversampling method.
[0100] S105: Input the training sample data set into a preset machine learning model for training processing to obtain a machine learning model with training convergence, and the machine learning model includes a random forest model and an artificial neural network model;
[0101] In the specific implementation process of the present invention, the step of inputting the training sample data set into a preset machine learning model for training processing to obtain a machine learning model with training convergence includes: when inputting the training sample data set into a preset machine learning model for training processing, adjusting the hyperparameters in the preset machine learning model through a grid search algorithm or a genetic optimization algorithm, selecting the best parameter combination for training processing to obtain a machine learning model with training convergence; when the preset machine learning model is a random forest model, the hyperparameters are the number of decision trees, the maximum depth of the decision tree, the minimum number of samples contained in the leaf nodes, and the minimum number of samples that can be divided at the nodes in the random forest model; when the preset machine learning model is an artificial neural network model, the hyperparameters are the number of layers and the number of nodes of the artificial neural network model.
[0102] Specifically, a random forest model and an artificial neural network model are selected as the machine learning models; the random forest model can effectively avoid the overfitting problem through the integration of multiple decision trees; the artificial neural network model has the ability to handle complex non-linear problems.
[0103] When inputting the training sample data set into a preset machine learning model for training, the hyperparameters in the preset machine learning model are adjusted by a grid search algorithm or a genetic optimization algorithm, and the best parameter combination is selected for training to obtain a machine learning model with training convergence; that is, by methods such as grid search or genetic optimization algorithm, the hyperparameters of the model (the number of decision trees, the maximum depth of decision trees, the minimum number of samples in leaf nodes, the minimum number of samples for node splitting, the number of layers and nodes of neural networks, etc.) are adjusted, and the best parameter combination is selected.
[0104] Among them, the artificial neural network model is a supervised feedforward neural network that maps a set of input vectors to output vectors through hidden layers and consists of three node layers: an input layer, one or more hidden layers, and an output layer; the role of the input layer is to receive predictor variables and pass them to the hidden layer. The hidden layer consists of one or more network layers, where each neuron is connected to the neurons in the previous and subsequent layers. The number of neurons in the output layer is the same as the number of categories in the actual classification problem, and it is essentially a multi-class logistic regression process. The hidden layer and the output layer calculate the output signal of this layer through an activation function based on the received signal from the previous layer; the calculation process of the artificial neural network is outlined as follows. First, calculate the output signal:
[0105] ;
[0106] Among them, represents the activation function; represents the weighted value; represents the input; b represents the bias term; in order to obtain the most suitable weights and bias terms, the backpropagation algorithm is used for iterative calculation; for this purpose, the loss function is defined:
[0107] ;
[0108] Among them, represents the true label (the probability that the known sample i belongs to the j-th class); represents the output of the model (the probability that the predicted sample i belongs to the j-th class). Calculate the gradients of the loss function with respect to the weights and bias terms layer by layer from back to front, that is:
[0109] ;
[0110] ;
[0111] Among them, and respectively represent the weights and bias terms of the n-th neuron in the l -th layer; denotes the learning rate; the neural network randomly initializes the weights and bias values, and then updates the weights and bias terms of each layer according to the above formula, continuously iterating the forward propagation and backward propagation processes until the stopping criterion is met, i.e., the set number of loops or error threshold, and finally achieving the purpose of solving.
[0112] The random forest model is an ensemble learning algorithm based on decision trees. By generating a large number of decision trees (weak classifiers), it independently learns and predicts, and then votes on the results of these decision trees to form a strong classifier. A decision tree is a tree structure, where each internal node represents an attribute judgment, each branch represents a judgment result, and each leaf node represents a classification result; common decision tree algorithms include C4.5, ID3, and CART, and CART decision trees are usually used in random forests. CART decision (Classification and Regression Tree) can be used for classification and regression tasks. If the dataset contains continuous variables, the decision tree is used as a regression tree and the mean of the leaf nodes is used for prediction; if the dependent variable is a discrete value, the decision tree is used as a classifier to solve the classification problem; in a random forest, the decision tree traverses all possible split points of the feature variables to find the split point with the smallest Gini impurity coefficient and divides the data into two subsets; this process is repeated continuously until the predefined stopping condition is met, thus completing the purification of a single decision tree. The random forest generates decision trees through the following two random processes:
[0113] Randomly generate the training set: Use the bootstrap method to randomly sample with replacement to generate the training subset data for each tree. This process avoids high correlations between trees and uses the bagging method to replace the original data for random sampling.
[0114] Randomly select feature attributes: When splitting the nodes of the tree, randomly select features without replacement to generate a feature subset, select the best feature suitable for splitting from it, and continue with feature selection and tree splitting.
[0115] These two random processes ensure the randomness of the entire algorithm and the independence of each decision tree, without any pruning, improving the classification accuracy. The calculation process of the random forest algorithm is briefly outlined as follows:
[0116] In the training data X, generate a subsample subset through random sampling, and select the corresponding feature subset , where n is the number of decision trees; the sample and the randomly selected features Train the decision tree ; the decision tree Votes for each class, and the voting result is used as the final classification value; the voting classification result of a certain sample can be expressed as:
[0117] ;
[0118] Among them, represents the lithology category predicted by the i-th decision tree; is its voting result; the random forest model can evaluate the importance of the predictive variables, and the corresponding calculation formula is calibrated as:
[0119] ;
[0120] Among them, t, k, and m are the number of nodes of each tree, the number of decision trees, and the number of predictive variables, respectively; is the reduction value of the Gini coefficient at the j-th node of the r-th predictive variable in the i-th tree.
[0121] Usually, the random forest will generate hundreds to thousands of decision trees and return the final result through voting; the goal of the random forest is not to pursue the optimal decision tree, but to assemble a set of predictors so that it can benefit from the extensive search of all tree predictors.
[0122] After training is completed, the performance of the trained machine learning model is evaluated through performance metrics to see if it meets the requirements; in the embodiments of this application, five metrics are used: accuracy, precision, recall, F1 score, and AUC value; usually, in order to describe these metrics, the following confusion matrix of binary classification data needs to be introduced, as shown in Table 1:
[0123] Table 1 Confusion Matrix of Binary Classification Data
[0124]
[0125] Among them, TP is the correctly classified positive sample, FP is the misclassified negative sample, FN is the misclassified positive sample, and TN is the correctly classified negative sample.
[0126] Accuracy: Measures the overall classification accuracy of the test set, indicating the ratio of the number of samples that match the label in the prediction results to the total number of samples. Its calculation formula is:
[0127] ;
[0128] Precision: Refers to the proportion of samples that are correctly predicted and classified among all samples predicted as positive. Its calculation formula is:
[0129] ;
[0130] Recall: It refers to the proportion of samples correctly classified by prediction among all samples that are truly positive. Its calculation formula is:
[0131] ;
[0132] F1 Score: It is a combination of precision and recall. Since there are contradictory situations in actual cases where precision is high while recall is low, the F1 score is introduced to reconcile the evaluation of the two. The calculation formula is:
[0133] ;
[0134] AUC Value: When dealing with imbalanced datasets, the performance of the random forest algorithm is not good, tending to give higher weights to the majority class; so it is necessary to introduce the Receiver Operating Characteristic (ROC) curve, which is a useful tool for evaluating classification performance; First, it is necessary to calculate the false positive rate (FPR) of misclassified negative samples and the true positive rate (TPR) of correctly classified positive samples. The values of both range from 0 to 1, and their calculation formulas are as follows:
[0135] ;
[0136] ;
[0137] Taking the FPR as the abscissa and the TPR as the ordinate, the ROC curve is established. The closer the curve is to the point (0, 1), that is, the TPR approaches 1 while the FPR approaches 0, the more perfect the classifier is. On the contrary, if the curve is closer to the point (1, 0), then the performance of this classifier is worse. The area value under the curve is the AUC value. Usually, its value range is between 0.5 and 1. The closer the AUC value is to 1, the better the performance of the classifier.
[0138] S106: Apply the trained and converged machine learning model to the unknown fracture target area for concealed fracture prediction processing, output the prediction results in the form of spatial coordinates, and generate a concealed fracture prediction distribution map covering the unknown fracture target area.
[0139] In the specific implementation process of the present invention, the step of outputting the prediction results in the form of spatial coordinates and generating a concealed fracture prediction distribution map covering the unknown fracture target area includes: outputting the prediction results in the form of spatial coordinates, and performing visual analysis on the output prediction results in the form of spatial coordinates in combination with geographic information system software, and generating a concealed fracture prediction distribution map covering the unknown fracture target area.
[0140] Specifically, apply the machine learning model with training convergence to an unknown research area for predicting hidden fractures in the unknown fracture target area; output the predicted results in the form of spatial coordinates, and combine with Geographic Information System (GIS) software to visually analyze the predicted results and generate a fracture distribution map of the coverage area.
[0141] In the embodiment of the present invention, by fusing multi-source detection data such as gravity and magnetism, the underground geological structure can be more comprehensively and deeply reflected; the superposition and comprehensive analysis of multi-source data make the fracture characteristics more obvious, greatly improving the recognition ability of hidden fractures; and through intelligent algorithms and automated analysis techniques, the operation steps are simplified, the dependence on high-end equipment and professional personnel is reduced, and the overall cost is significantly reduced; at the same time, by optimizing the analysis process, a large amount of data can be quickly processed, improving the working efficiency of fracture prediction.
[0142] Embodiment 2, please refer to Figure 2 , Figure 2 which is a schematic flowchart of a method for predicting hidden fractures based on multi-source detection data fusion in another embodiment of the present invention.
[0143] As Figure 2 shown, a method for predicting hidden fractures based on multi-source detection data fusion, the method includes:
[0144] S201: Conduct detection processing in a specified area based on gravity instrument equipment and magnetic instrument equipment to obtain gravity detection data and magnetic detection data, the specified area is a known fracture area, the gravity detection data is Bouguer gravity anomaly data, and the magnetic detection data is aeromagnetic anomaly data;
[0145] S202: Fuse the gravity detection data and the magnetic detection data to form multi-source fusion detection data;
[0146] S203: Perform enhancement processing on the multi-source fusion detection data based on a preset data enhancement algorithm to obtain enhanced multi-source fusion detection data;
[0147] S204: Perform sample synthesis processing on the enhanced multi-source fusion detection data based on the Synthetic Minority Over-sampling Technique (SMOTE) algorithm, and perform dataset construction processing based on the synthesized sample data to obtain the constructed dataset;
[0148] S205: Obtain the existing fracture distribution map and geological data of the specified area, and perform annotation processing on the constructed dataset based on the existing fracture distribution map and geological data to form a training sample dataset;
[0149] S206: Input the training sample data set into a preset machine learning model for training to obtain a trained and converged machine learning model, where the machine learning model includes a random forest model and an artificial neural network model;
[0150] S207: Apply the trained and converged machine learning model to an unknown fracture target area for buried fracture prediction, output the prediction result in the form of spatial coordinates, and generate a buried fracture prediction distribution map covering the unknown fracture target area.
[0151] Specifically, for the specific implementation manner of Embodiment 2, reference can be made to Embodiment 1, which will not be elaborated here.
[0152] Embodiment 3, please refer to Figure 3 , Figure 3 is a schematic structural composition diagram of a buried fracture prediction device based on multi-source detection data fusion in an embodiment of the present invention.
[0153] As Figure 3 shown, a buried fracture prediction device based on multi-source detection data fusion, the device includes:
[0154] Detection module 301: Used to perform detection processing in a specified area based on a gravity instrument device and a magnetic instrument device to obtain gravity detection data and magnetic detection data. The specified area is a known fracture area, the gravity detection data is Bouguer gravity anomaly data, and the magnetic detection data is aeromagnetic anomaly data;
[0155] In the specific implementation process of the present invention, a specified area needs to be specified, and this specified area is a known fracture area; then use a gravity instrument device and a magnetic instrument device to perform detection processing in the specified area to obtain gravity detection data and magnetic detection data; among them, the gravimeter provides Bouguer gravity anomaly data by measuring the minute changes in the earth's surface gravity field; the magnetometer records the magnetic field anomalies on the earth's surface to generate aeromagnetic anomaly data; the data is transmitted to a storage device through a data acquisition workstation for subsequent processing.
[0156] Data fusion module 302: Used to fuse the gravity detection data and the magnetic detection data to form multi-source fusion detection data;
[0157] In the specific implementation process of the present invention, the fusion process of the gravity detection data and the magnetic detection data to form multi-source fusion detection data includes: filtering the abnormal data in the gravity detection data and the magnetic detection data to obtain the filtered gravity detection data and the filtered magnetic detection data; sequentially performing smoothing and compensation processing on the filtered gravity detection data and the filtered magnetic detection data to obtain the compensated gravity detection data and the compensated magnetic detection data; and performing fusion processing on the compensated gravity detection data and the compensated magnetic detection data based on a data fusion algorithm to obtain multi-source fusion detection data.
[0158] Specifically, for the detected gravity detection data and magnetic detection data, preprocessing is first required. Here, the preprocessing includes data cleaning, filtering, smoothing, and compensation processing, etc.; that is, filtering the abnormal data in the gravity detection data and the magnetic detection data to obtain the filtered gravity detection data and the filtered magnetic detection data; then sequentially performing smoothing and compensation processing on the filtered gravity detection data and the filtered magnetic detection data to obtain the compensated gravity detection data and the compensated magnetic detection data; and finally performing fusion processing on the compensated gravity detection data and the compensated magnetic detection data through a data fusion algorithm to obtain multi-source fusion detection data.
[0159] Data enhancement module 303: used to perform enhancement processing on the multi-source fusion detection data based on a preset data enhancement algorithm to obtain enhanced multi-source fusion detection data;
[0160] In the specific implementation process of the present invention, the enhancement processing of the multi-source fusion detection data based on a preset data enhancement algorithm to obtain enhanced multi-source fusion detection data includes: using a variety of data enhancement algorithms to perform enhancement processing on the geological structure in the multi-source fusion detection data to obtain enhanced multi-source fusion detection data; where the variety of data enhancement algorithms include the total horizontal derivative method, the vertical derivative method, the vertical second derivative, the tilt angle method, the tilt angle horizontal derivative method, the analytical signal method, the Theat method, the total horizontal derivative tilt angle method, the horizontal tilt angle method, the enhanced analytical signal method, and the enhanced tilt angle method.
[0161] Specifically, after obtaining the multi-source fusion detection data, it is necessary to perform enhancement processing on the geological structure in the multi-source fusion detection data through a variety of data enhancement algorithms to obtain enhanced multi-source fusion detection data; where the variety of data enhancement algorithms include at least two combinations of the total horizontal derivative method, the vertical derivative method, the vertical second derivative, the tilt angle method, the tilt angle horizontal derivative method, the analytical signal method, the Theat method, the total horizontal derivative tilt angle method, the horizontal tilt angle method, the enhanced analytical signal method, and the enhanced tilt angle method.
[0162] Among them, in the total horizontal derivative method, the horizontal gradient of the gravity anomaly is first used for boundary identification. Later, this method was widely applied to gravity and magnetic anomalies. Its principle is to assume that the anomaly is the boundary of a vertical body, and the modulus of the horizontal gradient of the anomaly reaches its maximum value at the boundary, thereby determining the position of the boundary. However, as the burial depth of the geological body increases and the dip angle changes, the offset of the maximum value position from the boundary will also become larger. However, this offset can be ignored when studying the anomalies in the study area. The calculation formula of the total horizontal derivative method can be expressed as:
[0163] ;
[0164] The calculation formula of the vertical derivative method is:
[0165] ;
[0166] The vertical derivative method is different from the total horizontal derivative method. The vertical derivative method determines the boundary by finding the zero value position of the gravity and magnetic anomalies, and as the burial depth increases, its offset from the true position will also be more obvious. Usually, the implementation of the vertical derivative is carried out in the frequency domain. Therefore, for the stability of the results, methods such as upward continuation and Winer filtering, which are low-pass filtering methods, are generally used to suppress high-frequency interference. At the same time, the vertical second derivative method can also be used, but increasing the order of the derivative may cause the identified boundary to be closer to the center position; as follows:
[0167] ;
[0168] The tilt angle method, also known as the oblique derivative method, has the following calculation formula:
[0169] ;
[0170] Among them, 、 are obtained from the vertical derivative method and the total horizontal derivative method respectively; the tilt angle method is a boundary enhancement method with better balance; it uses the ratio of the vertical derivative and the total horizontal derivative of the anomaly, and its function is to well balance the anomalies with high amplitude and low amplitude, and appropriately stretch the geological bodies with different burial depths. The first two methods perform poorly in identifying deep-source anomalies, while the tilt angle method is insensitive to the depth of the field source due to its balance and can well identify the boundary; however, the tilt angle method also has defects, and the "analytical singularity" phenomenon may occur, affecting the results; further taking the total horizontal derivative of the tilt angle not only eliminates the "analytical singularity", but also determines the boundary using the maximum value position, improving the horizontal resolution of the identification; the improved calculation formula can be expressed as:
[0171] ;
[0172] The analytical signal method, also known as the total gradient modulus method, was initially only used for two-dimensional space processing, and later the three-dimensional analytical signal method was developed. The analytical signal method has its own advantages when processing magnetic anomalies. It is not affected or only slightly affected by the magnetization direction and magnetic anomaly components. However, its lateral resolution is very low compared to other methods. Therefore, improving its lateral resolution is also the focus of future research. The calculation formula of the analytical signal method is:
[0173] ;
[0174] where, 、 are respectively obtained from the above formula.
[0175] The calculation formula of the Theta map method is:
[0176] ;
[0177] where THDR and ASM are respectively obtained from the above.
[0178] The Theta map identifies the boundary by taking the maximum position, and also well balances the anomaly amplitude values from different depths, enhancing both shallow-source and deep-source anomalies. It is considered that this method is affected by magnetic anomaly components and magnetization direction to the same extent as the vertical derivative method and the tilt angle method. However, the Theta map also has the same "analytical singularity" problem as the tilt angle method, which can be improved by introducing a regularization factor.
[0179] The calculation formula of the horizontal tilt angle method is:
[0180] ;
[0181] where THDR and VDR are respectively given above; contrary to the normalized vertical derivative amplitude of the tilt angle method, the horizontal tilt angle method normalizes the amplitude of the horizontal derivative and shows equally good effects on deep and shallow anomalies in boundary recognition. It is a very useful method for enhancing features in a given direction.
[0182] The enhanced analytical signal tilt angle method is an equalization filter formed based on the tilt angle method and the analytical signal method, aiming to be able to handle the situation where positive and negative anomalies coexist. Its calculation formula is:
[0183] ;
[0184] where ASM is given by the above formula.
[0185] At the same time, a method to increase the resolution is also proposed, defining a filter using the vertical first-order derivative of the analytical signal amplitude:
[0186] ;
[0187] where VAS is the first vertical derivative of ASM, i.e.:
[0188] ;
[0189] Generally, it is considered that the TAAS and TAVAS methods identify more accurate boundaries than traditional methods, and when both positive and negative anomalies are included, no redundant and incorrect boundary information will be generated. However, due to the calculation of high-order derivatives, TAAS and TAVAS are susceptible to external noise. Therefore, in order to obtain accurate fracture information, a filtering operation needs to be performed before preprocessing.
[0190] The calculation formula of the enhanced tilt angle method is:
[0191] ;
[0192] The enhanced tilt angle method is similar to the tilt angle method. The zero value provides the potential field boundary and also achieves balance through the normalization of anomalies with different field source depths. Its greatest advantage lies in that it can improve the accuracy of identifying small-scale linear structures from superimposed anomalies and can highlight the linear features of weak changes in practical applications. Similarly, TAVDR is also affected by noise, so filtering should also be performed before enhancement.
[0193] In this embodiment, according to the discontinuity of gravity and magnetism at the geological boundary, the boundary is reasonably extracted by using the large change rate of gravity and magnetic anomalies at the boundary, providing a basis for the subsequent identification of faults. The common methods for gravity and magnetic boundary extraction can generally be divided into mathematical statistics methods (such as small sub-domain filtering method, normalized standard deviation method, etc.), numerical calculation methods (methods of various order derivatives of potential field data), and other methods (such as direction filtering method, etc.). The most commonly used numerical calculation method for boundary detection is used, including a total of 12 methods including the first-order derivative and high-order derivatives of the gravity and magnetic fields. The enhanced result will be used as the feature data for fault identification.
[0194] Data construction and annotation module 304: used to perform data set construction processing on the enhanced multi-source fusion detection data, and perform annotation processing on the constructed data set to form a training sample data set;
[0195] In the specific implementation process of the present invention, the enhanced multi-source fusion detection data is subjected to dataset construction processing, and the constructed dataset is subjected to annotation processing to form a training sample dataset, including: performing sample synthesis processing on the enhanced multi-source fusion detection data based on the Synthetic Minority Over-sampling Technique (SMOTE), and performing dataset construction processing based on the synthesized sample data to obtain the constructed dataset; obtaining the existing fracture distribution map and geological data of the specified area, and performing annotation processing on the constructed dataset based on the existing fracture distribution map and geological data to form a training sample dataset.
[0196] Further, the performing sample synthesis processing on the enhanced multi-source fusion detection data based on the Synthetic Minority Over-sampling Technique (SMOTE) includes:
[0197] For each minority class sample data X in the enhanced multi-source fusion detection data, new samples will be synthesized by calculating the difference between the minority class sample X and its nearest neighbor sample data according to the required oversampling quantity, specifically expressed as:
[0198] ;
[0199] Wherein, represents a random function; represents the synthesized sample data.
[0200] Specifically, for the enhanced multi-source fusion detection data, in this embodiment, the Synthetic Minority Over-sampling Technique (SMOTE) is used to address the problem of sample imbalance, that is, to increase the data volume of the minority class samples, so as to achieve the purpose of sample balance. The principle of the SMOTE method is to create synthetic samples instead of randomly replacing oversampled samples; in this embodiment, for each minority class sample X, according to the required oversampling quantity, taking each group of fracture points as the sampling benchmark, new samples will be synthesized by calculating the difference between the minority sample and its nearest neighbor sample data, specifically expressed as:
[0201] ;
[0202] Wherein, represents a random function; represents the synthesized sample data; the SMOTE method can not only fully learn the features of the minority samples and improve the accuracy of the classifier, but also avoid the overfitting phenomenon caused by the random oversampling method.
[0203] Model training module 305: It is used to input the training sample data set into a preset machine learning model for training processing to obtain a machine learning model with training convergence. The machine learning model includes a random forest model and an artificial neural network model;
[0204] In the specific implementation process of the present invention, the step of inputting the training sample data set into a preset machine learning model for training processing to obtain a machine learning model with training convergence includes: when inputting the training sample data set into a preset machine learning model for training processing, adjusting the hyperparameters in the preset machine learning model through a grid search algorithm or a genetic optimization algorithm, selecting the best parameter combination for training processing, and obtaining a machine learning model with training convergence; when the preset machine learning model is a random forest model, the hyperparameters are the number of decision trees, the maximum depth of the decision tree, the minimum number of samples contained in the leaf node, and the minimum number of samples that can be divided by the node in the random forest model; when the preset machine learning model is an artificial neural network model, the hyperparameters are the number of layers and the number of nodes of the artificial neural network model.
[0205] Specifically, a random forest model and an artificial neural network model are selected as the machine learning models; the random forest model can effectively avoid the overfitting problem through the integration of multiple decision trees; the artificial neural network model has the ability to handle complex non-linear problems.
[0206] When inputting the training sample data set into a preset machine learning model for training processing, adjusting the hyperparameters in the preset machine learning model through a grid search algorithm or a genetic optimization algorithm, selecting the best parameter combination for training processing, and obtaining a machine learning model with training convergence; that is, through methods such as grid search or genetic optimization algorithm, adjust the hyperparameters of the model (the number of decision trees, the maximum depth of the decision tree, the minimum number of samples contained in the leaf node, the minimum number of samples that can be divided by the node, the number of layers and the number of nodes of the neural network, etc.), and select the best parameter combination.
[0207] The artificial neural network model is a supervised feedforward neural network, which maps a set of input vectors to output vectors through hidden layers and consists of three node layers: an input layer, one or more hidden layers, and an output layer; the role of the input layer is to receive the predictor variables and pass them to the hidden layer. The hidden layer consists of one or more network layers, and each neuron in it is connected to the neurons in the previous and subsequent layers. The number of neurons in the output layer is the same as the number of categories of the actual classification problem. It is essentially a multi-class logistic regression process. The hidden layer and the output layer calculate the output signal of this layer through an activation function based on the received signal of the previous layer; the calculation process of the artificial neural network is outlined as follows. First, calculate the output signal:
[0208] ;
[0209] Among them, represents the activation function; represents the weight value; represents the input; b represents the bias term; in order to obtain the most suitable weights and bias terms, the backpropagation algorithm is used for iterative calculation; for this purpose, a loss function is defined:
[0210] ;
[0211] Among them, represents the true label (the probability that the known sample i belongs to the j-th class); represents the output of the model (the probability that the predicted sample i belongs to the j-th class). Calculate the gradients of the loss function with respect to the weights and bias terms layer by layer from back to front, that is:
[0212] ;
[0213] ;
[0214] Among them, and respectively represent the weights and bias terms of the n-th neuron in the l layer; represents the learning rate; the neural network randomly initializes the weight and bias term values, and then updates the weights and bias terms of each layer according to the above formula, continuously iterating the forward propagation and backpropagation processes until the stopping criterion is met, that is, the set number of loops or error threshold, and finally achieving the purpose of solving.
[0215] The random forest model is an ensemble learning algorithm based on decision trees. By generating a large number of decision trees (weak classifiers), it independently learns and predicts, and then votes on the results of these decision trees to form a strong classifier. A decision tree is a tree structure, where each internal node represents an attribute judgment, each branch represents a judgment result, and each leaf node represents a classification result; common decision tree algorithms include C4.5, ID3, and CART, among which CART decision trees are usually used in random forests. CART decision (Classification and Regression Tree) can be used for classification and regression tasks. If the data set contains continuous variables, the decision tree is used as a regression tree and the mean of the leaf nodes is used for prediction; if the dependent variable is a discrete value, the decision tree is used as a classifier to solve classification problems; the decision trees in the random forest search for the splitting point with the smallest Gini impurity coefficient by traversing all possible splitting points of the feature variables, and divide the data into two subsets; this process is repeated continuously until the predefined stopping condition is met, thus completing the purification of a single decision tree. The random forest generates decision trees through the following two random processes:
[0216] Randomly generate the training set: Use the Bootstrap method to randomly sample with replacement to generate the training subset data for each tree. This process avoids high correlations between trees and uses the Bagging method to replace the original data for random sampling.
[0217] Randomly select feature attributes: When splitting the nodes of the tree, randomly select features without replacement to generate a feature subset, select the best feature suitable for splitting from it, and continue with feature selection and tree splitting.
[0218] These two random processes ensure the randomness of the entire algorithm and the independence of each decision tree. Without any pruning, the classification accuracy is improved. The calculation process of the random forest algorithm is briefly outlined as follows:
[0219] In the training data X, generate a subsample subset through random sampling , and select the corresponding feature subset , where n is the number of decision trees; the sample and the randomly selected features Train the decision tree ; the decision tree Votes for each class, and use the voting result as the final classification value; the voting classification result of a certain sample can be expressed as:
[0220] ;
[0221] Among them, represents the lithology class predicted by the i-th decision tree; is its voting result; the random forest model can evaluate the importance of the predictor variables, and the corresponding calculation formula is calibrated as:
[0222] ;
[0223] Among them, t, k, and m are the number of nodes of each tree, the number of decision trees, and the number of predictor variables respectively; is the decrease in the Gini coefficient of the r-th predictor variable at the j-th node of the i-th tree.
[0224] Usually, the random forest will generate hundreds to thousands of decision trees and return the final result through voting; the goal of the random forest is not to pursue the optimal decision tree, but to assemble a set of predictors so that it can benefit from the extensive search of all tree predictors.
[0225] After training is completed, evaluate whether the performance of the trained machine learning model meets the requirements through performance metrics; in the embodiments of this application, five metrics are used: accuracy, precision, recall, F1 score, and AUC value; generally, in order to describe these metrics, the following confusion matrix for binary classification data needs to be introduced, as shown in Table 1:
[0226] Table 1 Confusion Matrix for Binary Classification Data
[0227]
[0228] Among them, TP is the correctly classified positive sample, FP is the misclassified negative sample, FN is the misclassified positive sample, and TN is the correctly classified negative sample.
[0229] Accuracy: Measures the overall classification accuracy of the test set, indicating the ratio of the number of samples that match the label in the prediction results to the total number of samples. Its calculation formula is:
[0230] ;
[0231] Precision: Refers to the proportion of samples that are correctly predicted and classified among all samples predicted as positive. Its calculation formula is:
[0232] ;
[0233] Recall: Refers to the proportion of samples that are correctly predicted and classified among all samples that are actually positive. Its calculation formula is:
[0234] ;
[0235] F1 Score: It is a combination of precision and recall. Since there are conflicting situations in actual cases where precision is high and recall is low, the F1 score is introduced to reconcile the evaluation of the two. The calculation formula is:
[0236] ;
[0237] AUC Value: When dealing with imbalanced datasets, the performance of the random forest algorithm is poor and it tends to give higher weights to the majority class; therefore, it is necessary to introduce the Receiver Operating Characteristic (ROC) curve, a useful tool for evaluating classification performance; first, it is necessary to calculate the false positive rate (FPR) of misclassified negative samples and the true positive rate (TPR) of correctly classified positive samples. The values of both range from 0 to 1, and their calculation formulas are as follows:
[0238] ;
[0239] ;
[0240] Taking FPR as the abscissa and TPR as the ordinate, an ROC curve is established. The closer the curve is to the point (0, 1), that is, the closer TPR approaches 1 and FPR approaches 0, the more perfect the classifier is. On the contrary, if the curve is closer to the point (1, 0), the performance of the classifier is worse. The area value under the curve is the AUC value. Usually, its value range is between 0.5 and 1. The closer the AUC value is to 1, the better the performance of the classifier.
[0241] Prediction generation module 306: It is used to apply the machine learning model with training convergence to the unknown fracture target area for concealed fracture prediction processing, output the prediction result in the form of spatial coordinates, and generate a concealed fracture prediction distribution map covering the unknown fracture target area.
[0242] In the specific implementation process of the present invention, outputting the prediction result in the form of spatial coordinates and generating a concealed fracture prediction distribution map covering the unknown fracture target area includes: outputting the prediction result in the form of spatial coordinates, and performing visual analysis on the prediction result in the form of spatial coordinates in combination with geographic information system software, and generating a concealed fracture prediction distribution map covering the unknown fracture target area.
[0243] Specifically, applying the machine learning model with training convergence to the unknown research area for concealed fracture prediction processing of the unknown fracture target area; outputting the prediction result in the form of spatial coordinates, and performing visual analysis on the prediction result in combination with geographic information system (GIS) software to generate a fracture distribution map of the coverage area.
[0244] In the embodiment of the present invention, by fusing multi-source detection data such as gravity and magnetism, the underground geological structure can be reflected more comprehensively and deeply; the superposition and comprehensive analysis of multi-source data make the fracture characteristics more obvious, greatly improving the recognition ability of concealed fractures; and through intelligent algorithms and automated analysis technologies, the operation steps are simplified, the dependence on high-end equipment and professional personnel is reduced, and the overall cost is significantly reduced; at the same time, by optimizing the analysis process, a large amount of data can be processed quickly, improving the working efficiency of fracture prediction.
[0245] A computer-readable storage medium provided by an embodiment of the present invention has a computer program stored thereon, and when the program is executed by a processor, it implements the hidden fracture prediction method of any one of the above embodiments. Among them, the computer-readable storage medium includes but is not limited to any type of disk (including floppy disks, hard disks, optical disks, CD-ROMs, and magneto-optical disks), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards or optical cards. That is, the storage device includes any medium that can store or transmit information in a readable form by a device (such as a computer, mobile phone), and can be a read-only memory, a magnetic disk or an optical disk, etc.
[0246] An embodiment of the present invention also provides a computer application program that runs on a computer and is used to execute the hidden fracture prediction method of any one of the above embodiments.
[0247] In addition, Figure 4 It is a schematic diagram of the structural composition of an electronic device in an embodiment of the present invention.
[0248] An embodiment of the present invention also provides an electronic device, as Figure 4 shown. The electronic device includes devices such as a processor 402, a memory 403, an input unit 404, and a display unit 405. Those skilled in the art can understand that Figure 4 the structural devices of the electronic device shown do not constitute a limitation on all devices, and may include more or fewer components than shown, or combine certain components. The memory 403 can be used to store the application program 401 and each functional module, and the processor 402 runs the application program 401 stored in the memory 403, thereby executing various functional applications and data processing of the device. The memory can be an internal memory or an external memory, or include both an internal memory and an external memory. The internal memory can include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a flash memory, or a random access memory. The external memory can include a hard disk, a floppy disk, a ZIP disk, a USB flash drive, a magnetic tape, etc. The memory disclosed in the present invention includes but is not limited to these types of memories. The memory disclosed in the present invention is only an example and not a limitation.
[0249] The input unit 404 is used to receive the input of signals and the keywords input by the user. The input unit 404 may include a touch panel and other input devices. The touch panel can collect the touch operations of the user on or near it (such as the operations of the user using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch panel), and drive the corresponding connection device according to a pre-set program; the other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (such as play control keys, switch keys, etc.), a trackball, a mouse, a joystick, etc. The display unit 405 can be used to display the information input by the user or the information provided to the user and various menus of the terminal device. The display unit 405 can be in the form of a liquid crystal display, an organic light-emitting diode, etc. The processor 402 is the control center of the terminal device, connecting various parts of the entire device through various interfaces and lines, and performing various functions and processing data by running or executing the software programs and / or modules stored in the memory 403, and calling the data stored in the memory.
[0250] As an embodiment, the electronic device includes: one or more processors 402, a memory 403, and one or more application programs 401, wherein the one or more application programs 401 are stored in the memory 403 and are configured to be executed by the one or more processors 402, and the one or more application programs 401 are configured to execute the hidden fracture prediction method in any one of the above embodiments.
[0251] In the embodiment of the present invention, by fusing multi-source detection data such as gravity and magnetism, the underground geological structure can be more comprehensively and deeply reflected; the superposition and comprehensive analysis of multi-source data make the fracture characteristics more obvious, greatly improving the recognition ability of hidden fractures; and through intelligent algorithms and automated analysis techniques, the operation steps are simplified, the dependence on high-end equipment and professionals is reduced, and the overall cost is significantly reduced; at the same time, by optimizing the analysis process, a large amount of data can be quickly processed, improving the working efficiency of fracture prediction.
[0252] In addition, the above has introduced in detail a hidden fracture prediction method and related devices based on the fusion of multi-source detection data provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for predicting concealed faults based on multi-source detection data fusion, characterized in that The method includes: Performing detection processing within a specified area based on a gravity instrument device and a magnetic instrument device to obtain gravity detection data and magnetic detection data. The specified area is a known fracture area. The gravity detection data is Bouguer gravity anomaly data, and the magnetic detection data is aeromagnetic anomaly data; Performing fusion processing on the gravity detection data and the magnetic detection data to form multi-source fusion detection data; Performing enhancement processing on the multi-source fusion detection data based on a preset data enhancement algorithm to obtain enhanced multi-source fusion detection data; Performing dataset construction processing on the enhanced multi-source fusion detection data and performing annotation processing on the constructed dataset to form a training sample dataset; Inputting the training sample dataset into a preset machine learning model for training processing to obtain a trained and converged machine learning model. The machine learning model includes a random forest model and an artificial neural network model; Applying the trained and converged machine learning model to an unknown fracture target area for concealed fracture prediction processing, outputting the prediction result in the form of spatial coordinates, and generating a concealed fracture prediction distribution map covering the unknown fracture target area.
2. The method for predicting hidden faults according to claim 1, characterized in that, The performing fusion processing on the gravity detection data and the magnetic detection data to form multi-source fusion detection data includes: Performing abnormal data filtering processing on the gravity detection data and the magnetic detection data to obtain filtered gravity detection data and filtered magnetic detection data; Successively performing smoothing and compensation processing on the filtered gravity detection data and the filtered magnetic detection data to obtain compensated gravity detection data and compensated magnetic detection data; Performing fusion processing on the compensated gravity detection data and the compensated magnetic detection data based on a data fusion algorithm to obtain multi-source fusion detection data.
3. The method for predicting hidden faults according to claim 1, wherein The performing enhancement processing on the multi-source fusion detection data based on a preset data enhancement algorithm to obtain enhanced multi-source fusion detection data includes: Using multiple data enhancement algorithms to perform enhancement processing on the geological structure in the multi-source fusion detection data to obtain enhanced multi-source fusion detection data; Wherein the multiple data enhancement algorithms include the total horizontal derivative method, the vertical derivative method, the vertical second derivative, the tilt angle method, the tilt angle horizontal derivative method, the analytical signal method, the Theat method, the total horizontal derivative tilt angle method, the horizontal tilt angle method, the enhanced analytical signal method, and the enhanced tilt angle method.
4. The method for predicting buried faults according to claim 1, characterized in that, The performing dataset construction processing on the enhanced multi-source fusion detection data and performing annotation processing on the constructed dataset to form a training sample dataset includes: Performing sample synthesis processing on the enhanced multi-source fusion detection data based on the synthetic minority over-sampling technique (SMOTE) and performing dataset construction processing based on the synthesized sample data to obtain a constructed dataset; Obtaining the existing fracture distribution map and geological data of the specified area, and performing annotation processing on the constructed dataset based on the existing fracture distribution map and geological data to form a training sample dataset.
5. The method for predicting concealed faults according to claim 4, wherein The performing sample synthesis processing on the enhanced multi-source fusion detection data based on the synthetic minority over-sampling technique (SMOTE) includes: For each minority-class sample data X in the enhanced multi-source fusion detection data, new samples will be synthesized by calculating the difference between the minority-class sample X and its nearest neighbor sample data according to the required oversampling quantity, which is specifically expressed as: ; Among them, represents a random function; represents the synthesized sample data.
6. The method for predicting buried faults according to claim 1, wherein Inputting the training sample data set into a preset machine learning model for training to obtain a trained and converged machine learning model includes: When inputting the training sample data set into a preset machine learning model for training, adjust the hyperparameters in the preset machine learning model through a grid search algorithm or a genetic optimization algorithm, select the best parameter combination for training, and obtain a trained and converged machine learning model; When the preset machine learning model is a random forest model, the hyperparameters are the number of decision trees, the maximum depth of the decision trees, the minimum number of samples in the leaf nodes, and the minimum number of samples for node splitting in the random forest model; when the preset machine learning model is an artificial neural network model, the hyperparameters are the number of layers and the number of nodes of the artificial neural network model.
7. The method for predicting buried faults according to claim 1, wherein Outputting the prediction result in the form of spatial coordinates and generating a concealed fault prediction distribution map covering the unknown fault target area includes: Outputting the prediction result in the form of spatial coordinates, and performing visual analysis on the output prediction result in the form of spatial coordinates in combination with geographic information system software, and generating a concealed fault prediction distribution map covering the unknown fault target area.
8. An apparatus for predicting concealed faults based on multi-source detection data fusion, characterized in that, The device includes: A detection module: used to perform detection processing in a specified area based on a gravity instrument device and a magnetic instrument device to obtain gravity detection data and magnetic detection data. The specified area is a known fault area, the gravity detection data is Bouguer gravity anomaly data, and the magnetic detection data is aeromagnetic anomaly data; A data fusion module: used to fuse the gravity detection data and the magnetic detection data to form multi-source fusion detection data; A data enhancement module: used to enhance the multi-source fusion detection data based on a preset data enhancement algorithm to obtain enhanced multi-source fusion detection data; A data construction and annotation module: used to perform data set construction processing on the enhanced multi-source fusion detection data and perform annotation processing on the constructed data set to form a training sample data set; A model training module: used to input the training sample data set into a preset machine learning model for training to obtain a trained and converged machine learning model. The machine learning model includes a random forest model and an artificial neural network model; A prediction generation module: used to apply the trained and converged machine learning model to an unknown fault target area for concealed fault prediction processing, output the prediction result in the form of spatial coordinates, and generate a concealed fault prediction distribution map covering the unknown fault target area.
9. An electronic device, comprising a processor and a memory, characterized in that, The processor runs the computer program or code stored in the memory to implement the concealed fault prediction method according to any one of claims 1 to 7.
10. A computer-readable storage medium for storing a computer program or code, characterized in that, When the computer program or code is executed by the processor, the concealed fault prediction method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Surveying and mapping method for geological exploration
CN117274520A
Data fusion detection method and equipment for detecting geological structure of iron ore area
CN117991405A