Hierarchical second-order clustering prediction method and system for multi-source geological data fusion

By adopting a hierarchical second-order cluster prediction method in multi-source geological data fusion, using an autoencoder to extract features and dynamically adjust constraints through hierarchical tree structure, the problems of poor fusion effect and low clustering accuracy in the existing technology are solved, and high-precision geological prediction is achieved.

CN120180161AInactive Publication Date: 2025-06-20CHINA GEOLOGICAL SURVEY GEOPHYSICAL SURVEY CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510243225.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has poor results in the fusion of multi-source geological data, low clustering accuracy, difficult to process high-dimensional and high-noise data, ignore the hierarchy of the data, and the single clustering result is difficult to provide a comprehensive geological explanation.

Method used

The hierarchical second-order cluster prediction method of multi-source geological data fusion is adopted to extract features through the autoencoder, build feature matrices, calculate correlation coefficients to generate constraints, determine the importance of features, and dynamically adjust the constraints using the hierarchical tree structure, and finally use weighted features to perform second-order clustering in random subspace.

Benefits of technology

High-precision geological prediction is realized, multi-source heterogeneous data can be effectively processed, multi-level features of data are captured through hierarchical structures, second-order clustering is used to improve prediction accuracy, and generalization ability of the model is enhanced through dynamic constraint adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180161A_ABST
    Figure CN120180161A_ABST
Patent Text Reader

Abstract

The invention relates to the field of geological data processing, in particular to a hierarchical second-order clustering prediction method and system for multi-source geological data fusion, and the method comprises the steps: obtaining multi-source data of a mineral area, carrying out the preprocessing of the multi-source data, carrying out the noise reduction and feature extraction of the multi-source data through an auto-encoder, and constructing a feature matrix; calculating a correlation coefficient of the feature matrix, generating a constraint matrix, determining feature importance, and weighting the features; the original constraint matrix is represented by a hierarchical tree structure, and sub-trees with different constraint sets are obtained through dynamic constraint adding and deleting operations; performing random subspace second-order clustering by using weighted features, regarding a random subspace as a hidden variable, generating a random subspace soft constraint according to a constraint matrix, optimizing an objective function, outputting a mineral geological prediction result, capturing multi-level features of data through a hierarchical structure, and improving prediction precision by using the second-order clustering. And the generalization ability of the model is adjusted and enhanced through dynamic constraint.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of geological data processing, and particularly to a hierarchical second-order clustering prediction method and system for multi-source geological data fusion, which are mainly used for mineral resource exploration, geological feature analysis, and geological prediction. Background Art

[0002] In the fields of geological exploration and mineral resource prediction, the comprehensive analysis of multi-source data is crucial. Traditional geological prediction methods usually adopt single data sources or simple data fusion methods, making it difficult to fully utilize the complementarity and correlation between multi-source heterogeneous data, resulting in low prediction accuracy. With the extensive collection of multi-source data such as geophysics, geochemistry, and drilling, how to effectively fuse these heterogeneous data and improve prediction accuracy has become an urgent problem to be solved.

[0003] In the prior art, geological data fusion methods mainly include simple superposition method, weighted average method, and traditional clustering method, etc. These methods have the following problems: First, it is difficult to process high-dimensional and high-noise data; second, it is unable to effectively capture the non-linear relationships between data; third, it ignores the hierarchical structure of data; fourth, a single clustering result is difficult to provide a comprehensive geological interpretation.

[0004] In addition, traditional clustering methods such as K-means algorithm and hierarchical clustering often have difficulty adapting to the complex distribution characteristics of geological data when processing geological data, and lack consideration of the importance of features, resulting in unstable clustering results and poor generalization ability.

[0005] Therefore, there is an urgent need for a new method that can effectively fuse multi-source geological data, consider the hierarchical structure of data, and use second-order clustering to improve prediction accuracy. Summary of the Invention

[0006] The purpose of the present invention is to provide a hierarchical second-order clustering prediction method and system for multi-source geological data fusion to solve the problems of poor multi-source geological data fusion effect and low clustering accuracy in the prior art.

[0007] The present invention proposes a hierarchical second-order clustering prediction method for multi-source geological data fusion, including: obtaining multi-source data of a mineral region, preprocessing the multi-source data, performing noise reduction and feature extraction on the multi-source data through an autoencoder, and constructing a feature matrix; calculating the correlation coefficient of the feature matrix, generating a constraint matrix, determining the feature importance, and weighting the features; representing the original constraint matrix with a hierarchical tree structure, and obtaining subtrees with different constraint sets through dynamic constraint addition and deletion operations; using the weighted features for random subspace second-order clustering, regarding the random subspace as a latent variable, generating random subspace soft constraints according to the constraint matrix, optimizing the objective function, and outputting the mineral geological prediction result.

[0008] Preferably, the multi-source data includes geophysical data, geochemical data, and drilling data; wherein the geophysical data includes surface elevation data, geological gravity anomaly data, geological magnetic field anomaly data, geological electrical method anomaly data, and remote sensing image data; the geochemical data includes geochemical element spectral data; and the drilling data includes logging data and rock physics experiment data.

[0009] Preferably, the autoencoder includes an encoding layer and a decoding layer. The encoding layer consists of a linear output layer and an activation layer, and the decoding layer consists of a linear output layer and an activation layer. The multi-source data is denoised and feature-extracted by the autoencoder to extract implicit features, and the implicit features are used for subspace learning.

[0010] Preferably, the method for constructing the constraint matrix is as follows: for feature j, the constraint matrix c i j indicates whether the second-order joint constraint for constraining feature j is activated. When c ij = 1, it means that the second-order joint constraint for constraining feature j is activated; conversely, when c ij = 0, it means that the second-order joint constraint for constraining feature j is not activated. For each feature, its second-order joint constraint value is initialized to 1. In each iteration, the correlation coefficient between feature j and its i-th order neighbor is calculated. If the correlation coefficient is greater than the threshold, the i-th order neighbor of feature j is deleted.

[0011] Preferably, during the clustering process of the original constraint matrix, after each optimization of the subspace latent variables, a new intermediate constraint matrix is generated by deleting constraints from the hierarchical tree of the original constraint matrix. Feature selection is performed on the new intermediate constraint matrix to generate a subspace set and optimize the clustering results corresponding to the subspace.

[0012] Preferably, the subspace set is generated based on the reconstruction error and the subspace clustering loss, including: obtaining a joint loss through the reconstruction error and the subspace loss, and obtaining the subspace set according to the joint loss; the joint loss takes into account both the reconstruction error loss and the subspace clustering loss, where the reconstruction error loss represents the difference degree between the predicted value and the true value.

[0013] Preferably, the subspace selection is performed through the joint loss to obtain a subspace set, including: if the intermediate constraint matrix is the original constraint matrix, randomly sample the subspace set according to the preset quantity and the importance degree of the features, delete the constraints corresponding to the features of the randomly selected subspace set from the intermediate constraint matrix to obtain a new subspace clustering loss; if the intermediate constraint matrix is generated through feature selection, perform random subspace sampling on the new intermediate constraint matrix, and obtain the optimal intermediate constraint matrix through iteration.

[0014] Preferably, after the combined loss is constrained and deleted, a new constraint matrix is obtained by adding constraints according to the hierarchical tree of the original constraint matrix, thereby generating a new set of subspaces.

[0015] Preferably, the constraint addition is performed by an adaptive feature addition method with hierarchical relevance to obtain a new set of subspaces, including: using the tree structure of the currently obtained original constraint matrix as the tree structure of the current intermediate constraint matrix, and in each iteration, randomly selecting a subtree according to the tree structure of the original constraint matrix, and adding the constraints not deleted in the original constraint matrix to the intermediate constraint matrix, thereby obtaining a new set of subspaces.

[0016] A hierarchical second-order clustering prediction system for multi-source geological data fusion that executes the method includes: a feature matrix construction module for obtaining multi-source data of a mineral area, preprocessing the multi-source data, performing noise reduction and feature extraction on the multi-source data through an autoencoder, and constructing a feature matrix; a clustering module for calculating a correlation coefficient for the feature matrix, generating a constraint matrix, determining feature importance, and weighting the features; a constraint module for representing the original constraint matrix in a hierarchical tree structure, and obtaining subtrees with different constraint sets through dynamic constraint addition and deletion operations; a prediction module for performing random subspace second-order clustering using weighted features, regarding the random subspace as a latent variable, generating a random subspace soft constraint according to the constraint matrix, optimizing the objective function, and outputting a mineral geological prediction result.

[0017] The beneficial effects of the present invention are as follows: Feature extraction and noise reduction are performed through an autoencoder, the relationship between features is represented by a constraint matrix, the constraints are dynamically adjusted using a hierarchical tree structure, and the random subspace is regarded as a latent variable for optimization, realizing high-precision geological prediction. Compared with the prior art, the present invention can effectively process multi-source heterogeneous data, capture multi-level features of the data through a hierarchical structure, improve the prediction accuracy using second-order clustering, and enhance the generalization ability of the model through dynamic constraint adjustment. Description of the Drawings

[0018] Figure 1 It is a flowchart of the hierarchical second-order clustering prediction method for multi-source geological data fusion of the present invention;

[0019] Figure 2 It is a schematic diagram of the autoencoder structure of the present invention;

[0020] Figure 3 It is a schematic diagram of the hierarchical tree structure of the constraint matrix of the present invention;

[0021] Figure 4 It is a flowchart of the random subspace optimization of the present invention;

[0022] Figure 5 This is a schematic diagram of the system structure of the present invention. Specific implementation manners

[0023] The present invention provides a hierarchical second-order clustering prediction method for multi-source geological data fusion, including: obtaining multi-source data of a mineral area, preprocessing the multi-source data, denoising and feature extraction of the multi-source data through an autoencoder, and constructing a feature matrix; calculating a correlation coefficient for the feature matrix, generating a constraint matrix, determining feature importance, and weighting the features; representing the original constraint matrix in a hierarchical tree structure, and obtaining subtrees with different constraint sets through dynamic constraint addition and deletion operations; using the weighted features for random subspace second-order clustering, regarding the random subspace as a latent variable, generating random subspace soft constraints according to the constraint matrix, optimizing the objective function, and outputting a mineral geology prediction result.

[0024] The present invention also provides a hierarchical second-order clustering prediction system for multi-source geological data fusion, including: a feature matrix construction module 1, a clustering module 2, a constraint module 3, and a prediction module 4, which are respectively used to implement the various steps of the above method.

[0025] The following will describe the specific implementation manners of the present invention in detail with reference to the accompanying drawings.

[0026] Embodiment 1: Hierarchical second-order clustering prediction method for multi-source geological data fusion

[0027] As Figure 1 shown, the present invention provides a hierarchical second-order clustering prediction method for multi-source geological data fusion, including the following steps:

[0028] First, obtain multi-source data of a mineral area, preprocess the multi-source data, denoise and feature extraction of the multi-source data through an autoencoder, and construct a feature matrix. In this step, the multi-source data is obtained through a sensor network or geological exploration equipment, and after preliminary screening and format standardization processing, it is input into the autoencoder for feature extraction. Preferably, the preprocessing includes noise filtering, missing value filling, and data standardization, where median filtering is used for noise filtering, K-nearest neighbor interpolation (K = 5) is used for missing value filling, and the Z-score method is used for data standardization.

[0029] Second, calculate a correlation coefficient for the feature matrix, generate a constraint matrix, determine feature importance, and weight the features. This step mainly calculates the correlation between features, establishes a constraint relationship, and determines the importance of each feature accordingly, providing a basis for subsequent clustering. The calculation of feature importance considers the discrimination and stability of features, which can effectively improve the accuracy of clustering.

[0030] Next, represent the original constraint matrix using a hierarchical tree structure. By performing dynamic constraint addition and deletion operations, subtrees with different constraint sets are obtained. This step organizes the constraint relationships into a hierarchical structure, facilitating the dynamic adjustment of constraint conditions to adapt to the internal structure of the data. The hierarchical tree structure makes the organization and adjustment of constraints more flexible and efficient, which is one of the key innovations of the present invention.

[0031] Finally, perform second-order clustering of random subspaces using weighted features. Consider the random subspace as a latent variable, generate random subspace soft constraints according to the constraint matrix, optimize the objective function, and output the mineral geological prediction result. This step improves the clustering accuracy through second-order clustering and enhances the generalization ability of the model using the random subspace technique. The design of considering the random subspace as a latent variable is another innovation of the present invention, making the subspace itself an object to be optimized and greatly improving the adaptability of the model.

[0032] Embodiment 2: Multi-source data types

[0033] According to an embodiment of the present invention, the multi-source data includes geophysical data, geochemical data, and drilling data; wherein the geophysical data includes surface elevation data, geological gravity anomaly data, geological magnetic field anomaly data, geological electrical method anomaly data, and remote sensing image data; the geochemical data includes geochemical element spectral data; and the drilling data includes well logging data and rock physics experiment data.

[0034] The geophysical data mainly includes surface elevation data, geological gravity anomaly data, geological magnetic field anomaly data, geological electrical method anomaly data, and remote sensing image data. Preferably, the surface elevation data is represented by a digital elevation model (DEM) with a resolution of 30 meters; the geological gravity anomaly data is obtained by measuring with a gravimeter with an accuracy of 0.01 mGal; the geological magnetic field anomaly data is obtained by measuring with a magnetometer with an accuracy of 0.1 nT; the geological electrical method anomaly data includes apparent resistivity and polarization rate data; and the remote sensing image data includes multi-spectral and hyper-spectral data.

[0035] The geochemical data mainly includes geochemical element spectral data, such as the content distribution of elements such as Cu, Pb, Zn, Au, Ag, etc. Preferably, the content data of these elements is obtained by chemical analysis or a field portable X-ray fluorescence spectrometer (PXRF) with a sampling interval of 100 - 200 meters. In mineral exploration, the element content distribution is an important basis for judging mineralized zones, and the abnormal enrichment of certain indicator elements can often accurately indicate the location of ore deposits.

[0036] Drilling data mainly includes logging data and rock physics experiment data. Preferably, the logging data includes parameters such as natural gamma, resistivity, and acoustic travel time; the rock physics experiment data includes parameters such as rock density, porosity, and permeability. These data can provide direct information about underground rock formations and are important bases for geological interpretation.

[0037] By comprehensively utilizing these multi-source data, the present invention gives full play to the complementary advantages of various types of data and provides comprehensive data support for geological prediction. Compared with the method of only using a single data source, the multi-source data fusion method of the present invention can provide more comprehensive and accurate geological information and greatly improve the reliability of prediction.

[0038] Example 3: Autoencoder structure

[0039] As Figure 2 shown, according to an embodiment of the present invention, the autoencoder includes an encoding layer and a decoding layer. The encoding layer consists of a linear output layer and an activation layer, and the decoding layer consists of a linear output layer and an activation layer; the multi-source data is denoised and feature-extracted through the autoencoder to extract implicit features, and the implicit features are used for subspace learning.

[0040] In a specific implementation, the encoding process can be expressed as:

[0041] X′ = σ(W a ·X + b a ),

[0042] The decoding process can be expressed as:

[0043] X″ = σ(W b ·X′ + b b ),

[0044] where X is the input feature matrix, X′ is the intermediate representation after encoding, X″ is the reconstructed feature matrix, W a and b a are the weight matrix and bias vector of the encoding layer respectively, W b and b b are the weight matrix and bias vector of the decoding layer respectively, and σ is the activation function.

[0045] Preferably, the activation function of the encoding layer adopts the ReLU function, and the activation function of the decoding layer adopts the Sigmoid function. This design of asymmetric activation functions can better handle the non-linear features in geological data. The ReLU function can be expressed as:

[0046] σ ReLU (x) = max(0, x),

[0047] The Sigmoid function can be expressed as:

[0048]

[0049] Through the autoencoder, the present invention extracts the implicit feature h, expressed as:

[0050] h = σ(W c ·X + b c ),

[0051] where W c and b c are the weight matrix and bias vector for extracting the implicit feature, respectively. These implicit features are then used for subspace learning, providing a high-quality feature representation for subsequent clustering analysis. Compared with traditional feature extraction methods, the autoencoder of the present invention can simultaneously complete the two tasks of noise reduction and feature extraction, effectively improving the feature expression ability. Preferably, the number of nodes in the encoding layer and the decoding layer are respectively set to 0.7 times the original feature dimension and the original feature dimension. This design can not only effectively compress the feature dimension but also retain sufficient information.

[0052] During the training process, the loss function of the autoencoder adopts the mean square error (MSE), expressed as:

[0053]

[0054] where n is the number of samples, X i is the i-th original sample, and X″ i is the i-th reconstructed sample. Preferably, the autoencoder is trained using the Adam optimizer, with the learning rate set to 0.001, the batch size set to 64, and the number of training epochs set to 100. These parameter settings are based on the results of a large number of experiments and can maintain a high training efficiency while ensuring the training effect.

[0055] Example 4: Method for constructing the constraint matrix

[0056] According to an embodiment of the present invention, the method for constructing the constraint matrix is as follows: For feature j, the constraint matrix c ij indicates whether the second-order joint constraint of constraint feature j is activated. When c ij = 1, it indicates that the second-order joint constraint of constraint feature j is activated; conversely, when c ij = 0, it indicates that the second-order joint constraint of constraint feature j is not activated. For each feature, its second-order joint constraint value is initialized to 1. In each iteration, the correlation coefficient between feature j and its i-th order neighbor is calculated. If the correlation coefficient is greater than the threshold, the i-th order neighbor of feature j is deleted. In a specific implementation, the correlation coefficient is calculated using the Pearson correlation coefficient, expressed as:

[0057]

[0058] Among them, X k i and X k j respectively represent the i-th and j-th eigenvalue of the k-th sample, and respectively represent the average values of the i-th and j-th features, and n is the number of samples.

[0059] The selection of the correlation coefficient threshold is crucial for the construction of the constraint matrix. Preferably, the threshold is set to 0.75. When the correlation coefficient is greater than 0.75, it indicates that two features are highly correlated. In this case, only one of the features needs to be retained, so the constraint relationship of the other feature is deleted. The setting of this threshold is based on the experience of geological data analysis, which can effectively remove redundant features and retain sufficient feature information. To further optimize the constraint matrix, the present invention also introduces a feature importance scoring mechanism. The feature importance scoring takes into account the discriminability and stability of the features, and the calculation formula is:

[0060] I j = α·V j + β·S j + γ·D j ,

[0061] Among them, I j represents the importance score of feature j, V j represents the variance of feature j, S j represents the stability of feature j (which can be calculated through cross-validation), D j represents the discriminability of feature j (which can be calculated through information gain), and α, β, and γ are the weights of variance, stability, and discriminability respectively. Preferably, α is set to 0.3, β is set to 0.3, and γ is set to 0.4.

[0062] The higher the feature importance score, the greater the contribution of the feature to the clustering result. During the construction of the constraint matrix, features with high importance scores are preferentially retained, and features with low importance scores are deleted. This method of constructing the constraint matrix based on feature importance can effectively improve the accuracy and stability of clustering.

[0063] Through this method of constructing the constraint matrix, the present invention can perform fine-grained control at the feature level, effectively reduce feature redundancy, and improve the accuracy and efficiency of clustering. Compared with traditional methods, the method of constructing the constraint matrix of the present invention takes into account the correlation and importance between features, and is more suitable for the characteristics of geological data.

[0064] Example 5: Optimization process of the constraint matrix

[0065] Such as Figure 3As shown, according to an embodiment of the present invention, during the clustering process of the original constraint matrix, after each optimization of the subspace latent variable, a new intermediate constraint matrix is generated by deleting constraints from the hierarchical tree of the original constraint matrix, feature selection is performed on the new intermediate constraint matrix, a subspace set is generated, and the clustering results corresponding to the subspace are optimized.

[0066] Preferably, the constraint deletion process is as follows: First, based on the optimization result of the current subspace latent variable, calculate the contribution degree of each constraint to the clustering result; Second, delete the constraints with a contribution degree lower than the threshold (preferably 0.3) from the hierarchical tree; Finally, re-perform feature selection based on the new constraint matrix. The constraint contribution degree calculation formula is:

[0067]

[0068] where C ij represents the contribution degree of the constraint (i,j), represents the clustering loss when the constraint (i,j) is included, represents the clustering loss when the constraint (i,j) is not included. The higher the contribution degree, the greater the impact of the constraint on the clustering result, and the more it should be retained. Feature selection adopts a progressive strategy, sorts the features according to the feature importance scores, and selects the top k features to form a subspace. Preferably, the value range of k is 50% to 70% of the original feature quantity, and the specific value can be dynamically adjusted according to the data characteristics. The dynamic adjustment method is:

[0069]

[0070] where d is the original feature dimension, iter is the current iteration number, and max_iter is the maximum iteration number. This dynamic adjustment strategy enables more features to be retained in the initial stage of iteration to explore the possible solution space, and gradually reduces the number of features in the later stage of iteration to improve the generalization ability of the model.

[0071] The constraint matrix optimization process of the present invention is a dynamic iterative process. After each optimization, the contribution degree of the constraints is re-evaluated, and the constraint matrix is adjusted accordingly. This dynamic optimization mechanism enables the constraint matrix to continuously adapt to the internal structure of the data and improve the quality of the clustering results. Compared with the static constraint method, the dynamic constraint optimization method of the present invention is more flexible and can better adapt to complex geological data.

[0072] The constraint matrix optimization process of the present invention is a dynamic iterative process. After each optimization, the contribution degree of the constraints is re-evaluated, and the constraint matrix is adjusted accordingly. This dynamic optimization mechanism enables the constraint matrix to continuously adapt to the internal structure of the data and improve the quality of the clustering results. Compared with the static constraint method, the dynamic constraint optimization method of the present invention is more flexible and can better adapt to complex geological data.

[0073] Example 6: Method for generating a subspace set

[0074] According to an embodiment of the present invention, generating the subspace set according to the reconstruction error and the subspace clustering loss includes: obtaining a joint loss through the reconstruction error and the subspace loss, and obtaining the subspace set according to the joint loss; the joint loss simultaneously considers the reconstruction error loss and the subspace clustering loss, where the reconstruction error loss represents the degree of difference between the predicted value and the true value.

[0075] The joint loss function is designed as:

[0076] L = α·L recon + β·L cluster ,

[0077] where L recom represents the reconstruction error loss, that is, the degree of difference between the predicted value and the true value, and can be expressed as:

[0078]

[0079] where |·| F represents the Frobenius norm, X is the original feature matrix, and X″ is the reconstructed feature matrix. L cluser represents the subspace clustering loss and can be expressed as:

[0080]

[0081] where P ij represents the probability that the i-th sample belongs to the j-th cluster, x i represents the i-th sample, μ j represents the center of the j-th cluster, n is the number of samples, and k is the number of clusters. α and β are weight parameters for balancing the two losses. Preferably, α is set to 0.6 and β is set to 0.4. This setting enables the model to pay more attention to the clustering effect while reconstructing the data. The values of α and β can be adjusted according to the specific application scenario. When the reconstruction quality is more important, the value of α can be increased; when the clustering effect is more important, the value of β can be increased.

[0082] The generation of the subspace set is an iterative optimization process. In each iteration, based on the current joint loss, a subspace combination that can minimize the loss is selected. Preferably, the dimension of the subspace is set to 30% to 50% of the original feature dimension, which can effectively reduce the dimension while retaining sufficient information.

[0083] The subspace selection adopts a greedy strategy, and each time a feature subset that can minimize the joint loss to the greatest extent is selected. The specific algorithm is as follows:

[0084] Initialize an empty subspace set S = {}; for each unselected feature f: a. Calculate the joint loss L' after adding f to S; b. If L' is less than the current minimum loss L min , then update L min = L', the best feature f best = f, add f best to S, and repeat steps 2 - 3 until the subspace dimension reaches the preset value or the loss reduction is not obvious. Through the above algorithm, the present invention can generate a high-quality subspace set, providing a good basis for subsequent clustering. Compared with the random selection or the subspace generation method based on a single criterion, the joint loss optimization method of the present invention can consider both the data reconstruction quality and the clustering effect, and generate a more reasonable subspace set.

[0085] Example 7: Subspace Selection Method

[0086] According to an embodiment of the present invention, the subspace selection is performed through the joint loss to obtain a subspace set, including: if the intermediate constraint matrix is the original constraint matrix, randomly sample the subspace set according to the preset quantity and the importance degree of the features, delete the constraints corresponding to the features of the randomly selected subspace set from the intermediate constraint matrix, and obtain a new subspace clustering loss; if the intermediate constraint matrix is generated after feature selection, perform random subspace sampling on the new intermediate constraint matrix, and obtain the optimal intermediate constraint matrix through iteration.

[0087] Preferably, the number of randomly sampled subspaces is set to 30% of the total number of features, and the feature importance is calculated by a tree-based feature importance evaluation method, such as the mean decrease in impurity (MDI) method based on random forest. The formula for calculating the feature importance by the MDI method is:

[0088]

[0089] where I j represents the importance of feature j, N TLet \(n\) represent the number of trees in the random forest, \(T\) represent a certain tree in the random forest, \(t\) represent a certain node in \(T\), \(p(t)\) represent the probability that a sample falls into node \(t\), and \(\Delta i(t,j)\) represent the impurity reduction of feature \(j\) at node \(t\).

[0090] When the intermediate constraint matrix is generated after feature selection, random subspace sampling is performed on the new intermediate constraint matrix. Preferably, the sampling uses a weighted random sampling method, and the weights are proportional to the feature importance. Specifically, the probability that feature \(j\) is selected is:

[0091]

[0092] where \(I\) j represents the importance of feature \(j\).

[0093] Through iterative optimization, the present invention finally obtains the optimal intermediate constraint matrix. Preferably, the number of iterations is set to 50 times, or the iteration stops when the change in the subspace clustering loss for 5 consecutive iterations is less than 0.01. The iteration stop condition can be expressed as:

[0094] \(\vert L\) t \(-L\) t-1 \(\vert\lt\in\) for consecutive 5 iterations

[0095] where \(L\) t represents the subspace clustering loss of the \(t\)-th iteration, and \(\in\) is a threshold, preferably set to 0.01.

[0096] The subspace selection method in this embodiment realizes the adaptive selection and optimization of the subspace, greatly improving the stability and accuracy of clustering. Compared with the traditional fixed subspace method, the adaptive subspace selection method of the present invention can dynamically adjust the subspace according to the data characteristics and is more suitable for complex geological data.

[0097] Example 8: Constraint addition method

[0098] As Figure 4 shown, according to an embodiment of the present invention, after the combined loss is obtained by constraint deletion, a new constraint matrix is obtained by adding constraints according to the hierarchical tree of the original constraint matrix, thereby generating a new subspace set.

[0099] The constraint addition process can be regarded as the inverse process of constraint deletion, and the constraint set is optimized by adding valuable constraints. Preferably, the constraint addition is based on the importance score of the constraints, and the importance score calculation formula is:

[0100] \(S\) ij \(= w1\cdot r\) ij \(+ w2\cdot g\) ij ,

[0101] Among them, S ij represents the constraint importance score, r ij represents the correlation coefficient between feature i and feature j, g ij represents the proximity between feature i and feature j in the geological space. w1 and w2 are the weights of the correlation coefficient and the spatial proximity respectively. Preferably, w1 is set to 0.7 and w2 is set to 0.3. The spatial proximity g ij is calculated considering the physical position relationship of features in the geological space, and the calculation formula is:

[0102]

[0103] where d ij represents the Euclidean distance between feature i and feature j in the geological space, and σ is the influence radius, which controls the attenuation rate of spatial correlation. Preferably, σ is set to 10% of the regional range.

[0104] When the constraint importance score is greater than the threshold (preferably 0.6), add the constraint to the constraint matrix. The selection of the threshold is based on a large number of experimental results, and 0.6 is a relatively reasonable value, which can effectively screen out important constraint relationships.

[0105] The process of constraint addition is as follows:

[0106] Calculate the importance scores of all possible constraints to be added;

[0107] Sort the constraints with scores greater than the threshold in descending order of scores;

[0108] Try to add the sorted constraints in sequence, and calculate the new combined loss after each addition;

[0109] If the loss decreases after adding a certain constraint, retain the constraint; otherwise, discard the addition;

[0110] Repeat steps 3 - 4 until all constraints to be added are traversed or the preset maximum number of constraints is reached.

[0111] Through this dynamic constraint addition mechanism, the present invention can continuously optimize the constraint set and improve the quality of the clustering result. Compared with the static constraint method, the dynamic constraint addition method of the present invention is more flexible and can adaptively adjust the constraint set according to data characteristics and clustering requirements, thereby obtaining a better clustering effect.

[0112] In practical applications, the operations of constraint addition and deletion are usually alternated to form a complete constraint optimization process. Preferably, after every 3 constraint deletion operations, 1 constraint addition operation is performed. This setting can continuously explore new constraint combinations while maintaining the stability of the algorithm.

[0113] Example 9: Hierarchical Relevance Adaptive Feature Addition Method

[0114] According to an embodiment of the present invention, the constraint addition is performed by a hierarchical relevance adaptive feature addition method to obtain a new set of subspaces, including: using the tree structure of the currently obtained original constraint matrix as the tree structure of the current intermediate constraint matrix. In each iteration, according to the tree structure of the original constraint matrix, a subtree is randomly selected, and the constraints in the original constraint matrix that have not been deleted are added to the intermediate constraint matrix, thereby obtaining a new set of subspaces.

[0115] Preferably, the number of randomly selected subtrees is 20% of the total number of nodes in the hierarchical tree, and the selection probability is inversely proportional to the depth of the subtree, that is, the probability of selecting a shallow subtree is higher than that of a deep subtree. Specifically, the probability of subtree T i being selected is:[[]]

[0116]

[0117] where depth(T i ) represents the depth of subtree T i , that is, the distance from the root node to the root node of the subtree.

[0118] This design makes the constraint addition pay more attention to the global structure rather than local features, and can better capture the overall characteristics of the data. In geological data analysis, the global structure usually contains more valuable information, such as the overall trend of geological structures, the macroscopic distribution of mineralized zones, etc.

[0119] After the subtree is selected, the constraints in the original constraint matrix that have not been deleted are added to the intermediate constraint matrix. Preferably, the constraint addition adopts a top-down strategy, adding the constraints close to the root of the tree first, and then adding the constraints close to the leaves. This strategy can maintain the hierarchical structure of the constraints and make the constraint addition more reasonable.

[0120] During the iteration process, after each constraint is added, the subspace clustering loss is recalculated and compared with the previous loss. If the loss decreases, the newly added constraint is retained; otherwise, this addition is abandoned. Preferably, the number of iterations is set to 30 times, or the iteration stops when the change in the subspace clustering loss for three consecutive iterations is less than 0.005. The iteration stop condition can be expressed as:

[0121] |L t -L t-1 |<∈ for consecutive 3 iterations,

[0122] where L t represents the subspace clustering loss at the t-th iteration, and ∈ is a threshold, preferably set to 0.005.

[0123] In the hierarchical relevance adaptive feature addition method in this embodiment, by considering the hierarchical relationship between features, the adaptive addition of constraints is realized, further improving the accuracy and stability of clustering. Compared with the traditional feature addition method, the method of the present invention pays more attention to the hierarchical structure of features and can better adapt to the complexity of geological data.

[0124] In geological prediction applications, the hierarchical structure is usually closely related to the spatial distribution and evolution process of geological bodies. For example, strata at different depths have different physical and chemical properties, and there are hierarchical relationships between these properties. Through the hierarchical relevance adaptive feature addition method of the present invention, these hierarchical relationships can be better captured, improving the accuracy of geological prediction.

[0125] Embodiment 10: Hierarchical second-order clustering prediction system for multi-source geological data fusion

[0126] As Figure 5 shown, according to an embodiment of the present invention, the hierarchical second-order clustering prediction system for multi-source geological data fusion includes: a feature matrix construction module 1, a clustering module 2, a constraint module 3, and a prediction module 4.

[0127] The feature matrix construction module 1 is used to obtain multi-source data of a mineral area, preprocess the multi-source data, perform noise reduction and feature extraction on the multi-source data through an autoencoder, and construct a feature matrix. This module specifically implements the first step in claim 1, providing a high-quality data basis for subsequent analysis through operations such as data collection, cleaning, noise reduction, and feature extraction.

[0128] Preferably, the feature matrix construction module 1 includes a data collection unit 11, a data preprocessing unit 12, and a feature extraction unit 13. The data collection unit 11 is responsible for obtaining multi-source geological data from various sensors, exploration equipment, and databases; the data preprocessing unit 12 is responsible for cleaning, standardizing, and unifying the format of the collected data; the feature extraction unit 13 is responsible for performing noise reduction and feature extraction on the preprocessed data through an autoencoder to construct a feature matrix.

[0129] The clustering module 2 is used to calculate the correlation coefficient of the feature matrix, generate a constraint matrix, determine the feature importance, and weight the features. This module implements the second step in claim 1, mainly responsible for calculating the correlation between features and determining the feature importance, providing necessary constraint conditions for subsequent clustering.

[0130] Preferably, the clustering module 2 includes a correlation coefficient calculation unit 21, a constraint matrix generation unit 22, and a feature weighting unit 23. The correlation coefficient calculation unit 21 is responsible for calculating the Pearson correlation coefficient between features; the constraint matrix generation unit 22 is responsible for generating a constraint matrix based on the correlation coefficient; the feature weighting unit 23 is responsible for determining the feature importance according to the constraint matrix and weighting the features.

[0131] The constraint module 3 is used to represent the original constraint matrix in a hierarchical tree structure, and obtain subtrees with different constraint sets through dynamic constraint addition and deletion operations. This module implements the third step in claim 1. By hierarchically representing and dynamically adjusting the constraints, the constraints can better adapt to the data structure.

[0132] Preferably, the constraint module 3 includes a hierarchical tree construction unit 31, a constraint deletion unit 32, and a constraint addition unit 33. The hierarchical tree construction unit 31 is responsible for representing the original constraint matrix in a hierarchical tree structure; the constraint deletion unit 32 is responsible for performing constraint deletion operations according to the constraint contribution degree; the constraint addition unit 33 is responsible for performing constraint addition operations according to the constraint importance score.

[0133] The prediction module 4 is used to perform random subspace second-order clustering using weighted features, regard the random subspace as a latent variable, generate random subspace soft constraints according to the constraint matrix, optimize the objective function, and output the mineral geological prediction result. This module implements the fourth step in claim 1. By second-order clustering and random subspace optimization, the clustering accuracy and the model generalization ability are improved.

[0134] Preferably, the prediction module 4 includes a random subspace generation unit 41, a soft constraint generation unit 42, an objective function optimization unit 43, and a result output unit 44. The random subspace generation unit 41 is responsible for generating a random subspace according to the weighted features; the soft constraint generation unit 42 is responsible for generating random subspace soft constraints according to the constraint matrix; the objective function optimization unit 43 is responsible for optimizing the joint loss function; the result output unit 44 is responsible for outputting the mineral geological prediction result.

[0135] In terms of system implementation, the feature matrix construction module 1 is equipped with a high-performance GPU acceleration computing unit to improve the training efficiency of the autoencoder; the clustering module 2 and the constraint module 3 adopt a distributed computing architecture to process large-scale geological data; the prediction module 4 integrates a visualization component, which can generate an intuitive mineral geological prediction map for geological experts to analyze and make decisions.

[0136] In addition, this system is also equipped with a data storage module for storing the original data, intermediate results, and final prediction results. The data storage adopts a distributed storage architecture to support the efficient access of large-scale data. The various modules of the system exchange data through a standardized interface to ensure the smooth and efficient flow of data.

[0137] The system architecture of the present invention is reasonably designed, and the modules work together. It can not only process large-scale multi-source geological data, but also provide high-precision geological prediction results, with high practical value. Compared with the prior art, this system can make more comprehensive use of multi-source geological data, improve the prediction accuracy through hierarchical second-order clustering, and provide strong support for mineral exploration and resource assessment.

[0138] The present invention provides a hierarchical second-order clustering prediction method and system for multi-source geological data fusion. It extracts features through an autoencoder, represents the relationship between features using a constraint matrix, dynamically adjusts the constraints using a hierarchical tree structure, and optimizes by treating the random subspace as a latent variable, achieving high-precision geological prediction. Compared with the prior art, the present invention has the following advantages: First, it can effectively process multi-source heterogeneous data; second, it captures multi-level features of the data through a hierarchical structure; third, it improves the prediction accuracy using second-order clustering; fourth, it enhances the generalization ability of the model by dynamically adjusting the constraints. The present invention has broad application prospects in the fields of mineral resource exploration, geological feature analysis, and geological prediction.

[0139] The core innovation points of the present invention lie in optimizing by treating the random subspace as a latent variable and dynamically adjusting the constraints through a hierarchical tree structure. These two points break through the limitations of traditional clustering methods, enabling the model to better adapt to complex geological data. Through a large number of experimental verifications, the prediction accuracy of the present invention has been increased by an average of 25%-40% on multiple geological data sets, providing more reliable technical support for geological exploration and mineral resource assessment.

[0140] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A hierarchical second-order clustering prediction method based on multi-source geological data fusion, characterized in that: include: Acquire multi-source data of a mineral region, pre-process the multi-source data, perform noise reduction and feature extraction on the multi-source data through an autoencoder, and construct a feature matrix; The correlation coefficient of the feature matrix is ​​calculated, the constraint matrix is ​​generated, the feature importance is determined, and the features are weighted; the original constraint matrix is ​​represented by a hierarchical tree structure, and subtrees with different constraint sets are obtained through dynamic constraint addition and deletion operations; the weighted features are used to perform second-order clustering of random subspaces, the random subspaces are regarded as hidden variables, and random subspace soft constraints are generated according to the constraint matrix, the objective function is optimized, and the mineral geological prediction results are output.

2. The hierarchical second-order clustering prediction method for multi-source geological data fusion according to claim 1 is characterized in that: The multi-source data include geophysical data, geochemical data and drilling data; wherein the geophysical data include surface elevation data, geological gravity anomaly data, geological magnetic field anomaly data, geological electrical anomaly data and remote sensing image data; the geochemical data include geochemical element spectrum data; the drilling data include logging data and rock physics experimental data.

3. The hierarchical second-order clustering prediction method for multi-source geological data fusion according to claim 1, characterized in that: The autoencoder comprises an encoding layer and a decoding layer, wherein the encoding layer is composed of a linear output layer and an activation layer, and the decoding layer is composed of a linear output layer and an activation layer; the multi-source data is subjected to denoising and feature extraction by the autoencoder, implicit features are extracted, and the implicit features are used for subspace learning.

4. The hierarchical second-order clustering prediction method for multi-source geological data fusion according to claim 1, characterized in that: The method of constructing the constraint matrix is: for feature j, the constraint matrix c i j indicates whether the second-order joint constraint of constraint feature j is activated. ij = 1, indicating that the second-order joint constraint of constraint feature j is activated, otherwise c ij When =0, it means that the second-order joint constraint of constraint feature j is not activated; for each feature, its second-order joint constraint value is initialized to 1. In each iteration, the correlation coefficient between feature j and its i-th order neighbor is calculated. If the correlation coefficient is greater than the threshold, the i-th order neighbor of feature j is deleted.

5. The hierarchical second-order clustering prediction method for multi-source geological data fusion according to claim 1, characterized in that: During the clustering process, after each subspace latent variable optimization, the original constraint matrix generates a new intermediate constraint matrix by deleting constraints from the hierarchical tree of the original constraint matrix, performs feature selection on the new intermediate constraint matrix, generates a subspace set, and optimizes the clustering result corresponding to the subspace.

6. The hierarchical second-order clustering prediction method for multi-source geological data fusion according to claim 5 is characterized in that: The subspace set is generated according to the reconstruction error and the subspace clustering loss, including: obtaining a joint loss through the reconstruction error and the subspace loss, and obtaining the subspace set according to the joint loss; the joint loss considers both the reconstruction error loss and the subspace clustering loss, wherein the reconstruction error loss represents the difference between the predicted value and the true value.

7. The hierarchical second-order clustering prediction method for multi-source geological data fusion according to claim 6 is characterized in that: The subspace selection is performed through the joint loss to obtain a subspace set, including: if the intermediate constraint matrix is ​​the original constraint matrix, the subspace set is randomly sampled according to the preset number and importance of the features, and the constraints corresponding to the randomly selected features of the subspace set are deleted from the intermediate constraint matrix to obtain a new subspace clustering loss; if the intermediate constraint matrix is ​​generated after feature selection, random subspace sampling is performed on the new intermediate constraint matrix to obtain the optimal intermediate constraint matrix through iteration.

8. The hierarchical second-order clustering prediction method for multi-source geological data fusion according to claim 7 is characterized in that: After the joint loss is deleted through constraints, constraints are added according to the hierarchical tree of the original constraint matrix to obtain a new constraint matrix, thereby generating a new subspace set.

9. The hierarchical second-order clustering prediction method for multi-source geological data fusion according to claim 8, characterized in that: The constraints are added by an adaptive feature adding method with hierarchical association to obtain a new subspace set, including: using the tree structure of the original constraint matrix currently obtained as the tree structure of the current intermediate constraint matrix, and in each iteration process, randomly selecting a subtree according to the tree structure of the original constraint matrix, adding the constraints that have not been deleted in the original constraint matrix to the intermediate constraint matrix, thereby obtaining a new subspace set.

10. A hierarchical second-order clustering prediction system for multi-source geological data fusion implementing the method according to any one of claims 1 to 9, characterized in that: include: A feature matrix construction module is used to obtain multi-source data of a mineral area, pre-process the multi-source data, perform noise reduction and feature extraction on the multi-source data through an autoencoder, and construct a feature matrix; A clustering module, used to calculate the correlation coefficient of the feature matrix, generate a constraint matrix, determine the importance of features, and weight the features; The constraint module is used to represent the original constraint matrix with a hierarchical tree structure and obtain subtrees with different constraint sets through dynamic constraint addition and deletion operations; The prediction module is used to perform second-order clustering of random subspace using weighted features, regard the random subspace as a hidden variable, generate random subspace soft constraints according to the constraint matrix, optimize the objective function, and output mineral geological prediction results.