Multi-element prospecting data cleaning and feature extraction system based on machine learning
By generating a multi-scale structural tensor field and using its structural direction vector to clean and extract features from non-seismic data, the problem of insufficient utilization of geological structural information in existing technologies is solved, and the quality of multivariate mineral exploration data and the predictive ability of machine learning models are improved.
Patent Information
- Application Number
- CN202511131094.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-18
AI Technical Summary
Existing technologies fail to effectively utilize the rich geological structural information contained in seismic data when processing diverse mineral exploration data, resulting in inaccurate feature extraction and loss of information fusion, thus severing the inherent physical connections between different data.
By generating a multi-scale structural tensor field, the non-seismic data is cleaned using the structural direction vectors it provides, and coupled features, including single-scale and cross-scale coupled features, are generated. Feature extraction is then performed in conjunction with geological structural information.
It improves data quality and geological consistency, enhances the accuracy of machine learning models in identifying mineralization-favorable areas, and enables the discovery of more concealed and complex mineralization patterns.
Smart Images

Figure CN120972280A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of geophysical exploration data processing technology, specifically a multivariate mineral exploration data cleaning and feature extraction system based on machine learning. Background Technology
[0002] With the increasing difficulty of deep mineral exploration, utilizing machine learning techniques to integrate multivariate geophysical and geochemical data for mineralization prediction has become a research hotspot and an important development direction in this field. In traditional machine learning-based mineral exploration processes, researchers typically spatially align 3D seismic attribute data with various non-seismic data (such as gravity, magnetic, electrical, and geochemical data), and simply combine them as independent feature dimensions to construct feature vectors for model training. However, this approach faces profound inherent challenges in practice.
[0003] Non-seismic data, especially borehole geochemical data, are often sparsely and irregularly distributed in their raw form, and inevitably contain noise and missing values. Existing data cleaning and interpolation methods, such as conventional inverse distance weighted interpolation or Kriging interpolation, are mostly isotropic mathematical fillers, which fail to fully consider and utilize the tectonic orientation inherent in geological bodies during processing. This results in the spatial distribution of the generated data volume potentially contradicting the actual geological structural characteristics, thus introducing artificial noise that does not conform to geological laws and reducing the quality of the input data.
[0004] More importantly, existing technical solutions typically stop at simply piling up multi-source data at the feature construction level. They overlook a core geological fact: mineralization is often strongly controlled by geological structures, and the distribution, morphology, and intensity of geophysical and geochemical anomalies are closely related to the scale, orientation, and continuity of faults, folds, and other structures. Traditional methods fail to generate features that can directly and quantitatively describe this "coupling relationship," forcing machine learning models to learn these complex and implicit patterns from massive amounts of data. This not only increases the difficulty of model training but also limits the upper limit of their final predictive ability.
[0005] Therefore, how to deeply integrate geological structural information into the entire process of cleaning and feature extraction of multivariate data, so as to provide higher quality and more geologically relevant input for machine learning models, is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0006] The technical problem to be solved by this invention is that, in processing multi-modal mineral exploration data, existing technologies typically clean and preprocess data of different modalities independently, and only perform simple spatial feature stacking during subsequent fusion. This process severs the inherent physical connection between different data, and in particular, fails to effectively utilize the rich geological structural information contained in seismic data to guide the processing and feature extraction of other data, resulting in inaccurate feature extraction and information fusion loss.
[0007] To solve the above-mentioned technical problems, the present invention provides the following solution: The first aspect of this invention provides a machine learning-based multivariate mineral exploration data cleaning and feature extraction system, comprising: The data receiving unit is used to receive three-dimensional seismic data and non-seismic data within the exploration area. The structural field generation unit is used to generate a multi-scale structural tensor field covering the exploration area based on the three-dimensional seismic data, wherein the multi-scale structural tensor field provides structural direction vectors for each spatial location in the exploration area at multiple preset scales. An adaptive cleaning unit is used to clean the non-seismic data using the structural direction vector provided by the multi-scale structural tensor field, so as to obtain cleaned non-seismic data. The coupling feature generation unit is used to generate coupling features based on the cleaned non-seismic data and the multi-scale structural tensor field; The feature output unit is used to output a final feature vector containing the coupled features for use in a machine learning model.
[0008] In one specific embodiment, the structure field generation unit generates the multi-scale structure tensor field in the following manner: For the three-dimensional seismic data Apply a set of different preset scales Gaussian kernel Perform convolution to generate a set of multi-scale smooth data volumes. ,in , The total number of preset dimensions; Calculate each of the multi-scale smoothed data volumes Spatial gradient The multi-scale structure tensor field is generated by locally weighting the dyadic product of the spatial gradients, wherein the structure tensor at each scale is: ; in, For preset scale The structure tensor below, For multi-scale smooth data volume Spatial gradient, This represents the dyadic product operation. Indicated on the integral scale Local spatial weighted average; The multi-scale structural tensor field is subjected to eigenvalue decomposition at each spatial location and at each preset scale to obtain a set of orthogonal structural direction vectors. .
[0009] In one specific embodiment, the adaptive cleaning unit cleans the non-seismic data by at least one of the following methods: Based on the structural direction vector defined by the structure, the validity of data points in the non-seismic data is determined, and data points identified as noise are processed. Based on the structural direction vector, an anisotropic interpolation method is used to fill in the missing values in the non-seismic data, wherein the interpolation weight along the structural direction defined by the structural direction vector is higher than the interpolation weight across the structural direction.
[0010] Preferably, the coupling features generated by the coupling feature generation unit include single-scale coupling features. The coupling feature generation unit calculates the cleaned non-seismic data. Spatial gradient The spatial gradient is then projected onto the multi-scale structure tensor field at at least one preset scale. The structure direction vector below Above, to generate the single-scale coupling feature: ; in, For preset scale Single-scale coupling features; The spatial gradient of the cleaned non-seismic data; For preset scale The next structural direction vector; This represents the vector dot product operation.
[0011] Preferably, the coupling feature further includes cross-scale coupling features. The coupling feature generation unit generates the cross-scale coupling features based on the structural direction vectors of the multi-scale structural tensor field at at least two different preset scales.
[0012] In one specific embodiment, the at least two different preset scales include a small scale. and regional scale The cross-scale coupling feature is a structural consistency feature. The coupling feature generation unit calculates the structural direction vector at the small scale. With the structural direction vector at the aforementioned regional scale The absolute value of the dot product is used to determine the structural consistency feature. : ; in, This is a structural consistency feature; This refers to the structural direction vector at the small scale. This refers to the structural direction vector at the scale of the region. This represents the vector dot product operation; This indicates the operation of taking the absolute value.
[0013] In one specific embodiment, the cross-scale coupling feature is further defined as a hierarchical gradient projection feature. The coupling feature generation unit is based on small-scale... The single-scale coupling features generated below and the structural consistency features Perform combination operations to generate the hierarchical gradient projection features: ; in, This represents hierarchical gradient projection features; This refers to the single-scale coupling features generated at the aforementioned small scale. The structural consistency feature; It represents the multiplication operation.
[0014] Preferably, the non-seismic data includes at least one of geochemical data, gravity data, magnetic data, and electrical data.
[0015] Preferably, the final feature vector output by the feature output unit further includes the cleaned non-seismic data and the seismic attributes extracted from the three-dimensional seismic data.
[0016] The second aspect of this invention provides a method for cleaning and extracting features from multivariate mineral exploration data based on machine learning, comprising the following steps: Receives 3D seismic and non-seismic data within the exploration area; Based on the three-dimensional seismic data, a multi-scale structural tensor field covering the exploration area is generated, wherein the multi-scale structural tensor field provides structural direction vectors for each spatial location in the exploration area at multiple preset scales; The non-seismic data is cleaned using the structural direction vector provided by the multi-scale structural tensor field to obtain cleaned non-seismic data. Based on the cleaned non-seismic data and the multi-scale structural tensor field, coupling features are generated; The output is a final feature vector containing the coupled features, which is then used in a machine learning model.
[0017] This invention provides a multivariate mineral exploration data cleaning and feature extraction system based on machine learning. It has the following beneficial effects: 1. This invention solves the problems of outlier misjudgment and interpolation results inconsistent with geological laws caused by independent data processing in existing technologies by constructing a multi-scale structural tensor field and utilizing the structural direction vector it provides to guide adaptive cleaning of non-seismic data. This scheme uses the structural information inherent in seismic data as prior knowledge, introducing geological structural constraints during the data preprocessing stage. This enables the data cleaning process to distinguish between structure-related geological anomalies and structure-independent noise, thereby improving the data quality and geological consistency input to the machine learning model.
[0018] 2. This invention, through a coupled feature generation unit, projects the spatial gradient of cleaned non-seismic data onto the structural direction vector determined by seismic data, generating single-scale and cross-scale coupled features. This scheme creates a new feature dimension that can directly characterize the coupling relationship between physical property changes and geological structures, overcoming the shortcomings of existing technologies that can only simply stack various original features and cannot express their inherent physical correlations. These newly generated coupled features provide machine learning models with direct quantitative indicators of mineralization processes (such as material migration and enrichment along structures), enhancing the accuracy of the model in identifying favorable mineralization areas.
[0019] 3. This invention, by generating cross-scale coupling features such as structural consistency features and hierarchical gradient projection features, enables the technical solution to quantitatively analyze the synergistic relationship between the regional macroscopic tectonic background and local microstructures. This solution solves the technical problem of existing technologies' difficulty in capturing and utilizing the multi-scale tectonic composite control of mineralization patterns. It allows machine learning models not only to identify isolated local anomalies but also to identify local tectonic anomalies with higher mineralization potential that develop under favorable regional tectonic backgrounds, thereby enabling the discovery of more concealed and complex mineralization patterns. Attached Figure Description
[0020] Figure 1 This is a structural block diagram of a machine learning-based multivariate mineral exploration data cleaning and feature extraction system according to an embodiment of the present invention. Figure 2 This is a flowchart of a machine learning-based multivariate mineral exploration data cleaning and feature extraction method according to an embodiment of the present invention.
[0021] Among them, 10 is the data receiving unit; 20 is the structure field generation unit; 30 is the adaptive cleaning unit; 40 is the coupled feature generation unit; and 50 is the feature output unit. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0023] See attached document Figure 1 , Figure 1 This is a structural block diagram of a machine learning-based multivariate mineral exploration data cleaning and feature extraction system according to an embodiment of the present invention. The system provided by the present invention can be implemented on a computing device. This computing device can be a server, workstation, or any other suitable computing device. The computing device includes at least one processor, a memory connected to the processor, and a bus for connecting different device components. The memory stores computer program instructions, which, when executed by the processor, are used to implement the system and method of the present invention.
[0024] See attached document Figure 1 The system provided by this invention can be functionally divided into multiple logical units. In this embodiment, the system includes: a data receiving unit 10, a structure field generation unit 20, an adaptive cleaning unit 30, a coupled feature generation unit 40, and a feature output unit 50.
[0025] The data receiving unit 10 is used to receive three-dimensional seismic data and non-seismic data within the exploration area. The three-dimensional seismic data is a three-dimensional volume data. Non-seismic data includes, but is not limited to, geochemical data, gravity data, magnetic data, or electrical data acquired in point, line, or area formats. The data receiving unit 10 transmits the received data to a memory for use by other units.
[0026] The structural field generation unit 20 generates a multi-scale structural tensor field covering the exploration area based on the three-dimensional seismic data provided by the data receiving unit 10. The multi-scale structural tensor field is a four-dimensional data structure that provides a set of structural tensors at multiple preset scales at each location within the three-dimensional spatial grid of the exploration area. The structural field generation unit 20 further performs feature decomposition on each structural tensor, thereby providing a set of orthogonal structural direction vectors for each spatial location at each preset scale. The output of this unit, namely the multi-scale structural tensor field and its decomposed structural direction vectors, is transmitted to the adaptive cleaning unit 30 and the coupled feature generation unit 40.
[0027] The adaptive cleaning unit 30 receives non-seismic data provided by the data receiving unit 10 and a multi-scale structural tensor field generated by the structural field generation unit 20. This unit uses the structural orientation vectors provided by the multi-scale structural tensor field to clean the non-seismic data, obtaining cleaned non-seismic data. The cleaning operation includes identifying and processing outlier data points in the non-seismic data based on the structural orientation vectors, and performing anisotropic interpolation on missing values in the non-seismic data based on the structural orientation vectors. This unit then transmits the processed cleaned non-seismic data to the coupled feature generation unit 40 and feature output unit 50.
[0028] The coupling feature generation unit 40 receives the cleaned non-seismic data output by the adaptive cleaning unit 30 and the multi-scale structural tensor field generated by the structural field generation unit 20. Based on these two input data, this unit generates coupling features through a series of preset mathematical operations. The coupling features include single-scale coupling features and cross-scale coupling features. The cross-scale coupling features may further include structural consistency features and hierarchical gradient projection features. This unit transmits all generated coupling features to the feature output unit 50.
[0029] The feature output unit 50 receives the cleaned non-seismic data output by the adaptive cleaning unit 30, all coupled features output by the coupled feature generation unit 40, and optionally, the original seismic attributes obtained from the data receiving unit 10. This unit combines all the received data at a unified spatial grid location to construct a high-dimensional final feature vector. This final feature vector is output to external storage or directly provided to subsequent machine learning models for processing.
[0030] See attached document Figure 2 , Figure 2 This is a flowchart of a machine learning-based multivariate mineral exploration data cleaning and feature extraction method according to an embodiment of the present invention. The method provided by the present invention can be executed by the aforementioned system and specifically includes the following steps: Step S210: Preprocessing of multivariate heterogeneous data.
[0031] The received 3D seismic data and various non-seismic data are unified into the same 3D coordinate system. Using the voxel grid of the 3D seismic data as a reference, the non-seismic data is spatially meshed and aligned to generate an initial non-seismic data volume that corresponds one-to-one with the 3D seismic data in spatial location.
[0032] Step S220: Construct a multi-scale seismic structure tensor field.
[0033] Based on the preprocessed 3D seismic data in step S210, a spatially continuous and scale-layered structural tensor field is generated through multi-scale analysis and tensor calculation, and structural direction vectors representing the local geological structure direction are extracted from this field.
[0034] Step S230: Structure-guided adaptive data cleaning.
[0035] Using the structural orientation vector generated in step S220 as guiding information, outlier detection and anisotropic missing value interpolation are performed on the initial non-seismic data volume generated in step S210 to obtain cleaned non-seismic data.
[0036] Step S240: Generate hierarchical coupling features.
[0037] Based on the cleaned non-seismic data obtained in step S230 and the multi-scale structural tensor field generated in step S220, new features, including single-scale coupling features and cross-scale coupling features, are generated through a series of preset projections and combined mathematical operations.
[0038] Step S250: Construct and output the final feature vector.
[0039] On a unified spatial grid, the original seismic attributes, cleaned non-seismic data, and all coupled features generated in step S240 are integrated to construct a high-dimensional final feature vector, which is then output for use by subsequent machine learning models.
[0040] The following provides a detailed description of each step of the method of the present invention.
[0041] In step S210, the aim is to establish a unified spatial benchmark for all subsequent analyses. This step receives 3D seismic data volumes within the exploration area from external data sources. and one or more non-seismic data .
[0042] This step begins with coordinate system unification. The system pre-defines a target 3D geographic coordinate system, such as a specific Universal Transverse Mercator (UTM) zone or a mining area coordinate grid specific to the exploration area. For each piece of data received, the system reads its metadata to obtain its original coordinate system information. If the original coordinate system is inconsistent with the target coordinate system, a standard geospatial transformation algorithm is invoked to accurately transform each spatial location point of the data from its original coordinate system to the target coordinate system.
[0043] Secondly, spatial grid alignment is performed based on a unified coordinate system. This process uses the 3D seismic data volume... A voxel mesh is used as the reference 3D mesh. The geometry of this reference 3D mesh is composed of... The dimensions (e.g., number of main survey lines, number of connecting survey lines, number of time or depth sampling points) and spatial sampling intervals (e.g., spacing between main survey lines, spacing between connecting survey lines, time or depth sampling intervals) are uniquely determined.
[0044] For point-distributed non-seismic data, such as borehole geochemical sampling data, the system maps the spatial coordinates of each sampling point to a reference 3D grid to determine its corresponding voxel cell. If multiple sampling points fall within a single voxel cell, a single value is assigned to that voxel according to preset rules (e.g., calculating the arithmetic mean, median, or maximum value of all point values). Voxel cells containing no sampling points are marked as invalid values for subsequent processing.
[0045] For non-seismic data that is already in a gridded form but has a different resolution than the reference 3D grid, such as 2D planar gravity and magnetic anomaly data or 3D electrical resistivity volumes, the system employs a resampling algorithm to map its data values onto the reference 3D grid. Depending on the required data fidelity, resampling algorithms can be selected, including but not limited to nearest neighbor interpolation, trilinear interpolation, or cubic convolution interpolation. For example, for trilinear interpolation, the value of the center point of any voxel in the reference 3D grid is calculated by weighted averaging the values of its eight nearest neighbors in the source data grid.
[0046] After the above processing, step S210 outputs a series of spatially aligned three-dimensional data volumes, including three-dimensional seismic data volumes and corresponding initial non-seismic data volumes with the same dimensions and geometric structure.
[0047] In step S220, a Gaussian scale space is first generated. The purpose of this operation is to transform the original three-dimensional seismic data volume... The data is decomposed into a series of different observation scales to simulate the observation of geological structures of different scales.
[0048] This operation first presets a set of discrete scale parameters. ,in Let be the total number of scale levels, and let the scale parameters satisfy . This set of scale parameters defines the observation range from local details to macroscopic regions. Smaller scale parameters, such as This corresponds to the analysis of fine geological structures; larger scale parameters, such as This corresponds to the analysis of regional, large-scale geological structures. This set of scale parameters can be preset by the user according to geological objectives and characteristics of the work area, for example, by selecting them in a geometric sequence.
[0049] For each preset scale parameter The system constructs a corresponding three-dimensional Gaussian kernel. The mathematical expression for a three-dimensional Gaussian kernel is: ; in, These are spatial coordinates. It is the base of the natural logarithm. This kernel function reaches its maximum value at the origin and decays smoothly with increasing distance from the origin; the rate of decay is determined by the scale parameter. control.
[0050] Subsequently, the system performs convolution operations to convert the preprocessed 3D seismic data volume... With each scale parameter Corresponding Gaussian kernel Perform 3D convolution to generate a corresponding multi-scale smooth data volume. The mathematical expression for this process is: ; In the discrete digital implementation, this integral is a summation operation. The effect of this convolution operation is equivalent to performing a weighted average filter on the original seismic data volume; the weights are determined by the Gaussian kernel function, and the filtering range is determined by the scale parameter. Decide.
[0051] After performing convolution operations on all preset scale parameters, the final output of this step is a set of... Multi-scale smoothed data volume This set of data constitutes the Gaussian scale space of the original seismic data. Among them, It is a data volume that has undergone the slightest smoothing process, retaining the most high-frequency detail information; It is a data volume that has undergone the most powerful smoothing process, suppressing local details and highlighting regional low-frequency macro trends.
[0052] Obtain a set of multi-scale smoothed data volumes in Gaussian scale space. Next, the operation of calculating the structure tensor is performed. This operation is performed for each data volume in scale space. For each voxel location in the 3D grid of the exploration area, calculate a 3x3 matrix to describe the local structural geometry at that location.
[0053] First, for each multi-scale smoothed data volume The system calculates its three-dimensional gradient vector. In a discrete 3D mesh, this gradient vector is approximated using the finite difference method. For example, central difference, the Sobel operator, or the Scharr operator can be used to compute the data volume. Partial derivatives in the three orthogonal directions x, y, z , , Therefore, at each voxel location, the gradient vector is represented as: ; The gradient vector points to the data volume. The direction in which the rate of change of value is greatest and the direction is steepest at that point.
[0054] Next, the system calculates the dyadic product (or outer product) of the gradient vectors. At each voxel location, the gradient vector at that point is calculated. Its own transpose Multiplying these results in a 3x3 initial tensor matrix. This matrix captures the gradient direction information at that single location point. Its mathematical form is: ; To obtain a stable statistical description of the structure within the local neighborhood, rather than relying solely on information from a single voxel, a local spatial weighted average of the initial tensor matrix is required. This operation introduces an integral scale parameter. To achieve this, the system constructs a system with a standard deviation of... Integral Gaussian kernel The integral Gaussian kernel is then convolved with each component of the initial tensor matrix. This process is equivalent to performing a Gaussian weighted summation on the initial tensor in the neighborhood of each voxel.
[0055] Through the above operations, the final result is obtained at the scale. The structure tensor below Its complete mathematical expression is: ; in, Indicated by the integral Gaussian kernel The defined local weighted average operation. The resulting structure tensor. It is a symmetric positive semi-definite matrix that combines the properties of different scales. Lower, Integral Neighborhood Within the range, earthquake data volume Gradient distribution information.
[0056] For all Multi-scale smoothed data volume Repeating the above calculation process, the final output of this step is the multi-scale structure tensor field covering the entire exploration area's three-dimensional grid. .
[0057] The multi-scale structure tensor field was calculated. Then, it is further decomposed into features to extract a quantitative structural direction vector.
[0058] For each structure tensor matrix in the multiscale structure tensor field (Located at a specific location on the three-dimensional grid of the exploration area, and corresponding to the scale) The system performs eigenvalue decomposition on it. Due to the structure tensor... It is a 3x3 real symmetric matrix, and its eigenvalue decomposition will necessarily produce three real eigenvalues and three mutually orthogonal unit eigenvectors.
[0059] This decomposition process involves solving the following characteristic equation: ; in, It is an eigenvalue. These are its corresponding eigenvectors. Solving the equation yields three eigenvalues. and their corresponding three feature vectors .
[0060] To ensure consistency in the analysis, the system sorts the feature values by size, as follows: ; Feature vector With eigenvalues One-to-one correspondence. These three orthogonal eigenvectors Together they form a local coordinate system, the geometric meaning of which is as follows: Feature vector Corresponding to the largest eigenvalue It points to the scale The direction of the greatest change in seismic data signal within a local neighborhood is the local gradient direction.
[0061] Feature vector Corresponding to the smallest eigenvalue It points in the direction where the signal change is most gradual. In the presence of local planar or linear structures (such as stratigraphic interfaces or fault planes), the direction of this vector is perpendicular to the structural surface or parallel to the structural line. Therefore, It is defined as the principal structural direction vector, or normal vector, at that location and scale.
[0062] Feature vector and and Orthogonal, pointing in the direction of signal change, etc. It is related to... They jointly spanned the direction vector of the principal structure. A vertical plane.
[0063] For every location in the three-dimensional grid of the entire exploration area, and for every scale The above feature decomposition and sorting process is repeated for each case.
[0064] The final output of this step is a complete set of multi-scale structural direction vector fields covering the entire region. Specifically, for any point in the 3D mesh, this field provides... Group of orthogonal structural direction vector basis This set of directional information will serve as input for subsequent adaptive cleaning and coupled feature generation. Specifically, the main structure direction vector... (i.e., the vector corresponding to the smallest eigenvalue) is the most frequently used vector in subsequent steps.
[0065] In step S230, structure-guided adaptive data cleaning is performed. First, outlier detection is performed based on structural orientation. This operation utilizes the structural orientation vector extracted in step S220 to provide an objective geometric criterion for distinguishing between random noise and geologically derived anomalies.
[0066] For the initial non-seismic data volume to be cleaned, the system selects one or more analysis scales. The choice of scale should match the expected size of the geological phenomenon being analyzed. For example, a smaller scale is chosen when identifying geochemical anomalies associated with localized small faults. The corresponding structural direction vector field.
[0067] For each voxel in the non-seismic data volume The system defines a 3D local neighborhood centered on the point, such as a 3×3×3 or 5×5×5 voxel window. Within this neighborhood, a robust local background value is computed, such as the arithmetic median of all valid voxel values within the neighborhood. Then, the central voxel point is calculated. The absolute deviation between the value and its local background value .
[0068] The system sets an initial deviation threshold. Anything that satisfies voxel points All were initially marked as "potential outliers". This threshold... It can be determined based on the statistical distribution of the deviation values across the entire data volume (e.g., set as a specific multiple of the standard deviation).
[0069] Next, for each voxel point marked as a "potential outlier" Perform a structural consistency review. The system extracts voxel points from the results of step S220. Location, at the selected scale The main structure direction vector below This vector indicates the most important local structural distribution direction at that point.
[0070] The system is based on vectors The direction is determined by the point in the 3D mesh. The two nearest voxel points that are adjacent and best fit the direction of this vector are denoted as . and The system then checks whether these two neighboring voxel points are also "potential outliers," that is, it calculates their absolute deviations separately. and and with threshold Compare them.
[0071] The final judgment criteria are as follows: If the voxel point It is a "potential outlier" and has at least one neighboring point in the construction direction ( or If a point is also a "potential outlier", then the voxel point is determined. The anomaly exhibits spatial tectonic extension and has been identified as a "geogenic anomaly," with its original values preserved.
[0072] Conversely, if voxel points If a point is identified as a "potential anomaly," but none of its neighboring points in the structural direction are "potential anomalies," then the anomaly is determined to be inconsistent with the local structural morphology and lacks geological origin support, and is thus identified as a "random noise point."
[0073] For all voxels identified as "random noise points," the system modifies their values to a predefined invalid data label (e.g., NaN) for processing in subsequent missing value imputation steps. Through this step, isolated high or low values unrelated to geological structures in the original data are removed, while anomalous zones extending along structural directions are effectively preserved.
[0074] In step S230, missing values are further imputed based on anisotropic distance. This operation aims to fill in all voxels marked as invalid values in the initial non-seismic data volume (including original missing data and points identified as random noise in the previous step), ensuring that the filling results conform to the local geological structures revealed by the seismic data.
[0075] For each target voxel that needs interpolation The system first defines a three-dimensional search neighborhood around it, for example, a neighborhood of size . The cube region is searched, and all neighboring voxels with valid data values are searched within this neighborhood. Simultaneously, the system extracts the target voxel from the results of step S220. Location, at the selected analytical scale The main structure direction vector below .
[0076] The key to this method lies in calculating the target voxel. With any effective neighboring voxel When determining the distance between components, instead of the conventional Euclidean distance, an anisotropic distance that takes into account the structural orientation is used. This anisotropic distance... Calculated in the following way: First, calculate from arrive geometric displacement vector .
[0077] Then, define an anisotropy factor. Its value is greater than 1, for example It can take a value of 5 or 10. This factor is used to quantify the degree of "penalty" for crossing the construction direction.
[0078] Square of anisotropic distance Defined as: ; in, It is the square of the Euclidean distance; It is a displacement vector In the direction vector of the main structure The projected length on the surface represents the length from... arrive The vertical component that traverses the structure; It is a preset anisotropy factor.
[0079] From this definition, it can be seen that if the displacement vector With the main structure direction vector Orthogonal (i.e., voxel) Located in (Points on the same structural plane), then At this point, the anisotropic distance degenerates into the Euclidean distance. Conversely, if the displacement vector... With the main structure in the direction Parallel (i.e., voxel) If the term is located in the direction of the traversing structure, then the term reaches its maximum value, which significantly amplifies the anisotropic distance.
[0080] After calculating the target voxel With all its effective neighboring voxels After determining the anisotropic distance between them, the system uses the inverse distance weighted (IDW) method to calculate... Point interpolation. Point value Determined by the following formula: ; in, yes Effective neighboring voxel set, It is a neighboring voxel Effective value, weight Determined by anisotropic distance: ; in It is a positive real power, usually with a value of 2.
[0081] Through the above steps, interpolation calculations are repeated for all invalid voxel points until there are no more invalid values in the data volume. The final output of this step is a cleaned non-seismic data volume with complete data and numerical distribution that strictly follows the spatial orientation of local geological structures.
[0082] In step S240, the first specific operation in generating hierarchical coupling features is to generate single-scale coupling features. The purpose of this operation is to quantify the correlation between changes in non-seismic physicochemical properties and geological structural orientation at a specific scale.
[0083] This operation receives two input data: The first is the cleaned, fully processed non-seismic data volume generated in step S230. ; Secondly, there is the multi-scale structural direction vector field generated in step S220, covering the entire region, especially the vector field with the minimum eigenvalue. The corresponding principal structural direction vector representing the direction of structural development. .
[0084] First, the system calculates the cleaned non-seismic data volume. 3D spatial gradient This calculation is performed at each voxel location in the 3D mesh, using either the 3D Sobel operator or the central difference method, to obtain... The partial derivatives in the x, y, and z directions. The resulting gradient vector. The direction pointing to the location of the voxel is the direction with the largest local rate of change of non-seismic attribute values and the steepest direction.
[0085] Next, the system will perform analysis on each preset scale. Perform one generation of coupled features. For any voxel point in the 3D mesh, the system extracts the non-seismic data gradient vector of that point. and corresponding scale The main structure direction vector below As before, vectors It is related to the structure tensor The smallest eigenvalue λi,3 is associated with the eigenvector whose direction is parallel to the scale. The distribution direction of the local planar or linear structures identified below.
[0086] The system obtains single-scale coupling features by calculating the dot product of the two vectors mentioned above. The calculation is performed at each voxel point, and its mathematical expression is: ; in, That is, at scale Single-scale coupled feature values generated below; This represents the spatial gradient vector of the cleaned non-seismic data. For scale The main structure direction vector below; This represents the vector dot product operation.
[0087] The physical result of this dot product operation is to project the gradient vector of non-seismic data onto the distribution direction of the local structure. The resulting scalar value... The magnitude of the value directly measures the rate of change of non-seismic attribute values along the tectonic direction. For example, a large positive value indicates a significant increase in non-seismic attributes (such as the content of a certain metallic element) along the tectonic direction.
[0088] Repeat this calculation for all voxel points in the 3D mesh to generate a complete 3D single-scale coupled feature data volume. If multiple preset scales are used... Performing this operation on all cases will result in a final output consisting of a set of single-scale coupled feature data volumes. Each data volume represents the coupling relationship between changes in physical properties and structure at a specific scale.
[0089] In step S240, in addition to generating single-scale coupling features, cross-scale coupling features are also generated. One type of feature is structural consistency. This operation aims to quantify the directional stability and synergy of geological structures at different observation scales.
[0090] The input to this operation is the multi-scale principal structure direction vector field covering the entire area generated in step S220. This operation first requires pre-setting at least one pair of analytical scale combinations. ,in Represents a smaller local observation scale. Represents a large regional observation scale and satisfies Multiple different scale combinations can be preset to explore the relationships between different levels of structure.
[0091] For each preset scale pair The system performs the following calculations at each voxel location in the 3D mesh: First, extract the directions corresponding to the small scales at this location from the multi-scale principal structure direction vector field. and large scale The two principal structural direction vectors are denoted as... and .vector The vector represents the local fine-structure direction of the point. It represents the direction of regional tectonic trends within a larger area including that point.
[0092] Then, the system calculates the absolute value of the dot product of the two unit vectors to generate the structural consistency eigenvalue for that point. Its mathematical expression is: ; in, It corresponds to the scale pair The structural consistency characteristic value; and These are the principal structure direction unit vectors at small and large scales, respectively; This represents the vector dot product operation; This indicates taking the absolute value.
[0093] The calculation result It is a scalar value between 0 and 1. Its magnitude reflects the degree of parallelism between structural directions at two different scales. A value close to 1 indicates that the local structural direction is highly consistent with the regional structural direction, indicating the continuity and stability of the structure at that location. A value close to 0 indicates that the local structural direction is nearly orthogonal to the regional structural direction, indicating a possible abrupt change in structural direction, such as the core of a fold, the intersection of a fault, or a transition zone between structural styles.
[0094] Repeat this calculation for all voxel points in the 3D mesh to generate a complete 3D structural consistency feature data volume. If multiple scale combinations are preset, the final output is a set of structural consistency feature data volumes. Each data volume quantifies the structural complexity of the entire region from a specific hierarchical perspective.
[0095] In step S240, the generated cross-scale coupling features also include hierarchical gradient projection features. The purpose of this operation is to quantify the directional relationship between local variations of non-seismic attributes and regional, large-scale geological structural trends.
[0096] This operation requires two input data: The first is the cleaned non-seismic data volume calculated in the previous steps. 3D spatial gradient ; Secondly, the multi-scale principal structure direction vector field generated in step S220, covering the entire region. .
[0097] First, the system needs to select one or more large-scale parameters from a set of preset scale parameters. This large-scale parameter The corresponding structural tensor field reflects the regional, macroscopic structural distribution characteristics, rather than local details.
[0098] For each selected regional scale The system performs the following calculations at each voxel location in the 3D mesh: The system extracts the gradient vector of non-seismic data at this location. This vector indicates the direction of the fastest local change in non-seismic attribute values at that point.
[0099] At the same time, the system extracts the location corresponding to a large scale. Principal structure direction vector This vector represents the macroscopic distribution direction of the regional structure at that point.
[0100] The system obtains the hierarchical gradient projection features by calculating the dot product of the two vectors mentioned above. The calculation is performed at each voxel point, and its mathematical expression is: ; in, It corresponds to the regional scale. Hierarchical gradient projection eigenvalues; This represents the spatial gradient vector of the cleaned non-seismic data. For regional scale The main structure direction vector below; This represents the vector dot product operation.
[0101] The output of this operation is a scalar value, which physically represents the projected component of the local gradient of non-seismic data along the regional tectonic direction. A large absolute value indicates that the direction of rapid local change in non-seismic attributes closely matches the regional tectonic direction. For example, a large positive value means that the non-seismic attribute value exhibits a significant increasing trend along the regional tectonic direction.
[0102] Repeat this calculation for all voxel points in the 3D mesh to generate a complete 3D hierarchical gradient projection feature data volume. If multiple different regional scales are selected... This will generate a corresponding data volume. These features, from a cross-scale perspective, reveal the degree to which local geophysical and geochemical anomalies are controlled by regional tectonic structures.
[0103] Step S250 aims to construct and output a high-dimensional final feature vector. This step integrates all the information generated in the previous steps on a unified spatial benchmark, providing a structured and comprehensive input for subsequent machine learning analysis.
[0104] This step first aggregates multi-source data. All data to be integrated has been processed in previous steps onto the same spatial grid as the baseline 3D seismic data volume. The dataset to be integrated includes the following categories: 1. Raw seismic attribute data volume: This category includes one or more raw 3D seismic data volumes. Standard seismic properties are calculated directly. These properties include instantaneous amplitude, instantaneous frequency, coherence volume, curvature, etc. Each property constitutes an independent three-dimensional data volume, whose dimensions are completely consistent with the reference grid.
[0105] 2. Cleaned non-seismic data volumes: This category includes one or more cleaned non-seismic data volumes output from step S230. Each non-seismic data point (e.g., different geochemical element contents, gravity anomalies, magnetic anomalies) is represented as an independent three-dimensional data volume.
[0106] 3. Generated Coupled Feature Data Volume: This category includes all features generated in step S240. Specifically, it can be further subdivided into: Single-scale coupled feature body Each scale Each corresponds to a three-dimensional data volume.
[0107] Structural consistency feature Each pair of scale combinations Each corresponds to a three-dimensional data volume.
[0108] Hierarchical gradient projection feature : Each regional scale Each corresponds to a three-dimensional data volume.
[0109] Next, the eigenvector construction operation is performed. This operation is performed at each voxel location of the baseline 3D mesh. Performed independently. For any voxel position. The system extracts the scalar value at that precise location from all the aforementioned data volumes.
[0110] All extracted scalar values are arranged and connected in a pre-defined, fixed order to form a one-dimensional numerical vector, representing the position of the voxel. The final feature vector For example, the structure of this vector can be represented as: ; in, It is the first Earthquake attributes at point The value, It is the first The cleaned non-seismic data is located at [point]. The value of follows the same pattern. The dimension of this vector is equal to the sum of the number of all input data volumes.
[0111] This process is repeated for all voxels in the 3D mesh, ultimately generating a high-dimensional feature dataset. Logically, this dataset can be viewed as a four-dimensional array, where the first three dimensions are spatial coordinates and the fourth dimension is the feature dimension. Alternatively, it can be organized as a two-dimensional data table, where each row represents a voxel and each column represents a feature dimension.
[0112] Finally, this step outputs the constructed high-dimensional feature dataset. The output can be written to a disk file in a specific format, such as HDF5 or CSV, or it can be passed directly in memory as a data structure to subsequent machine learning processing units. This final feature vector set can be directly used to train classification or regression models for target prediction across the entire exploration area in three-dimensional space.
[0113] Example: To further illustrate the technical effects of the method proposed in this invention, a specific application scenario is described below. This embodiment aims to predict the mineralization favorability of a deep porphyry copper deposit in a certain exploration area in three dimensions.
[0114] Example Implementation Background and Data: The geological structure of this exploration area is complex, and it is known that mineralization is closely related to fault systems and rock mass contact zones. Existing data includes: Three-dimensional seismic data volume: High-precision three-dimensional seismic reflection data covering the entire exploration area, with a voxel grid size of 25m × 25m × 10m, serving as the spatial reference for this embodiment.
[0115] Non-seismic data: Borehole geochemical data: Analysis data from core samples from multiple boreholes, providing concentration values of copper (Cu) and molybdenum (Mo) in point form.
[0116] Airborne magnetic data: Two-dimensional gridded total magnetic field intensity anomaly (TMI) data covering the exploration area.
[0117] According to the embodiments of the present invention Figure 2 The process shown above involves processing and analyzing the data. The specific steps are as follows: Step S210, Preprocessing of multivariate heterogeneous data: First, the coordinate system of all data was unified to the Universal Transverse Mercator (UTM) coordinate system for this region. Then, the non-seismic data were spatially aligned using a 25m × 25m × 10m voxel grid of the 3D seismic data volume as a reference. For borehole geochemical data, the Cu and Mo concentration values of each sample are assigned to the voxel containing it based on its three-dimensional coordinates. If a single voxel falls across multiple sample points, the arithmetic mean of its concentration values is taken. Voxels not covered by sample points are marked as invalid (NaN) for their Cu and Mo values.
[0118] For 2D aeromagnetic TMI data, nearest neighbor interpolation is used for resampling, mapping the data values to each (x,y) cylinder of the baseline 3D grid, ensuring that all voxels within each cylinder have the same TMI value. This step yields an initial Cu data volume, an initial Mo data volume, and a complete TMI data volume that spatially correspond one-to-one with the seismic data volume, containing invalid values.
[0119] Step S220: Construct a multi-scale seismic structure tensor field: Generate Gaussian scale space: To capture structures of different scales, from local fractures to regional main faults, a set of scale parameters with standard deviations is preset. The values were set to 1.0, 2.0, 4.0, and 8.0 (unit: voxels), respectively. Four multi-scale smoothed data volumes with different smoothing levels were generated by convolving a 3D Gaussian kernel with the seismic data volume. .
[0120] Computing the structure tensor and eigenvalue decomposition: for each smooth data volume Calculate its three-dimensional gradient and use an integral scale. Gaussian weighted averaging of voxels yields the multi-scale structure tensor field. Then, eigenvalue decomposition is performed on each tensor matrix to extract the corresponding minimum eigenvalue. Principal structure direction vector .
[0121] Step S230: Perform structure-guided adaptive data cleaning. This step is primarily for initial Cu and Mo data volumes containing a large number of invalid values and potential noise.
[0122] Outlier identification: using a medium scale ( The main structure direction vector To guide the process, outlier checks are performed on the Cu and Mo data volumes. If the concentration value of a voxel deviates from the median of its neighborhood by more than a preset threshold, but its orientation is within the structural direction... If the adjacent voxels on the point are not abnormal, then the point is determined to be random noise, and its value is reset to an invalid value.
[0123] Anisotropic missing value imputation: Imputation is performed on all invalid values (original missing values and noise identified in the previous step). During interpolation, anisotropy factors are used. and the main structure direction vector Anisotropic distances were calculated and filled using an inverse distance-weighted method. This operation ensured that the spatial distribution of Cu and Mo concentrations was consistent with the orientation of geological structures at the intermediate scale.
[0124] Step S240: Generate hierarchical coupling features: Generate single-scale coupled features: Calculate the spatial gradient of the cleaned Cu and Mo data volumes and project them onto the principal structure direction vectors of all four scales. A total of 2×4=8 single-scale coupled feature bodies were generated.
[0125] Generate cross-scale coupled features: Structural consistency characteristics: Select two pairs of scale combinations (small scale) and mesoscale Mesoscale With large scale By calculating the absolute value of the dot product of the corresponding main structure direction vectors, two structural consistency feature bodies are generated to characterize the continuity of the construction at different scales.
[0126] Hierarchical gradient projection features: Projecting the local gradients of Cu and Mo onto a regional large scale respectively ( The main structure direction vector Two hierarchical gradient projection feature volumes are generated to characterize the synergistic relationship between local enrichment and regional structure.
[0127] Step S250: Construct and output the final feature vector: At each voxel location, all the following data are integrated: 1. Original earthquake attributes: coherence volume and instantaneous amplitude (2 characteristics).
[0128] 2. Cleaned non-seismic data: Cleaned Cu and Mo data volumes, and resampled TMI data volumes (3 features).
[0129] 3. Generated coupling features: 8 single-scale coupling features, 2 structural consistency features, and 2 hierarchical gradient projection features (a total of 12 features).
[0130] Finally, a high-dimensional feature vector with 2+3+12=17 dimensions was constructed at each voxel location. All feature vectors from the entire exploration area were aggregated to form a comprehensive feature dataset. This dataset can be directly input into machine learning models such as gradient boosting decision trees, using known borehole mineralization information as labels for training. Ultimately, a mineralization favorability prediction map covering the entire 3D exploration area is generated, providing a quantitative scientific basis for subsequent drilling target area selection.
[0131] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multi-source mineral exploration data cleaning and feature extraction system based on machine learning, characterized in that, include: The data receiving unit is used to receive three-dimensional seismic data and non-seismic data within the exploration area. The structural field generation unit is used to generate a multi-scale structural tensor field covering the exploration area based on the three-dimensional seismic data, wherein the multi-scale structural tensor field provides structural direction vectors for each spatial location in the exploration area at multiple preset scales. An adaptive cleaning unit is used to clean the non-seismic data using the structural direction vector provided by the multi-scale structural tensor field, so as to obtain cleaned non-seismic data. The coupling feature generation unit is used to generate coupling features based on the cleaned non-seismic data and the multi-scale structural tensor field; The feature output unit is used to output a final feature vector containing the coupled features for use in a machine learning model.
2. The machine learning-based multivariate mineral exploration data cleaning and feature extraction system according to claim 1, characterized in that, The structural field generation unit is specifically used for: The three-dimensional seismic data is convolved with Gaussian kernels of multiple preset scales to generate a multi-scale smooth data volume. The gradient is calculated based on the multi-scale smooth data volume, and the multi-scale structure tensor field is generated. The multi-scale structural tensor field is decomposed at each spatial location and at each preset scale to obtain the structural direction vector.
3. The multi-source mineral exploration data cleaning and feature extraction system based on machine learning according to claim 2, characterized in that, The adaptive cleaning unit is specifically used for: Based on the structural direction defined by the structural direction vector, the validity of data points in the non-seismic data is determined, and data points identified as noise are processed; or Based on the structural direction vector, anisotropic interpolation is used to fill in the missing values in the non-seismic data.
4. The multi-source mineral exploration data cleaning and feature extraction system based on machine learning according to claim 1, characterized in that, The coupling features generated by the coupling feature generation unit include single-scale coupling features; The coupling feature generation unit is specifically used for: Calculate the spatial gradient of the cleaned non-seismic data; The spatial gradient is projected onto the structural direction vector of the multi-scale structural tensor field at at least one preset scale to generate the single-scale coupling feature.
5. The machine learning-based multivariate mineral exploration data cleaning and feature extraction system according to claim 4, characterized in that, The coupling features also include cross-scale coupling features; The coupling feature generation unit is also used for: The cross-scale coupling feature is generated based on the structural direction vectors of the multi-scale structural tensor field at at least two different preset scales.
6. The multi-source mineral exploration data cleaning and feature extraction system based on machine learning according to claim 5, characterized in that, The at least two different preset scales include small scale and regional scale; The cross-scale coupling feature is a structural consistency feature; The coupling feature generation unit is also specifically used for: The structural consistency feature is determined by calculating the dot product of the structural direction vector at the small scale and the structural direction vector at the regional scale.
7. The multi-source mineral exploration data cleaning and feature extraction system based on machine learning according to claim 6, characterized in that, The cross-scale coupling feature is also a hierarchical gradient projection feature; The coupling feature generation unit is also specifically used for: The hierarchical gradient projection features are generated by combining the single-scale coupling features and the structural consistency features.
8. The multi-source mineral exploration data cleaning and feature extraction system based on machine learning according to claim 1, characterized in that, The non-seismic data includes at least one of geochemical data, gravity data, magnetic data, and electrical data.
9. The multi-source mineral exploration data cleaning and feature extraction system based on machine learning according to claim 1, characterized in that, The final feature vector also includes the cleaned non-seismic data and the seismic attributes extracted from the three-dimensional seismic data.
10. A method for cleaning and extracting features from multivariate mineral exploration data based on machine learning, implemented by the system described in any one of claims 1-9, characterized in that, Includes the following steps: Receives 3D seismic and non-seismic data within the exploration area; Based on the three-dimensional seismic data, a multi-scale structural tensor field covering the exploration area is generated, wherein the multi-scale structural tensor field provides structural direction vectors for each spatial location in the exploration area at multiple preset scales; The non-seismic data is cleaned using the structural direction vector provided by the multi-scale structural tensor field to obtain cleaned non-seismic data. Based on the cleaned non-seismic data and the multi-scale structural tensor field, coupling features are generated; The output is a final feature vector containing the coupled features, which is then used in a machine learning model.
Citation Information
Cited By
Flange forge piece surface defect detection method and system based on image processing
CN122023235A
Flange forging surface defect detection method and system based on image processing
CN122023235B