A spatial data labeling method for delineating a prospecting target area
Through the combination of multi-source data fusion, optimization algorithm and deep learning algorithm, the problem of inaccurate delineation of prospecting targets in mineral exploration has been solved, and efficient and accurate prospecting target area marking has been achieved.
Patent Information
- Application Number
- CN202510838211.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-06-23
AI Technical Summary
Existing technologies lack scientific classification and targeted processing of complex geological data in mineral exploration, resulting in insufficient accuracy and reliability in the delineation of prospecting targets. Traditional methods fail to fully utilize the powerful spatial feature extraction capabilities of deep learning algorithms, affecting the efficiency and accuracy of subsequent analysis.
Multi-source mining area data is collected for preprocessing, and geographic information system technology and bilinear interpolation method are used for data fusion to construct three-dimensional Thiessen polygons. Automatic regularized Gauss-Newton method and support vector machine are combined for optimization, and an evidence layer is constructed. The polygon vector elements are trained through deep learning algorithms to output visual annotation results.
It significantly improves the accuracy and reliability of prospecting target area delineation, provides a clear basis for decision-making on prospecting potential distribution, improves data processing efficiency and retains key geological features in complex areas.
Smart Images

Figure CN120354288B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data annotation, and in particular to a spatial data annotation method for delineating a prospecting target area. Background Art
[0002] In the field of mineral exploration, accurately delineating prospecting targets is crucial to improving the efficiency and accuracy of mineral resource exploration. With the development of geological exploration technology, the limitations of traditional methods that rely solely on remote sensing images or geochemical data have become increasingly prominent. It is impossible to fully explore the geological feature correlations behind the data, resulting in insufficient accuracy and reliability in the delineation of prospecting targets.
[0003] When processing complex geological data, existing spatial data annotation methods lack scientific classification and targeted processing of the annotated areas, affecting the efficiency and accuracy of subsequent analysis. In the model training stage, traditional methods fail to fully utilize the powerful extraction capabilities of deep learning algorithms for spatial features, resulting in insufficient learning of the relationship between geological features and mineralization potential by the annotation model. The output annotation results are difficult to accurately guide mineral exploration practices. Summary of the Invention
[0004] The purpose of the present invention is to provide a spatial data annotation method for delineating prospecting target areas.
[0005] To achieve the above object, the present invention is implemented according to the following technical solutions:
[0006] A first aspect of the present invention provides a spatial data annotation method for delineating a prospecting target area, comprising:
[0007] Collect mining area data and pre-process them to obtain spatial data, which includes remote sensing images, geochemical data, geological structure data, stratigraphic data, lithology data, alteration data, geophysical data, and rock mass data;
[0008] Based on the spatial data points and depth information of the ore body distribution projected onto the surface, 3D Thiessen polygons are constructed. The spatial data is divided using 3D Thiessen polygons to obtain a model area of a typical ore deposit research area to be targeted. The geological structure characteristics, stratigraphic characteristics, rock mass characteristics, geophysical characteristics, geochemical element characteristics, and lithologic characteristics of the model area of the typical ore deposit research area to be targeted are extracted to obtain initial values.
[0009] Randomly generate multiple sets of initial weight values for the initial values, use the automatic regularized Gauss-Newton method to perform linear approximation optimization on the initial weight values, iteratively calculate to obtain the first index, and perform nonlinear optimization on the initial values and initial weight values based on the support vector machine to obtain the second index;
[0010] The first index is used as the input parameter of the evidence weight method to construct the evidence layer, the first label and the second label are obtained according to the threshold difference of the evidence layer, and the first label is corrected based on the second index to obtain the corrected label;
[0011] Use correction labels and second labels to mark the model area of the typical mineral deposit research area to obtain the marked area, represent the marked area with polygonal vector elements, use deep learning algorithms to train polygonal vector elements and marked areas to obtain a spatial data marking model, and use the spatial data marking model to output visual marking results.
[0012] As a further method, the method of collecting mining area data and preprocessing to obtain spatial data includes:
[0013] Collect elevation remote sensing images, collect geochemical data with drilling depth information, collect geological structure data including underground structure extension depth information, collect stratigraphic data, collect rock mass data, collect geophysical property data, collect lithological data, collect alteration data, collect mineralization data, collect ore body data, use geographic information system technology to use elevation remote sensing images as spatial coordinate base, superimpose them on the coordinate space coordinate base in the form of point elements according to the coordinates of the sampling points of geochemical data, superimpose geological structure data on the spatial coordinate base in the form of line elements or surface elements according to geographic space coordinates, use bilinear interpolation method to unify the resolution of the spatial coordinate base, and form spatial data.
[0014] As a further method, the method of constructing Thiessen polygons based on spatial data points and depth information projected onto the surface of the ore body distribution, using Thiessen polygons to divide the spatial data to obtain a model area of a typical ore deposit research area to be marked, and extracting geological structure characteristics, stratigraphic characteristics, rock mass characteristics, geophysical characteristics, geochemical element characteristics, and lithologic characteristics of the data of the model area of the typical ore deposit research area to be marked to obtain initial values includes:
[0015] Extract spatial data points and the depth information corresponding to the spatial data points based on spatial data, import the spatial data points into the geographic information system, construct different spatial data point planes based on the depth information of different spatial data points, calculate the Euclidean distance between any two points in the spatial data point plane, use the Delaunay triangulation algorithm to generate the perpendicular bisectors of the lines connecting adjacent spatial data points, use the perpendicular bisectors to intersect to form the Thiessen polygon boundary, and divide the spatial data point plane into different typical mineral deposit research target area model areas based on the Thiessen polygon boundary;
[0016] The method extracts the geological structure characteristics and element characteristics of the model area data of the typical mineral deposit research area to be marked to obtain the initial value, obtains the fault intersection density based on the geological structure data of the model area of the typical mineral deposit research area to be marked, counts the element content outliers and the proportion of outliers in the geochemical exploration data in the model area of the typical mineral deposit research area to be marked, normalizes the outliers and uses them as weight values, and performs weighted summation on the fault intersection density and the proportion of outliers to obtain the initial value.
[0017] As a further method, the method of randomly generating multiple groups of initial weight values for the initial values, performing linear approximation optimization on the initial weight values using the automatic regularized Gauss-Newton method, and iteratively calculating to obtain the first exponent includes:
[0018] For the initial value, multiple groups of initial weight values are randomly generated according to the data category of the model area of the typical mineral deposit research area to be marked, and the initial value of the model area of the typical mineral deposit research area to be marked and the data of the model area of the typical mineral deposit research area to be marked are used to construct a first function through the ridge regression algorithm. Each group of initial weight values is used as the starting point of the iteration, and the parameter weights of the first function are iteratively optimized by the automatic regularized Gauss-Newton method. The Jacobian matrix is used to adjust the parameter weights during the iteration until the first function converges to obtain the optimized parameter weights. The formula of the automatic regularized Gauss-Newton method is:
[0019]
[0020] in For the The weight vector for the iteration, For the The weight vector for the iteration, is the Jacobian matrix, is the transposed matrix, which is multiplied with the Jacobian matrix to form the approximate Hessian matrix. For the The regularization parameter for the iteration, For the In the iteration The residual between the initial value of the model area and the predicted value of the first function in the target area of the typical mineral deposit study is For the In the iteration The residual between the initial value of the model area and the predicted value of the first function in the target area of the typical mineral deposit study is is a very small positive number to prevent the denominator from being zero. is the identity matrix, For the The residual vector of the iteration, is the total number of model areas to be benchmarked for typical mineral deposit research;
[0021] The first exponent is obtained by multiplying the optimization parameter weight by the first function parameter value and then summing them.
[0022] As a further method, the method of performing nonlinear optimization on the initial value and the initial weight value based on the support vector machine to obtain the second index includes:
[0023] The initial value is used as the target variable, the initial weight value corresponding to the initial value is used as the characteristic variable, and the training data set is obtained based on the initial values of the model area of the target area of all typical mineral deposits in the spatial data;
[0024] The support vector machine model is selected based on the kernel function. If the Pearson coefficient between the initial values is greater than 0.7, the linear kernel is selected. If the global Moran index between the initial values is greater than 0.5, the polynomial kernel function is selected, and if it is less than 0.3, the radial basis kernel function is selected. The support vector machine is trained using the training data set, and the support vector machine parameters are optimized by reducing the structural risk using the sequential minimum optimization algorithm to obtain a trained support vector machine model. The initial weight values of the model area of the typical mineral deposit research area to be marked are input into the trained support vector machine model to obtain the output value, and the output value is normalized to obtain the second index.
[0025] As a further method, the method of using the first index as an input parameter of the evidence weight method to construct an evidence layer and obtaining the first label and the second label according to the difference in the evidence layer threshold value includes:
[0026] Based on the model area of the typical mineral deposit research target area, the first index is used as the input parameter of the evidence weight method. The remote sensing images, geochemical data, geological structure data, stratigraphic data, lithological data, alteration data, geophysical data, and rock mass data of the model area of the typical mineral deposit research target area are used as evidence factors. The conditional probability of the evidence factors under the mineralization and non-metallization states is statistically analyzed. The natural logarithm of the ratio of the probability of the evidence factors under the mineralization and non-metallization conditions is used to calculate the weight of the evidence factors. The geographic information system is used to perform spatial overlay calculation on the evidence factor weight layers to form an evidence layer.
[0027] Frequency statistics are performed on the evidence layer values, and the potential threshold is set using the natural break point classification method based on the data distribution. The labels are divided according to the comparison results between the evidence layer values and the potential threshold. If the evidence layer value is greater than the potential threshold, it is classified as the first label; if it is less than the potential threshold, it is classified as the second label.
[0028] As a further method, the method of correcting the first label based on the second index to obtain a corrected label includes:
[0029] The element content abnormal data of the model area of the typical mineral deposit to be studied were interpolated using the IDW and global kriging methods to generate an element content grid map. The mean and standard deviation of the grid map were calculated. The lower limit of the abnormal value was obtained by adding twice the standard deviation to the element mean. The first correction parameter was obtained by dividing the single element content by the element mean using the contrast value method.
[0030] The fault length per unit area of the model area of the typical mineral deposit research area is calculated to extract the fault density. The line density analysis of the fold axis data is performed to count the number of folds per unit distance to extract the fold frequency. The slope grid is generated using the DEM to calculate the slope standard deviation. The fault density, fold frequency, and slope standard deviation are normalized and then weighted and summed to obtain the second correction parameter.
[0031] The first label is corrected using the second index and the first and second correction parameters through a correction formula, where the correction formula is:
[0032]
[0033] in To correct the label, is the first index, is the second index, is the second correction parameter, is the maximum value of the second correction parameter, is the first correction parameter, is the maximum value of the first correction parameter, is the first label, 、 、 、 is the coefficient of variation, is the subscript of the coefficient of variation.
[0034] As a further method, the method of using the correction label and the second label to mark the model area of the typical mineral deposit research area to obtain the marked area, and representing the marked area with polygonal vector elements includes:
[0035] The calibration label and the second label are used to label the model area of the typical mineral deposit research area to obtain the labeling area, and the variance of the data at the boundary of the labeling area is calculated. Based on the variance, the complexity threshold is set as the sum of the 75% quantile of the variance and 1.5 times the interquartile range using the quartile method. The complexity threshold is used to divide the low-complexity labeling area into the high-complexity labeling area, and the low-complexity labeling area is decomposed using the row-column convex decomposition formula. The formula is:
[0036]
[0037] in is the set of all convex polygons after decomposition, represent Any convex polygon in is an independent convex area unit after decomposition. Represents a convex polygon Any vertex in is used to limit the object range of the internal angle calculation. vertex an interior angle of the triangle, a vector pointing from the vertex to the vertex , a vector pointing from the vertex to the vertex , a low complexity annotation region, a k-th cutting operation performed on , a union operation;
[0038] the low complexity annotation region is a regular polygon vector element; and the high complexity annotation region is extracted to obtain a polygon vector element by connecting boundary coordinate points in sequence to form a closed figure.
[0039] As a further method, the method for training a spatial data annotation model using a deep learning algorithm and outputting a visual annotation result includes:
[0040] normalizing the geometric coordinates of the polygon vector element, one-hot encoding the annotation region category, generating a learning data set based on the polygon vector element and the one-hot encoding, and dividing the learning data set into a training set, a validation set, and a test set;
[0041] inputting the training set into a deep learning model, extracting spatial features of the polygon vector element through convolution layers and pooling layers, performing preliminary prediction based on the spatial features, calculating the difference between the prediction result and the one-hot encoded label using a cross-entropy loss function, inputting the validation set into the model to verify the overfitting state of the model through the loss function value and the accuracy rate change, adjusting the model parameters by feeding back the difference to each layer of the model through a back propagation algorithm, and continuously reducing the difference until the model training is completed, inputting the test set into the trained deep learning model, evaluating the model performance using the accuracy rate and the recall rate, obtaining a spatial data annotation model, and outputting a visual annotation result using the spatial data annotation model.
[0042] The second aspect of the present application provides a spatial data annotation system for delineating a prospecting target area, comprising:
[0043] a data acquisition module for acquiring mine data for preprocessing to obtain spatial data, wherein the spatial data includes remote sensing images, geochemical exploration data, geological structure data, stratum data, rock mass data, geophysical exploration data, and lithology data;
[0044] The module for generating the model area of the typical mineral deposit research area to be marked is used to construct Thiessen polygons based on spatial data points and depth information, use Thiessen polygons to divide spatial data to obtain the model area of the typical mineral deposit research area to be marked, and extract the geological structure characteristics and element characteristics of the model area data of the typical mineral deposit research area to be marked to obtain initial values;
[0045] The index calculation module is used to randomly generate multiple sets of initial weight values for the initial values, use the automatic regularized Gauss-Newton method to perform linear approximation optimization on the initial weight values, iteratively calculate to obtain the first index, and perform nonlinear optimization on the initial values and initial weight values based on the support vector machine to obtain the second index;
[0046] a label generation module, configured to construct an evidence layer using the first index as an input parameter of the weight of evidence method, obtain a first label and a second label according to a threshold difference of the evidence layer, and correct the first label based on the second index to obtain a corrected label;
[0047] The label model generation module is used to use the correction label and the second label to label the model area of the typical mineral deposit research area to obtain the labeled area, represent the labeled area by polygonal vector elements, use the deep learning algorithm to train the polygonal vector elements and the labeled area to obtain the spatial data labeling model, and use the spatial data labeling model to output the visual labeling results.
[0048] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects:
[0049] (1) The present invention collects multi-source mining area data such as remote sensing images, geochemical data, geological structure data, stratigraphic data, rock mass data, geophysical data, and lithologic data, and uses geographic information system technology and bilinear interpolation method for preprocessing to achieve the organic fusion of multi-source data, more comprehensively reflect the geological characteristics of the mining area, and significantly improve the accuracy and reliability of the delineation of the prospecting target area.
[0050] (2) The present invention uses the automatic regularized Gauss-Newton method and support vector machine to perform linear approximation optimization and nonlinear optimization respectively, combines the evidence weight method to construct the evidence layer and corrects the labels through multiple parameters to realize the judgment of mineralization potential.
[0051] (3) The present invention divides the labeled area into complexity groups, adopts the row-column convex decomposition formula for the low-complexity labeled area to obtain regular vector elements, retains the original outline of the high-complexity labeled area, and combines the deep learning algorithm to extract the spatial features of the vector elements and train the model. This differentiated processing not only improves the data processing efficiency, but also ensures that the key geological features of the complex area are not lost. The final output visual labeling results can intuitively present the distribution of prospecting potential and provide a clear decision-making basis for geological exploration. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 The present invention is a flowchart of the steps of a spatial data annotation method for delineating a prospecting target area in an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0054] Reference Figure 1 As shown, the present invention provides a spatial data annotation method for delineating a prospecting target area, comprising:
[0055] Collect typical ore deposit data of the mining area and pre-process them to obtain spatial data, which includes remote sensing images, geochemical data, geological structure data, stratigraphic data, rock mass data, geophysical data, and lithologic data;
[0056] During the actual assessment, remote sensing imagery was collected to obtain multispectral images of the mining area, including four bands: blue, green, red, and near-infrared. Geochemical data was also collected, and data from 1,200 soil sampling points covering an area of 80 square kilometers were collected. Eight elements, including Cu, Pb, Zn, Au, and Ag, were detected at a sampling depth of 0-20 cm, and drilling verification was carried out to a depth of 300 meters. Geological structural data were also collected, and 28 major faults were interpreted using geophysical seismic reflection technology, including their strike, dip, and underground extension depths of 50-800 meters. Twelve fold axes were extracted based on geological mapping, and the axis positions and flank attitudes were recorded. ArcGIS Pro was used to spatially register the remote sensing images with the geochemical points and geological structural lines in the WGS84 coordinate system. Radiometric calibration and atmospheric correction were performed on the remote sensing images to enhance mineralization and alteration information. Bilinear interpolation was used to resample all data to a 5-meter resolution to generate spatial data, including remote sensing bands, geochemical elements, and geological structural attributes.
[0057] Based on the spatial data points and depth information of the ore body distribution projected onto the surface, 3D Thiessen polygons are constructed. The spatial data is divided using 3D Thiessen polygons to obtain a model area of a typical ore deposit research area to be targeted. The geological structure characteristics and element characteristics of the model area of the typical ore deposit research area to be targeted are extracted to obtain initial values.
[0058] In the actual evaluation, the geochemical point coordinates and depth information are extracted, and 0-100 m, 100-200 m, and 200-300 m depth layers are constructed. In each depth layer, Thiessen polygons are generated using the CreateThiessenPolygons tool, and a total of 896 typical deposit research target area model areas are divided in the spatial data. The number of fault intersection points within 1 km² of each typical deposit research target area model area is counted, for example, the number of fault intersection points in a certain typical deposit research target area model area is 12 / km², the fold axis length per unit area in the typical deposit research target area model area is calculated, and the element characteristics of the typical deposit research target area model area are extracted. Taking Cu element as an example, the mean value is 80 ppm, the standard deviation is 25 ppm, and the outlier value is defined as the mean value plus twice the standard deviation (> 130 ppm). The proportion of outliers in a certain typical deposit research target area model area is calculated, for example, there are 15 geochemical points, of which 3 are outliers, accounting for 20%, and the normalized weight value is 0.2. The fault intersection density and the proportion of outliers are weighted and summed to obtain an initial value of 7.28.
[0059] A plurality of sets of initial weight values are randomly generated, and the automatic regularization Gauss-Newton method is used to linearly approximate and optimize the initial weight values. Iterative calculation obtains a first index, and a support vector machine is used to nonlinearly optimize the initial value and the initial weight value to obtain a second index.
[0060] It should be explained that there are some linear relationships in the spatial data, such as some geological variables within a certain range showing an approximate linear correlation with the mineralization potential. The automatic regularization Gauss-Newton method can efficiently optimize these linear relationships. It uses the Jacobian matrix to construct an approximate Hessian matrix through iterative calculation, and gradually adjusts the weight to make the model quickly and accurately fit the linear part in the spatial data, so as to obtain a more accurate first index for representing the linear characteristics of the mineralization potential. There are also complex nonlinear relationships between the characteristics of the spatial data, such as the comprehensive influence of geological structure and element distribution on the mineralization potential. The support vector machine can map the data to a high-dimensional feature space through the kernel function, and convert the originally non-linearly separable data in the low-dimensional space into linearly separable data, so as to effectively handle these nonlinear relationships and obtain a second index for mining the nonlinear characteristics of the mineralization potential.
[0061] In the actual evaluation, the regularization parameter λ is set to 0.3, and the first function is constructed as follows: wherein is the fault intersection density, is the proportion of outliers, the automatic regularization Gauss-Newton method is used to linearly approximate and optimize the initial weight value, the initial weight value is randomly generated, for example is 0.5, is 0.3, the maximum number of iterations is set to 100, the convergence threshold is 1e-6, the regularization parameter is dynamically adjusted, the initial α is 0.1, and the weight parameter is obtained after optimization. is 0.65, is 0.43, is 0.2, and the first exponent is calculated to be 8.176.
[0062] In the actual evaluation, the initial value is used as the target variable, and the initial weight value corresponding to the initial value is used as the characteristic variable. The training data set is obtained based on the initial values of the model area of all typical mineral deposits in the spatial data. The global Moran index between the initial values is calculated to be 0.2 less than 0.3, then the radial basis kernel function RBF is selected, and C is obtained as 1 and γ is 0.1 through grid search. The support vector machine is trained using the training data set. For example, if the initial weight value is input as [0.5, 0.3], the second index after output normalization is 0.82.
[0063] The first index is used as the input parameter of the evidence weight method to construct the evidence layer, the first label and the second label are obtained according to the threshold difference of the evidence layer, and the first label is corrected based on the second index to obtain the corrected label;
[0064] In the actual assessment, the first index is used as the input parameter of the evidence weight method based on the model area of the typical mineral deposit research area to be marked. The remote sensing images, geochemical data and geological structure data of the model area of the typical mineral deposit research area to be marked are used as evidence factors. The conditional probability of fault density in mineralization or non-mineralization areas is statistically analyzed. For example, the probability of fault density greater than 10 / km² in mineralization areas is 0.75, and that in non-mineralization areas is 0.3. The natural logarithm of the ratio of the probability of occurrence of evidence factors under mineralization and non-mineralization conditions is used to calculate the evidence factor weight of 0.916. ArcGIS Pro is used to spatially overlay and calculate the evidence layer to generate a continuous grid. The natural break point classification method is used to divide the evidence layer into 5 categories. The potential threshold is set to the fourth category with an upper limit of 0.75, and the first label is greater than 0.75 and the second label is less than or equal to 0.75.
[0065] In the actual assessment, the element content anomaly data of the model area of the typical mineral deposit research area to be marked were interpolated using IDW, global kriging and other methods to generate an element content grid map, and the mean and standard deviation of the grid map were calculated. For example, the Cu element mean of 80ppm plus twice the standard deviation of 50ppm was used to obtain an anomaly lower limit of 130ppm. The contrast value method was used to divide the single element content by the element mean to obtain the first correction parameter of 1.625. The geological data of the model area of the typical mineral deposit research area to be marked were used to calculate the fault length per unit area to extract the fault density. The line density analysis of the fold axis data was performed to count the number of folds per unit distance to extract the fold frequency. The slope grid was generated using DEM to calculate the slope standard deviation. The fault density of 12 / km², the fold frequency of 500m / km², and the slope standard deviation of 15° were normalized to obtain the weighted sum of 0.9, 0.7, and 0.8, and the second correction parameter was obtained as 0.81. The second index and the first and second correction parameters were used to correct the first label through the correction formula to obtain a corrected label of 0.13
[0066] Use correction labels and second labels to mark the model area of the typical mineral deposit research area to obtain the marked area, represent the marked area with polygonal vector elements, use deep learning algorithms to train polygonal vector elements and marked areas to obtain a spatial data marking model, and use the spatial data marking model to output visual marking results.
[0067] In the actual assessment, the correction label and the second label are used to mark the model area of the typical mineral deposit research area to obtain the marked area, and the variance of the marked area boundary data is calculated. For example, the variance of the model area of the typical mineral deposit research area to be marked is 95. Based on the variance, the quartile method is used to set the complexity threshold as the sum of the 75% quantile of the variance and 1.5 times the interquartile range, and the complexity threshold is 210. The complexity threshold is used to divide the low-complexity marked area and the high-complexity marked area. The low-complexity area with a variance lower than 210 is decomposed into 4 convex polygons using the row-column convex decomposition formula. For example, the vertex coordinate sequence of a convex polygon is [(100,200),(150,250),(200,220),(120,180),(180,190),(140,230)]. The high-complexity area with a variance higher than 210 is directly extracted with 12 coordinate points to form a closed curve. The coordinate points are connected in sequence to form a closed figure to obtain polygonal vector elements.
[0068] In the actual evaluation, the geometric coordinates of the polygonal vector features are normalized, the labeled area categories are one-hot encoded, and a learning dataset is generated based on the polygonal vector features and one-hot encoding. The learning dataset is divided into training set, validation set, and test set according to the ratio of 7:2:1.
[0069] Select the U-Net model, set the input layer to normalized coordinates (x_norm, y_norm), the encoding path to four 3×3 convolutional layers, the maximum pooling to 2×2, the decoding path to a jump connection between the upsampling layer and the encoding layer, and the output layer to 3 The training parameters are set as follows: batch size is 16, learning rate is 0.0001, iteration is 200 rounds, loss function is set as cross entropy, Adam is selected as optimizer, and training is stopped when the accuracy of the validation set is greater than or equal to 95%. The final test set indicators are calculated, with an accuracy of 96.2%, a recall rate of 0.94, and an F1 of 0.95. The spatial data annotation model is obtained and the visual annotation results are output using the spatial data annotation model. The visualization module of the prospecting prediction spatial data annotation system is used to render the annotation results into thematic maps. High potential area (red): 12% of the area is marked as Class I target area, medium potential area (yellow): 25% of the area is marked as Class II target area, low potential area (blue): 63% of the area is marked as Class III area, remote sensing images and geological structure lines are superimposed to generate a three-dimensional visualization scene with contour lines.
[0070] In this embodiment, the method of collecting mining area data and preprocessing to obtain spatial data includes:
[0071] Collect elevation remote sensing images, collect geochemical data with drilling depth information, collect geological structure data including underground structure extension depth information, collect stratigraphic data, collect rock mass data, collect geophysical property data, collect lithologic data, collect alteration data, collect mineralization data, collect ore body data, use geographic information system technology to use elevation remote sensing images as spatial coordinate base, superimpose them on the coordinate space coordinate base in the form of point elements based on the coordinates of the geochemical data sampling points, superimpose geological structure and other data on the spatial coordinate base in the form of line elements or surface elements according to geographic coordinates, use bilinear interpolation to unify the resolution of the spatial coordinate base, and form spatial data.
[0072] In this embodiment, the method of constructing Thiessen polygons based on spatial data points and depth information, using Thiessen polygons to divide spatial data to obtain a model area of a typical mineral deposit research area to be marked, and extracting geological structure characteristics, stratigraphic characteristics, rock mass characteristics, geophysical characteristics, geochemical element characteristics, lithologic characteristics, etc. of the model area of the typical mineral deposit research area to be marked to obtain initial values includes:
[0073] Extract spatial data points and the depth information corresponding to the spatial data points based on spatial data, import the spatial data points into the geographic information system, construct different spatial data point planes based on the depth information of different spatial data points, calculate the Euclidean distance between any two points in the spatial data point plane, use the Delaunay triangulation algorithm to generate the perpendicular bisectors of the lines connecting adjacent spatial data points, use the perpendicular bisectors to intersect to form the Thiessen polygon boundary, and divide the spatial data point plane into different typical mineral deposit research target area model areas based on the Thiessen polygon boundary;
[0074] The method extracts the geological structure characteristics and element characteristics of the model area data of the typical mineral deposit research area to be marked to obtain the initial value, obtains the fault intersection density based on the geological structure data of the model area of the typical mineral deposit research area to be marked, counts the element content outliers and the proportion of outliers in the geochemical exploration data in the model area of the typical mineral deposit research area to be marked, normalizes the outliers and uses them as weight values, and performs weighted summation on the fault intersection density and the proportion of outliers to obtain the initial value.
[0075] In this embodiment, the method of randomly generating multiple groups of initial weight values for the initial values, performing linear approximation optimization on the initial weight values using the automatic regularized Gauss-Newton method, and iteratively calculating to obtain the first index includes:
[0076] For the initial value, multiple groups of initial weight values are randomly generated according to the data category of the model area of the typical mineral deposit research area to be marked, and the initial value of the model area of the typical mineral deposit research area to be marked and the data of the model area of the typical mineral deposit research area to be marked are used to construct a first function through the ridge regression algorithm. Each group of initial weight values is used as the starting point of the iteration, and the parameter weights of the first function are iteratively optimized by the automatic regularized Gauss-Newton method. The Jacobian matrix is used to adjust the parameter weights during the iteration until the first function converges to obtain the optimized parameter weights. The formula of the automatic regularized Gauss-Newton method is:
[0077]
[0078] in For the The weight vector for the iteration, For the The weight vector for the iteration, is the Jacobian matrix, is the transposed matrix, which is multiplied with the Jacobian matrix to form the approximate Hessian matrix. For the The regularization parameter for the iteration, For the In the iteration The residual between the initial value of the model area and the predicted value of the first function in the typical mineral deposit study area is For the In the iteration The residual between the initial value of the model area and the predicted value of the first function in the target area of the typical mineral deposit study is is a very small positive number to prevent the denominator from being zero. is the identity matrix, For the The residual vector of the iteration, is the total number of model areas to be benchmarked for typical mineral deposit research;
[0079] The first exponent is obtained by multiplying the optimization parameter weight by the first function parameter value and then summing them.
[0080] In this embodiment, the method for obtaining the second index by performing nonlinear optimization on the initial value and the initial weight value based on the support vector machine includes:
[0081] The initial value is used as the target variable, the initial weight value corresponding to the initial value is used as the characteristic variable, and the training data set is obtained based on the initial values of the model area of the target area of all typical mineral deposits in the spatial data;
[0082] The support vector machine model is selected based on the kernel function. If the Pearson coefficient between the initial values is greater than 0.7, the linear kernel is selected. If the global Moran index between the initial values is greater than 0.5, the polynomial kernel function is selected, and if it is less than 0.3, the radial basis kernel function is selected. The support vector machine is trained using the training data set, and the support vector machine parameters are optimized by reducing the structural risk using the sequential minimum optimization algorithm to obtain a trained support vector machine model. The initial weight values of the model area of the typical mineral deposit research area to be marked are input into the trained support vector machine model to obtain the output value, and the output value is normalized to obtain the second index.
[0083] In this embodiment, the method of using the first index as an input parameter of the weight of evidence method to construct an evidence layer and obtaining the first label and the second label according to the difference in the threshold value of the evidence layer includes:
[0084] Based on the model area of the typical mineral deposit research target area, the first index is used as the input parameter of the evidence weight method. The remote sensing images, geochemical data and geological structure data of the model area of the typical mineral deposit research target area are used as evidence factors. The conditional probability of the evidence factors under the mineralization and non-metallization states is statistically analyzed. The natural logarithm of the ratio of the probability of the evidence factors under the mineralization and non-metallization conditions is used to calculate the evidence factor weight. The geographic information system is used to perform spatial overlay calculation on the evidence factor weight layer to form an evidence layer.
[0085] Frequency statistics are performed on the evidence layer values, and the potential threshold is set using the natural break point classification method based on the data distribution. The labels are divided according to the comparison results between the evidence layer values and the potential threshold. If the evidence layer value is greater than the potential threshold, it is classified as the first label; if it is less than the potential threshold, it is classified as the second label.
[0086] In this embodiment, the method of correcting the first label based on the second index to obtain a corrected label includes:
[0087] The element content abnormal data of the model area of the typical mineral deposit research area to be marked are interpolated using methods such as IDW and global kriging to generate element content grid maps, and the mean and standard deviation of the grid maps are calculated. The lower limit of the abnormal value is obtained by adding twice the standard deviation to the element mean, and the first correction parameter is obtained by dividing the single element content by the element mean using the contrast value method.
[0088] The fault length per unit area of the model area of the typical mineral deposit research area is calculated to extract the fault density. The line density analysis of the fold axis data is performed to count the number of folds per unit distance to extract the fold frequency. The slope grid is generated using the DEM to calculate the slope standard deviation. The fault density, fold frequency, and slope standard deviation are normalized and then weighted and summed to obtain the second correction parameter.
[0089] The first label is corrected using the second index and the first and second correction parameters through a correction formula, where the correction formula is:
[0090]
[0091] in To correct the label, is the first index, is the second index, is the second correction parameter, is the maximum value of the second correction parameter, is the first correction parameter, is the maximum value of the first correction parameter, is the first label, 、 、 、 is the coefficient of variation, is the subscript of the coefficient of variation.
[0092] In this embodiment, the method of using the correction label and the second label to mark the model area of the typical mineral deposit research area to obtain the marked area, and representing the marked area with polygonal vector elements includes:
[0093] The calibration label and the second label are used to label the model area of the typical mineral deposit research area to obtain the labeling area, and the variance of the data at the boundary of the labeling area is calculated. Based on the variance, the complexity threshold is set as the sum of the 75% quantile of the variance and 1.5 times the interquartile range using the quartile method. The complexity threshold is used to divide the low-complexity labeling area into the high-complexity labeling area, and the low-complexity labeling area is decomposed using the row-column convex decomposition formula. The formula is:
[0094]
[0095] wherein is a set composed of all convex polygons after decomposition, represents any one convex polygon in is a convex region unit after decomposition, represents a convex polygon any vertex in is used to limit the object range of the inner angle calculation, vertex inner angle at is a vector from vertex to , is a vector from vertex to , is a low-complexity labeling area, is the kth cutting operation on , is a union operation;
[0096] The low-complexity labeling area is a regular polygon vector element; for the high-complexity labeling area, the boundary coordinate points are extracted, the coordinate points are sequentially connected to form a closed figure, and a polygon vector element is obtained.
[0097] In the embodiment, the method for training a spatial data labeling model using a deep learning algorithm, and outputting a visual labeling result using the spatial data labeling model, comprises the following steps:
[0098] The geometric coordinates of the polygon vector element are normalized, the labeling area categories are one-hot encoded, a learning data set is generated based on the polygon vector element and the one-hot encoding, and the learning data set is divided into a training set, a validation set and a test set;
[0099] The training set is input into a deep learning model, the model extracts spatial features of the polygon vector element through a convolution layer and a pooling layer, performs a preliminary prediction according to the spatial features, calculates the difference between the prediction result and the one-hot encoded label using a cross-entropy loss function, inputs the validation set into the model during training to verify the overfitting state of the model through the loss function value and the accuracy change, feeds the difference back to the layers of the model through a back propagation algorithm to adjust the model parameters, so that the difference is continuously reduced until the model training is completed, inputs the test set into the trained deep learning model, evaluates the model performance using the accuracy and the recall rate, obtains a spatial data labeling model, and outputs a visual labeling result using the spatial data labeling model.
[0100] The second aspect of the application also provides a spatial data labeling system for delineating a prospecting target area, comprising:
[0101] A data collection module is configured to collect mine data for preprocessing to obtain spatial data, which includes spatial data of structure, stratum, rock mass, geochemical exploration, geophysical exploration, remote sensing, heavy sand, ore body, mineralization, alteration, lithology, etc.
[0102] A typical deposit research area to be labeled model area generation module is configured to construct a three-dimensional Thiessen polygon based on spatial data points and depth information of the ore body distribution projected to the surface, divide the spatial data using the three-dimensional Thiessen polygon to obtain an ore body model, and extract geological structure features, stratum features, rock mass features, geophysical exploration features, geochemical exploration element features, and lithology features of the typical deposit research area to be labeled model area data to obtain initial values.
[0103] An index calculation module is configured to randomly generate multiple sets of initial weight values for the initial values, perform linear approximation optimization on the initial weight values using an automatic regularization Gauss-Newton method, and iteratively calculate to obtain a first index, and perform nonlinear optimization on the initial values and the initial weight values based on a support vector machine to obtain a second index.
[0104] A label generation module is configured to use the first index as an input parameter of an evidence weight method to construct an evidence layer, obtain a first label and a second label according to threshold differences of the evidence layer, and correct the first label based on the second index to obtain a corrected label.
[0105] A label model generation module is configured to use the corrected label and the second label to label a typical deposit research area to be labeled model area to obtain a labeled area, represent the labeled area through a polygon vector element, train the polygon vector element and the labeled area using a deep learning algorithm to obtain a spatial data labeling model, and output a visual labeling result using the spatial data labeling model.
[0106] The above content is merely an example and description of the structure of the present application, and those skilled in the art can make various modifications or supplements or use similar ways to replace the described specific embodiments, as long as they do not deviate from the structure of the present application or exceed the scope defined by the present claims, and should be within the protection scope of the present application.
Claims
1. A spatial data annotation method for delineating a prospecting target area, characterized in that: The following steps are involved: Collect mining area data and pre-process them to obtain spatial data, which includes remote sensing images, geochemical data and geological structure data; Construct Thiessen polygons based on spatial data points and depth information, use Thiessen polygons to divide spatial data to obtain sub-target areas, and extract geological structural characteristics and element characteristics of sub-target area data to obtain initial values; Randomly generate multiple sets of initial weight values for the initial values, use the automatic regularized Gauss-Newton method to perform linear approximation optimization on the initial weight values, iteratively calculate to obtain the first index, and perform nonlinear optimization on the initial values and initial weight values based on the support vector machine to obtain the second index; The first index is used as the input parameter of the evidence weight method to construct the evidence layer, the first label and the second label are obtained according to the threshold difference of the evidence layer, and the first label is corrected based on the second index to obtain the corrected label; Use the correction label and the second label to label the sub-target area to obtain the labeled area, represent the labeled area by polygonal vector elements, use the deep learning algorithm to train the polygonal vector elements and the labeled area to obtain the spatial data labeling model, and use the spatial data labeling model to output the visual labeling results.
2. A spatial data annotation method for delineating a prospecting target area according to claim 1, characterized in that: The method for collecting mining area data and performing preprocessing to obtain spatial data includes: Collect elevation remote sensing images, collect geochemical data with drilling depth information, collect geological structure data including underground structure extension depth information, use geographic information system technology to use elevation remote sensing images as spatial coordinate base, superimpose them on the coordinate space coordinate base in the form of point elements according to the coordinates of the geochemical data sampling points, superimpose geological structure data on the spatial coordinate base in the form of line elements or surface elements according to geographic coordinates, use bilinear interpolation to unify the resolution of the spatial coordinate base, and form spatial data.
3. The spatial data annotation method for delineating a prospecting target area according to claim 1, characterized in that: The method of constructing Thiessen polygons based on spatial data points and depth information, using Thiessen polygons to divide spatial data to obtain sub-target areas, and extracting geological structural characteristics and element characteristics of the sub-target area data to obtain initial values includes: Extract spatial data points and their corresponding depth information based on spatial data, import the spatial data points into a geographic information system, construct different spatial data point planes based on the depth information of different spatial data points, calculate the Euclidean distance between any two points in the spatial data point plane, use the Delaunay triangulation algorithm to generate perpendicular bisectors of the lines connecting adjacent spatial data points, intersect the perpendicular bisectors to form a Thiessen polygon boundary, and divide the spatial data point plane into different sub-target areas based on the Thiessen polygon boundary; The geological structure characteristics and element characteristics of the sub-target area data are extracted to obtain initial values, the fault intersection density is obtained based on the geological structure data of the sub-target area, the element content outliers and the proportion of outliers in the geochemical exploration data in the sub-target area are counted, the outliers are normalized as weight values, and the fault intersection density and the proportion of outliers are weighted summed to obtain the initial value.
4. The spatial data annotation method for delineating a prospecting target area according to claim 1, characterized in that: The method of randomly generating multiple groups of initial weight values for the initial values, performing linear approximation optimization on the initial weight values using the automatic regularized Gauss-Newton method, and iteratively calculating to obtain the first exponent includes: For the initial value, multiple groups of initial weight values are randomly generated according to the sub-target area data category. The sub-target area initial value and the sub-target area data are used to construct the first function through the ridge regression algorithm. Each group of initial weight values is used as the iterative starting point. The parameter weights of the first function are iteratively optimized by the automatic regularized Gauss-Newton method. The Jacobian matrix is used to adjust the parameter weights during the iteration until the first function converges to obtain the optimized parameter weights. The formula of the automatic regularized Gauss-Newton method is: in For the The weight vector for the iteration, For the The weight vector for the iteration, is the Jacobian matrix, is the transposed matrix, which is multiplied with the Jacobian matrix to form the approximate Hessian matrix. For the The regularization parameter for the iteration, For the In the iteration The residual between the initial value of the sub-target area and the predicted value of the first function, For the In the iteration The residual between the initial value of the sub-target area and the predicted value of the first function, is a very small positive number to prevent the denominator from being zero. is the identity matrix, For the The residual vector of the iteration, represents the total number of sub-target areas; The first exponent is obtained by multiplying the optimization parameter weight by the first function parameter value and then summing them.
5. The spatial data annotation method for delineating a prospecting target area according to claim 1, characterized in that: The method for obtaining the second index by performing nonlinear optimization on the initial value and the initial weight value based on the support vector machine includes: The initial value is used as the target variable, the initial weight value corresponding to the initial value is used as the feature variable, and the training data set is obtained based on the initial values of all sub-target areas of the spatial data; The support vector machine model was selected based on the kernel function. If the Pearson coefficient between the initial values was greater than 0.7, a linear kernel was selected. If the global Moran index between the initial values was greater than 0.5, a polynomial kernel function was selected, and if it was less than 0.3, a radial basis kernel function was selected. The support vector machine was trained using the training data set. The support vector machine parameters were optimized using the sequential minimum optimization algorithm by reducing the structural risk to obtain a trained support vector machine model. The initial weight values of the sub-target areas were input into the trained support vector machine model to obtain the output values, and the output values were normalized to obtain the second index.
6. The spatial data annotation method for delineating a prospecting target area according to claim 1, characterized in that: The method of using the first index as an input parameter of the weight of evidence method to construct an evidence layer and obtaining a first label and a second label according to a difference in thresholds of the evidence layer includes: Based on the sub-target area, the first index is used as the input parameter of the evidence weight method. The remote sensing images, geochemical data and geological structure data of the sub-target area are used as evidence factors. The conditional probability of the evidence factors under the mineralization and non-metallization states is statistically analyzed. The natural logarithm of the ratio of the probability of the evidence factors under the mineralization and non-metallization conditions is used to calculate the evidence factor weight. The geographic information system is used to perform spatial overlay calculation on the evidence factor weight layer to form an evidence layer. Frequency statistics are performed on the evidence layer values, and the potential threshold is set using the natural break point classification method based on the data distribution. The labels are divided according to the comparison results between the evidence layer values and the potential threshold. If the evidence layer value is greater than the potential threshold, it is classified as the first label; if it is less than the potential threshold, it is classified as the second label.
7. The spatial data annotation method for delineating a prospecting target area according to claim 1, characterized in that: The method of correcting the first label based on the second index to obtain a corrected label includes: The element content abnormal data of the sub-target area were interpolated using IDW to generate an element content grid map. The mean and standard deviation of the grid map were calculated. The lower limit of the abnormal value was obtained by adding twice the standard deviation to the element mean. The first correction parameter was obtained by dividing the single element content by the element mean using the contrast value method. The fault length per unit area of the sub-target area's geological data was calculated to extract the fault density. Linear density analysis was performed on the fold axis data to count the number of folds per unit distance and extract the fold frequency. The slope grid was generated using the DEM to calculate the slope standard deviation. The fault density, fold frequency, and slope standard deviation were normalized and then weighted and summed to obtain the second correction parameter. The first label is corrected using the second index and the first and second correction parameters through a correction formula, where the correction formula is: in To correct the label, is the first index, is the second index, is the second correction parameter, is the maximum value of the second correction parameter, is the first correction parameter, is the maximum value of the first correction parameter, is the first label, 、 、 、 is the coefficient of variation, is the subscript of the coefficient of variation.
8. The spatial data annotation method for delineating a prospecting target area according to claim 1, characterized in that: The method of using the correction label and the second label to label the sub-target area to obtain the labeled area, and representing the labeled area with a polygonal vector element, includes: The correction label and the second label are used to label the sub-target area to obtain the labeled area. The variance of the data at the boundary of the labeled area is calculated. Based on the variance, the complexity threshold is set as the sum of the 75% quantile of the variance and 1.5 times the interquartile range using the quartile method. The complexity threshold is used to divide the low-complexity labeled area into the high-complexity labeled area. The low-complexity labeled area is decomposed using the row-column convex decomposition formula, which is: in is the set of all convex polygons after decomposition, represent Any convex polygon in is an independent convex area unit after decomposition. Represents a convex polygon Any vertex in is used to limit the object range of the internal angle calculation. vertex The inner angle of From the vertex point to vector, From the vertex point to vector, It is a low-complexity annotation area. For The kth cutting operation is performed, It is a union operation; The low-complexity annotation area is divided into regular polygonal vector elements; for the high-complexity annotation area, the boundary coordinate points are extracted, and the coordinate points are connected in sequence to form a closed figure to obtain the polygonal vector elements.
9. The spatial data annotation method for delineating a prospecting target area according to claim 1, characterized in that: The method of using a deep learning algorithm to train polygonal vector elements and annotation areas to obtain a spatial data annotation model, and using the spatial data annotation model to output visual annotation results, includes: Normalize the geometric coordinates of polygonal vector features, perform one-hot encoding on the labeled area categories, generate a learning dataset based on the polygonal vector features and one-hot encoding, and divide the learning dataset into training set, validation set, and test set; The training set is input into the deep learning model. The model extracts the spatial features of polygonal vector elements through convolutional layers and pooling layers, makes preliminary predictions based on the spatial features, and uses the cross-entropy loss function to calculate the difference between the prediction result and the label after one-hot encoding. During training, the validation set is input into the model to verify the overfitting state of the model through the change of loss function value and accuracy. The difference is fed back to each layer of the model through the back propagation algorithm to adjust the model parameters so that the difference continues to narrow until the end of model training. The test set is input into the trained deep learning model, and the accuracy and recall rate are used to evaluate the model performance to obtain the spatial data annotation model, which is then used to output the visual annotation results.
10. A spatial data annotation system for delineating a prospecting target area, used to execute the spatial data annotation method for delineating a prospecting target area according to any one of claims 1 to 9, characterized in that: The system comprises: A data acquisition module is used to collect mining area data and pre-process it to obtain spatial data, which includes remote sensing images, geochemical data and geological structure data; The sub-target area generation module is used to construct Thiessen polygons based on spatial data points and depth information, use Thiessen polygons to divide spatial data to obtain sub-target areas, and extract geological structure characteristics and element characteristics of sub-target area data to obtain initial values; The index calculation module is used to randomly generate multiple sets of initial weight values for the initial values, use the automatic regularized Gauss-Newton method to perform linear approximation optimization on the initial weight values, iteratively calculate to obtain the first index, and perform nonlinear optimization on the initial values and initial weight values based on the support vector machine to obtain the second index; a label generation module, configured to construct an evidence layer using the first index as an input parameter of the weight of evidence method, obtain a first label and a second label according to a threshold difference of the evidence layer, and correct the first label based on the second index to obtain a corrected label; The label model generation module is used to use the correction label and the second label to label the sub-target area to obtain the labeled area, represent the labeled area by polygonal vector elements, use the deep learning algorithm to train the polygonal vector elements and the labeled area to obtain the spatial data labeling model, and use the spatial data labeling model to output the visual labeling results.
Citation Information
Patent Citations
Method for delineating prospecting target area through multi-source heterogeneous information based on machine learning
CN116611697A
Method and system for rapidly delineating tin ore prospecting target area
CN119165040A