Geochemical data reconstruction method and system based on deep learning

The deep learning model combines a variety of geological data to reconstruct geochemical data, which solves the problems of high data loss rate and reduced model accuracy in the existing technology, and achieves efficient and accurate data reconstruction effect.

CN120375178APending Publication Date: 2025-07-25CHINA UNIV OF GEOSCIENCES (BEIJING)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510275837.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art has insufficient sample density, high data missing rate, and difficulty in digging deep into potential physical associations between data in geochemical data reconstruction, resulting in a decrease in model accuracy, especially when the missing rate is high, it is difficult to effectively integrate geological structure and topographic landform information.

Method used

The geochemical data reconstruction method based on deep learning is adopted, and the spatial data set of the target area is obtained for preprocessing. The geochemical data reconstruction model is trained using the deep learning model, and the missing data is reconstructed by combining digital elevation model, fault data and lithologic data.

Benefits of technology

It realizes efficient and accurate filling and accurate reconstruction of geochemical data in the case of large-scale data missing, improving reconstruction performance, especially when the loss rate is high, it can still maintain high robustness and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375178A_ABST
    Figure CN120375178A_ABST
Patent Text Reader

Abstract

The invention provides a geochemical data reconstruction method and system based on deep learning, and belongs to the technical field of geological data processing.The geochemical data reconstruction method comprises the steps that a spatial data set of a target area is obtained, the spatial data set of the target area is preprocessed, the spatial data set comprises geochemical data, and the geochemical data is obtained; and at least one of digital elevation model data, fault data and lithology data; training a deep learning model based on the preprocessed spatial data set of the target area to obtain a geochemical data reconstruction model; and reconstructing the geochemical data of the target area by using the geochemical data reconstruction model, thereby realizing accurate filling and accurate reconstruction of the missing geochemical data of the target area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of geological data processing, and in particular to a geochemical data reconstruction method and system based on deep learning. Background Art

[0002] Geochemical data, as fundamental data describing the material composition of the earth's surface layer, plays a crucial role in fields such as mineral exploration, agriculture, and environmental protection. The formation of the distribution pattern of surface geochemical elements involves complex primary geological processes and secondary geological processes, with significant non-linear and stochastic characteristics. Especially in the dynamic process of element leaching-migration-enrichment, the interactive influence of geological processes at multiple spatio-temporal scales makes it a huge challenge to model the distribution of geochemical elements.

[0003] Currently, the acquisition of geochemical data mainly relies on sampling and analysis of surface materials such as soil and river sediments. However, due to factors such as the high cost of high-precision chemical analysis and the limited accessibility of sampling under complex terrain conditions, there are generally problems such as insufficient sample density and high data missing rate. Therefore, a series of methods are needed to reconstruct the missing data. For the data reconstruction requirement, the existing technologies mainly include the following methods: (1) Traditional geostatistical methods: Represented by inverse distance weighted interpolation (IDW), Kriging interpolation, and data interpolation empirical orthogonal function (DINEOF), a univariate interpolation model is constructed based on the principle of spatial autocorrelation. Taking ordinary Kriging method as an example, it realizes spatial prediction through semi-variogram modeling and performs well in small-scale missing data reconstruction. However, due to their simplified processing in physical processes, they can only capture the relatively superficial correlation between the missing data and the existing data, and fail to deeply explore the potential physical associations between the data. When the missing rate exceeds 30%, the model accuracy drops significantly, and it is difficult to effectively integrate auxiliary variable information such as geological structures and landforms, resulting in the loss of core spatial structure characteristics; (2) Machine learning enhancement methods: Improve the spatial prediction model by introducing algorithms such as random forest (RF) and support vector machine (SVM), but are limited by the shallow learning mechanism and have bottlenecks in modeling complex non-linear relationships.

[0004] Therefore, providing a method that can maintain the geochemical spatial distribution law and has high robustness to more accurately reconstruct data has become a key technical requirement for improving the prediction accuracy of mineral resources and optimizing the ecological environment assessment. Summary of the Invention

[0005] Aiming at the above problems, the present invention provides a geochemical data reconstruction method and system based on deep learning to accurately reconstruct the missing geochemical data in the target area.

[0006] First, the present invention provides a method for reconstructing geochemical data based on deep learning, including: Obtain a spatial data set of a target area, and preprocess the spatial data set of the target area, where the spatial data set includes geochemical data and at least one of digital elevation model data, fault data, and lithology data; Train a deep learning model based on the preprocessed spatial data set of the target area to obtain a geochemical data reconstruction model; Use the geochemical data reconstruction model to reconstruct the geochemical data of the target area.

[0007] Optionally, the spatial data set includes geochemical data, digital elevation model data, fault data, and lithology data.

[0008] Optionally, obtaining the spatial data set of the target area and preprocessing the spatial data set of the target area includes converting the geochemical data and at least one of digital elevation model data, fault data, and lithology data into raster data available for a deep learning model.

[0009] Optionally, converting the geochemical data and at least one of digital elevation model data, fault data, and lithology data into raster data available for a deep learning model includes: Level and remove steps from the geochemical data, and interpolate the leveled geochemical data to obtain geochemical raster data with a set raster accuracy; Resample the digital elevation model data to obtain digital elevation model raster data with a set raster accuracy; Use the circular buffer method to divide the fault data into buffers according to distance, and represent different buffers with different values to obtain fault raster data with a set raster accuracy; Represent different lithologies in the lithology data with different values to obtain lithology raster data with a set raster accuracy; Wherein, the geochemical raster data, digital elevation model raster data, fault raster data, and lithology raster data have the same set raster accuracy.

[0010] Optionally, leveling and removing steps from the geochemical data, and interpolating the leveled geochemical data to obtain geochemical raster data with a set raster accuracy includes interpolating the leveled geochemical data using the inverse distance weighted interpolation method.

[0011] Optionally, the fault data is divided into buffers according to distance using the circular buffer method, and different buffers are represented by different values to obtain fault grid data with a set grid accuracy, including using the circular buffer method to create a buffer every 500 m, successively creating buffers of 500 m, 1000 m, 1500 m, 2000 m, and the remaining part, and then successively setting them to corresponding values of 1-5.

[0012] Optionally, the set grid accuracy is 1 km × 1 km.

[0013] Optionally, training the deep learning model based on the preprocessed spatial dataset of the target area to obtain a geochemical data reconstruction model includes training the RFR deep learning model based on the preprocessed spatial dataset of the target area to obtain a geochemical data reconstruction model.

[0014] Optionally, training the RFR deep learning model based on the preprocessed spatial dataset of the target area to obtain a geochemical data reconstruction model includes: Setting a data missing area in the target area, performing a large-area continuous missing operation on the data in the spatial dataset of the data missing area of the target area including at least one of geochemical data, digital elevation model data, fault data, and lithology data according to a set ratio, using the spatial dataset after the large-area continuous missing operation to make samples, and then dividing the samples into a training set and a test set according to a set ratio; Inputting the training set into the RFR deep learning model for training to obtain the geochemical data reconstruction model.

[0015] Optionally, performing a large-area continuous missing operation on the data in the spatial dataset of the data missing area of the target area including at least one of geochemical data, digital elevation model data, fault data, and lithology data according to a set ratio includes performing a large-area continuous missing operation on the data in the spatial dataset of the data missing area of the target area including at least one of geochemical data, digital elevation model data, fault data, and lithology data according to a ratio of 25%, 50%, or 75%.

[0016] Optionally, using the spatial dataset after the large-area continuous missing operation to make samples, and then dividing the samples into a training set and a test set according to a set ratio includes using the spatial dataset after the large-area continuous missing operation to make samples, and then dividing the samples into a training set and a test set according to a ratio of 4:1.

[0017] Optionally, the spatial data set after large-area continuous deletion operation is used to make samples, and then the samples are divided into a training set and a test set according to a set ratio, including making samples by using a sliding window method, and the sample coverage rate is 90%.

[0018] Optionally, inputting the training set into the RFR deep learning model for training to obtain the geochemical data reconstruction model includes setting training parameters. The number of mini-batch images used in each iteration of training the model is 16, the image size is 128×128, and the number of iterations is 100,000.

[0019] Optionally, when the spatial data set includes geochemical data, digital elevation model data, fault data, and lithology data, inputting the training set into the RFR deep learning model for training to obtain the geochemical data reconstruction model includes fusing the digital elevation model data, fault data, and lithology data of each sample corresponding to the processed position into a numerical matrix with three channels, where each channel represents a variable.

[0020] Second, the present invention provides a geochemical data reconstruction system based on deep learning. The system includes a spatial data set acquisition and preprocessing module, a geochemical data reconstruction model construction module, and a geochemical data reconstruction module; Among them, the spatial data set acquisition and preprocessing module is used to acquire the spatial data set of the target area and preprocess the spatial data set of the target area. Among them, the spatial data set includes geochemical data, and at least one of digital elevation model data, fault data, and lithology data; The geochemical data reconstruction model construction module is used to train a deep learning model based on the preprocessed spatial data set of the target area to obtain a geochemical data reconstruction model; The geochemical data reconstruction module is used to reconstruct the geochemical data of the target area by using the geochemical data reconstruction model.

[0021] Optionally, the spatial data set acquisition and preprocessing module includes a raster data conversion unit for converting the acquired spatial data set into raster data available for the deep learning model.

[0022] Optionally, the geochemical data reconstruction model construction module includes a sample making unit and a model training unit. Among them, the sample making unit is used to make samples based on the preprocessed spatial data set and divide the made samples into a training set and a test set; the model training unit is used to train a deep learning model according to the training set to obtain the geochemical data reconstruction model.

[0023] Third, the present invention provides a computer electronic device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the geochemical data reconstruction method based on deep learning.

[0024] Fourth, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the geochemical data reconstruction method based on deep learning are implemented.

[0025] By providing a geochemical data reconstruction method and system based on deep learning, the present invention first obtains a spatial data set of a target area, preprocesses the spatial data set of the target area, and then trains a deep learning model based on the preprocessed spatial data set of the target area to obtain a geochemical data reconstruction model; finally, uses the geochemical data reconstruction model to reconstruct the geochemical data of the target area, makes full use of the powerful capabilities of the deep learning model in data analysis and fusion, learns the relationships and rules between geochemical data and other different types of physical data through the deep learning model, realizes efficient and accurate filling and accurate reconstruction of the missing geochemical data in the target area, especially in the case of a large data missing range, can also be well reconstructed, and compared with the traditional reconstruction methods that only interpolate geochemical data or use machine learning, its reconstruction performance has been significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is a schematic flowchart of the geochemical data reconstruction method based on deep learning according to an embodiment of the present invention.

[0027] Figure 2 is a schematic diagram of the deep learning model structure of the geochemical data reconstruction method based on deep learning according to an embodiment of the present invention.

[0028] Figure 3 is a schematic diagram of the spatial data of a certain area of the geochemical data reconstruction method based on deep learning according to an embodiment of the present invention.

[0029] Figure 4 is a schematic diagram of the data missing area of a certain area of the geochemical data reconstruction method based on deep learning according to an embodiment of the present invention.

[0030] Figure 5 is a schematic diagram of the reconstruction results of various models for geochemical data V at a 25% missing rate.

[0031] Figure 6 is a schematic diagram of the reconstruction results of various models for geochemical data V at a 50% missing rate.

[0032] Figure 7Schematic diagram of the reconstruction results of various models for geochemical data V at a 75% missing rate.

[0033] Figure 8 Schematic diagram of the difference between the reconstruction results of various models and the true values at a 25% missing rate.

[0034] Figure 9 Schematic diagram of the difference between the reconstruction results of various models and the true values at a 50% missing rate.

[0035] Figure 10 Schematic diagram of the difference between the reconstruction results of various models and the true values at a 75% missing rate. Detailed implementation manners

[0036] The following combines specific embodiments and appended Figure 1-10 to make a detailed description of the invention, so that those skilled in the art can more fully understand the purpose, features and effects of the invention.

[0037] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the art to which the present invention belongs. When the definition of a term in the present invention conflicts with the meaning commonly understood by those skilled in the art, the definition described in the present invention shall prevail.

[0038] The present invention provides a geochemical data reconstruction method and system based on deep learning, and makes corresponding improvements to the existing geochemical data reconstruction method, thereby greatly improving the reconstruction quality of geochemical data in the target area.

[0039] Embodiment 1 As a specific embodiment of the present invention, this embodiment provides a geochemical data reconstruction method and system based on deep learning. Referring to Figure 1 , the specific steps are as follows: S100. Obtain the spatial data set of the target area, and preprocess the spatial data set of the target area Specifically, in S100, first obtain the spatial data set of the target area, where the spatial data set at least includes geochemical data. Preferably, the spatial data set includes geochemical data, digital elevation model (DEM) data, fault data, and lithology data, and preprocess the data in the spatial data set to convert it into raster data available for the deep learning model. Among them, in this embodiment, the geochemical data is defined as the reconstruction variable, and the digital elevation model data, fault data, and lithology data are defined as the covariables.

[0040] S200. Train the deep learning model based on the preprocessed spatial data set of the target area to obtain a geochemical data reconstruction model Specifically, in S200, a data missing area is set in the target area, and large-area continuous missing operations are performed on the data in the spatial data set of the data missing area in the target area that at least includes geochemical data according to a set ratio. The spatial data set after the large-area continuous missing operations is used to make samples, and then the samples are divided into a training set and a test set according to a set ratio. The area of the data missing area is smaller than that of the target area.

[0041] Among them, performing large-area continuous missing operations on the data in the spatial data set of the data missing area in the target area that at least includes geochemical data according to a set ratio includes: Performing large-area continuous missing operations on the geochemical data in the spatial data set of the data missing area in the target area according to a set ratio. The spatial data set after the large-area continuous missing operations on the geochemical data is used to make samples, and then the samples are divided into a training set and a test set according to a set ratio. The covariate digital elevation model data, fault data, and lithology data are not subjected to missing operations; Or performing large-area continuous missing operations on the geochemical data and at least one covariate data among the digital elevation model data, fault data, and lithology data in the spatial data set of the data missing area in the target area. The spatial data set after the large-area continuous missing operations on the geochemical data and at least one covariate data among the digital elevation model data, fault data, and lithology data is used to make samples, and then the samples are divided into a training set and a test set according to a set ratio; Among them, the combinations for performing large-area continuous missing operations can be geochemical data and digital elevation model data, geochemical data and fault data, geochemical data and lithology data, geochemical data and digital elevation model data and fault data, geochemical data and digital elevation model data and lithology data, geochemical data and fault data and lithology data, geochemical data and digital elevation model data, fault data, and lithology data.

[0042] The training set is input into the deep learning model for training to obtain the geochemical data reconstruction model, and the test set is used to test the geochemical data reconstruction model.

[0043] S300. Using the geochemical data reconstruction model, reconstruct the geochemical data in the target area.

[0044] Embodiment 2 This embodiment takes the reconstruction of geochemical data in a certain area as an example to further elaborate on the above method in detail and compare it with other data reconstruction methods in the prior art. This area is a porphyry copper deposit enrichment area with complex and diverse mineralization processes.

[0045] Referring to the geochemical data reconstruction method based on deep learning in Embodiment 1, in this embodiment, first, a spatial data set including geochemical data, digital elevation model data, fault data, and lithology data in this area is obtained. Among them, the geochemical data is point data containing multiple elements, with a scale of 1:200,000; the digital elevation model data is ASTER GDEM 30M resolution digital elevation model data from the Geospatial Data Cloud. In this embodiment, the vanadium (V) element, which has a relatively large dynamic range, obvious texture, and is relatively easy to extract features in this area, is taken as an example for further processing.

[0046] During the rasterization preprocessing, different treatments are performed on different data in the spatial data set. Specifically, for the geochemical data, leveling is performed to remove the existing step problem, and the leveled geochemical data is interpolated into raster data with a resolution of 1 km × 1 km using the interpolation method. Among them, the inverse distance weighting (IDW) interpolation method is preferably used. In its feasible embodiments, other interpolation methods can also be used to interpolate the geochemical data; for the digital elevation model data, resampling is performed to obtain 1 km × 1 km raster data; for the fault data, the method of circular buffer is used. A buffer is established every 500 m, and buffers of 500 m, 1000 m, 1500 m, 2000 m, and the remaining part are established in sequence, and then they are set to corresponding values of 1-5 in sequence and converted into 1 km × 1 km raster data; for the lithology data, in this area, the strata are divided into eight categories according to lithology: carbonate rock, clastic rock, intermediate-basic volcanic rock, basic and ultrabasic intrusive rock, metamorphic rock, Quaternary loose sediment, Proterozoic basement, and intermediate-acid intrusive rock, and they are set to corresponding values of 1-8 in sequence and converted into 1 km × 1 km raster data.

[0047] After preprocessing the spatial data set, four 1 km × 1 km tif raster data with a specification of 811 × 602 are obtained.

[0048] Referring to Figure 3 , where 3(a) is geochemical data, 3(b) is digital elevation model data, 3(c) is fault data, and 3(d) is lithology data.

[0049] After rasterizing the spatial data set in this embodiment, the data in the data missing area of the spatial data set including at least the geochemical V element data is continuously missing in a large area according to a set ratio, and the spatial data set after the large-area continuous missing operation is used to make samples. Then, the samples are divided into a training set and a test set according to a set ratio to train the deep learning model. Preferably, in this embodiment, a recursive feature inference network (RFR) model is used as the deep learning model.

[0050] Refer to Figure 2 , the RFR model includes a region recognition module, a feature inference module, and a feature merging module. The region recognition module is responsible for accurately identifying the regions to be inferred in each recursion, providing a clear target for subsequent feature inference. The feature inference module then uses deep learning techniques to efficiently infer the image content within the identified regions based on the recognized regions. Finally, the feature merging module merges all the intermediate feature maps generated during the inference process to generate a feature map with a fixed number of channels, providing strong support for the final image drawing.

[0051] Partial convolution is adopted in the region recognition module: Let represent the feature map output by the partial convolution layer, where represents the feature value at the position ( x, y ) and the z th channel. represents the z th convolutional kernel in this layer. For any position ( x, y ), and represent the input feature patch and the input mask patch centered at this position (with the same size as the convolutional kernel), respectively. On this basis, the feature map calculated by the partial convolution layer can be expressed as:

[0052] Similarly, the new mask value at the generated position ( i, j ) can be expressed as:

[0053] Based on the above equations, a gradually shrinking new mask can be obtained after each partial convolution layer. After being processed by the partial convolution layer, the feature map will be further processed by the normalization layer and the activation function, and then passed to the feature inference module.

[0054] In the feature merging module, during the process of generating the output feature map, values are only extracted from those feature maps whose corresponding positions have been effectively filled for calculation.

[0055] The value at the output feature map is defined as:

[0056] represents the th feature map generated by the feature inference module. F represents the value at the position ( x, y, z ) in the feature map F ). Represents a feature map of a binary mask that is used to mark which positions have been effectively filled, where N is the number of feature maps. In this way, any number of feature maps can be merged, making it possible for the RFR to fill a larger area.

[0057] In the RFR model of this embodiment, for image generation learning, the perceptual loss and style loss of the pre-trained and fixed VGG-16 are used.

[0058] The perceptual loss function is as follows:

[0059] The style loss function is as follows:

[0060]

[0061] φ pooli Represents the feature map of the i th pooling layer in the fixed VGG-16, H i , W i , C i respectively represent the height, weight, and channel size of the i th feature map.

[0062] In addition, L valid and L hole are also used in the model, that is, the valid loss and the hole loss functions, to calculate the L 1 difference between the unoccluded area and the occluded area.

[0063] The total loss function is:

[0064] Specifically, without introducing covariates, a large-area continuous missing operation is performed on the geochemical V element data in the data missing area of the target area, where the large-area continuous missing operation can be performed artificially. Refer to Figure 4, the area of the data missing area determined in the target area is 128 km × 128 km, and the missing ratios are 25%, 50% and 75% of the data missing area, representing the cases of 25% data missing, 50% data missing and 75% data missing in the data missing area respectively, and the covariates are not missing. The geochemical data after the large-area continuous missing operation is used to make samples. Preferably, the method of using a sliding window is used to make samples, the sample coverage rate is 90%, and the samples containing nodata are removed. Finally, 124 samples of 128×128 are obtained, and the ratio of the training set to the test set is 4:1.

[0065] The number of mini-batch images used in each iteration of the training model is 16, the image size is 128×128, and the number of iterations is 100,000. When starting training, the software environment is Python 3.9, Pytorch 2.3.0, Windows 10 Microsoft operating system, and the hardware environment is CPU Intel 8358P-32C 256GB and GPU 40GB Nvidia-A800.

[0066] Furthermore, when training the RFR model, in order to ensure the randomness of training, a completely random method is adopted to generate the mask. The initial form of the mask is a numerical matrix that is exactly the same as the layer size, and all elements are initialized to 1. In the process of generating the mask, first randomly select a position as the starting point, and then, starting from this starting point, randomly select a direction to move up, down, left or right at each step; whenever moving to a new position, the mask value at the current position is set to 0. In order to control the overall missing rate, it is monitored by calculating the sum of all elements in the mask. When the missing rate reaches the predetermined level, the process terminates. Finally, the generated mask is multiplied by the geochemical data to obtain the geochemical data containing missing values. The above method not only ensures the randomness of the experiment but also ensures the credibility of the results.

[0067] Furthermore, in order to enable the deep learning model to use more data for training, any one of the geochemical V element data and the digital elevation model data, fault data, and lithology data in the covariates is combined respectively for large-area continuous missing operation. Refer to Figure 4, the area of the data missing area determined in the target area is 128 km × 128 km, and the missing ratios are 25%, 50% and 75% of the data missing area respectively. The geochemical data and covariate data after large-area continuous missing operations are used to make samples. Preferably, the sliding window method is used to make samples, the sample coverage rate is 90%, and the samples containing nodata are removed. Finally, 124 samples of 128×128 are obtained, and the ratio of the training set to the test set is 4:1.

[0068] When training the deep learning model, modify the input of the model. When the sample is input into the model, the single covariate and the reconstructed variable with the same area are input one by one. Finally, 3 pre-trained models are trained respectively to obtain the geochemical data reconstruction models based on single covariates, namely the geochemical data reconstruction model based on the combination of geochemical data and digital elevation model data, the geochemical data reconstruction model based on the combination of geochemical data and fault data, and the geochemical data reconstruction model based on the combination of geochemical data and lithology data.

[0069] The number of mini-batch images used in each iteration of the training model is 16, the image size is 128×128, and the number of iterations is 100000. When starting the training, the software environment is Python 3.9, Pytorch 2.3.0, Windows 10 Microsoft operating system, and the hardware environment is CPU Intel 8358P-32C 256GB and GPU 40GB Nvidia-A800.

[0070] Use the geochemical data reconstruction model based on the combination of geochemical data and digital elevation model data, the geochemical data reconstruction model based on the combination of geochemical data and fault data, and the geochemical data reconstruction model based on the combination of geochemical data and lithology data to reconstruct the geochemical data of the target area respectively.

[0071] The geochemical data reconstruction model constructed without introducing covariates can successfully reconstruct the missing values of geochemical data. It deeply captures the global spatial features of the data through convolution operations. Compared with the traditional processing methods that mainly rely on local information, the geochemical data reconstruction model in this embodiment has achieved a significant improvement in reconstruction accuracy. However, since the learning ability of the RFR model algorithm is limited by the existing data set, and there is a complex physical mechanism behind the changes in geochemical data, which is the interaction of various physical factors that jointly drive the evolution of geochemical variables, the geochemical data reconstruction model trained by fusing covariate data is undoubtedly crucial for the accurate reconstruction of missing values of geochemical data.

[0072] When the RFR algorithm without introducing covariates processes data, it uses 0 as the initial filling value for the masked area, and thus cannot obtain the specific impact of covariates on model performance. There is a large deviation between this artificially masked data filled directly with 0 and the real ground observation values. This deviation increases the learning difficulty of the neural network and has an adverse impact on the final reconstruction result. The geochemical data reconstruction model based on a single covariate helps to obtain a more accurate reconstruction result. The geochemical data reconstruction model based on a single covariate introduces digital elevation model data, fault data, and lithology data as covariates for the missing area respectively. In this way, the model can use more data for learning and make full use of the powerful capabilities of the deep learning model in data analysis and fusion to achieve high-quality reconstruction of the missing data. It should be noted that during the construction of the geochemical data reconstruction model based on a single covariate, the processing of the mask is the same as described above.

[0073] Furthermore, to make the obtained geochemical data reconstruction model more accurate, the geochemical V element data is fused with the digital elevation model data, fault data, and lithology data in the covariates, and a large-area continuous missing operation is performed. Refer to Figure 4 , in the target area, the area of the data missing area determined is 128 km × 128 km, and the missing ratios are 25%, 50%, and 75% of the data missing area respectively. The geochemical data and covariate data after the large-area continuous missing operation are used to make samples. Preferably, the sliding window method is used to make samples, the sample coverage rate is 90%, and the samples containing nodata are removed. Finally, 124 samples of 128 × 128 are obtained, and the ratio of the training set to the test set is 4:1.

[0074] When training the deep learning model, modify the input of the model. When the samples are input into the model, the three covariates with the same area and the reconstruction variables are input one by one. Finally, train the model to obtain a geochemical data reconstruction model based on the combination of geochemical data, digital elevation model data, fault data, and lithology data.

[0075] The number of mini-batch images used in each iteration of training the model is 16, the image size is 128 × 128, and the number of iterations is 100000. When starting training, the software environment is Python 3.9, Pytorch 2.3.0, Windows 10 Microsoft operating system, and the hardware environment is CPU Intel 8358P-32C 256GB and GPU 40GB Nvidia-A800.

[0076] The geochemical data reconstruction model based on the combination of geochemical data, digital elevation model data, fault data, and lithology data reconstructs the geochemical data of the target area.

[0077] Preferably, in PyTorch, the digital elevation model data, fault data, and lithology data of each sample corresponding to the processed positions are fused into a numerical matrix with three channels using the cat function. In this matrix, each channel represents a covariate. The multi-covariate numerical matrix in this embodiment forms a more diverse and detailed data structure through the superposition of digital elevation model data, fault data, and lithology data.

[0078] Considering that the regional geochemical data is affected by multiple covariates rather than a single covariate and is in a complex earth system where the covariates have direct or indirect relationships, it is not comprehensive enough to only consider a single variable. Therefore, the effect of reconstructing the data in the area to be filled based on multiple covariates will be more excellent. By learning the rules between multiple covariates and the dependent variable through the deep learning model RFR model, the reconstruction performance of the model can be improved.

[0079] To better demonstrate the effect of this embodiment, the geochemical data of this area was reconstructed using traditional geostatistics and machine learning. Among them, the traditional geostatistical methods include Kriging and Co-kriging, and machine learning includes Random Forest (RF), Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost) models. Co-kriging interpolation combines digital elevation data, fault data, and lithology data as covariates, and machine learning also combines digital elevation data, fault data, and lithology data as covariates, with geochemical data as the prediction target.

[0080] When training the three machine learning models, a pixel-based method is adopted. Specifically, the digital elevation model data, fault data, and lithology data are used as covariates, while the geochemical data is set as the reconstruction variable. In data division, the data outside the area to be filled is used as the training set and validation set so that the model can learn the internal rules and characteristics of the data; while the data within the area to be filled is used as the test set to evaluate the prediction performance of the model. To find the optimal parameter configuration for each machine learning method, a grid search technique is adopted to find the optimal model parameters by exhaustively listing the specified parameter combinations.

[0081] To compare the performance of different geochemical reconstruction methods, this embodiment uses four performance evaluation indicators: Mean Absolute Error (MAE), Root Mean Square Error (RMSE), Coefficient of Determination (R²), and Concordance Correlation Coefficient (CCC). Tables 1 - 3 and Figures 5-7Shows the performance of reconstructing geochemical data by five different methods (including Kriging and Co-kriging methods based on traditional geostatistics, SVM, RF, and XGBoost methods based on machine learning, and RFR model based on deep learning) at three different missing rates of 25%, 50%, and 75%. Among them, RFR-0 is the geochemical data reconstruction model constructed without introducing covariates, RFR-dem is the geochemical data reconstruction model constructed by introducing digital elevation model data, RFR-fault is the geochemical data reconstruction model constructed by introducing fault data, RFR-litho is the geochemical data reconstruction model constructed by introducing lithology data, RFR-all is the geochemical data reconstruction model constructed by introducing digital elevation model data, fault data, and lithology data, and gt is the original data.

[0082] Table 1 Comparison of evaluation indexes of reconstruction results of various models at 25% missing rate

[0083] Table 2 Comparison of evaluation indexes of reconstruction results of various models at 50% missing rate

[0084] Table 3 Comparison of evaluation indexes of reconstruction results of various models at 75% missing rate

[0085] Generally speaking, the effect of the Kriging interpolation method is the most inferior. Visually, its reconstruction result appears rather blurred and lacks the original texture features, thus losing valuable spatial information. This is mainly because the Kriging method relies on the average value of surrounding data to predict missing values, and in the case of large-area continuous data missing, the remaining available data is not sufficient to support the Kriging model to accurately predict the data in the missing area. In contrast, the RFR model based on deep learning performs the best in data reconstruction. Whether from the visual texture feature reconstruction effect or from the evaluation indexes at the three missing rates, the RFR model has the best effect when fusing multiple variables, followed by the univariate case, while the case based on 0 variables is relatively poor. In addition, the reconstruction effects of the three machine learning models (SVM, RF, and XGBoost) are between the geostatistical method and the deep learning method.

[0086] In the results of model reconstruction based on traditional geostatistics, due to the limitations of its principle, the Kriging interpolation method exhibits relatively single characteristics under three different missing states, that is, simply smoothing the surrounding data to the central missing area. When other covariates are introduced for Co-Kriging interpolation, although in terms of various evaluation indicators, its effect is improved compared with the Kriging method, and some texture features similar to the real image are visually presented, generally speaking, the reconstruction result still fails to reach the expected ideal state.

[0087] In the results of model reconstruction based on machine learning, since the three machine learning models adopted in the present invention perform statistical analysis and model fitting based on individual pixels, and the prediction results are also for individual pixels, the reconstruction results lack continuous texture features visually. Specifically, in the cases of 25% and 50% missing rates, the two ensemble learning models - RF and XGBoost show better effects compared with SVM. A significant defect of the SVM model is its high computational complexity and large memory consumption. When dealing with large-scale data sets, the required computing resources and time will increase sharply, so it is not very suitable for large-scale data processing. If the computational cost of SVM is attempted to be reduced when dealing with large-scale data sets, its accuracy often decreases. And in the case of a 75% missing rate, the missing data may have a greater impact on the fitting effect of the RF model, thus reducing the prediction accuracy of RF.

[0088] In the results of model reconstruction based on deep learning models, the RFR model fully learns all the information of the entire spatial data distribution and uses these comprehensive valid values to reconstruct the data in the missing area, rather than simply relying on the data around the area to be reconstructed. From the visual effect of the reconstruction results, the RFR model can generate smooth and fluent textures in different situations, which is significantly different from the single numerical results generated by geostatistical methods and the non-smooth phenomenon of pixel-by-pixel numerical values in machine learning methods. Through the evaluation of the coefficient of determination R², in all cases, the deep learning-based RFR model shows higher accuracy than the other two major types of models. In various application scenarios of the RFR model, when the missing rate is 25%, the model based on all variables has the best effect, with an R² value as high as 0.6854. Followed by the model based on lithology single variable, then in turn are the models based on fault single variable, DEM single variable, and 0 value. When the missing rate increases to 50%, although the model based on lithology single variable has the maximum value in the coefficient of determination, the model based on all variables has a very small gap with it, which may be due to the accidental influence of the model on a specific test set. And when the missing rate is as high as 75%, the effect of the model is the same as that when the missing rate is 25%.

[0089] Figure 8- The figure shows the differences between the reconstruction results of various models and the true values. The difference value = reconstruction result - true value. It can be clearly seen from the difference figure that the deep learning model is superior in terms of effect compared to the other several models, and its difference performance is the smallest. Specifically, when the missing rate is 25%, the difference generated by the RFR model based on 0 values is the most significant. The RFR model based on univariate corrects this error to a certain extent, although there is still a certain deviation. Further, the RFR model based on multivariate optimizes and corrects the error again. This series of comparisons fully demonstrates that the RFR model incorporating multivariate shows more significant effectiveness in geochemical data reconstruction compared to other models.

[0090] The geochemical data reconstruction method based on deep learning in this embodiment is trained by using the RFR model in deep learning to obtain a model with higher accuracy, stronger generalization ability, greater flexibility, faster speed, and the ability to fill a larger range, efficiently filling the missing geochemical data in the target area. Applying it in ore-forming prediction has important significance for the reconstruction of geochemical missing values. In addition, it can not only play a crucial role in ore-forming prediction, but also have a profound impact on making decisions regarding interactions with the natural environment.

[0091] Embodiment III As another specific embodiment of the present invention, this embodiment provides a geochemical data reconstruction system based on deep learning, including a spatial data set acquisition and preprocessing module, a geochemical data reconstruction model construction module, and a geochemical data reconstruction module; Among them, the spatial data set acquisition and preprocessing module is used to acquire the spatial data set of the target area and preprocess the spatial data set of the target area. Among them, the spatial data set includes geochemical data and at least one of digital elevation model data, fault data, and lithology data; The geochemical data reconstruction model construction module is used to train the deep learning model based on the preprocessed spatial data set of the target area to obtain a geochemical data reconstruction model; The geochemical data reconstruction module is used to reconstruct the geochemical data of the target area by using the geochemical data reconstruction model.

[0092] Further, the spatial data set acquisition and preprocessing module includes a raster data conversion unit for converting the acquired spatial data set into raster data available for the deep learning model.

[0093] The geochemical data reconstruction model construction module includes a sample production unit and a model training unit. Among them, the sample production unit is used to produce samples based on the preprocessed spatial data set, and divide the produced samples into a training set and a test set; the model training unit is used to train a deep learning model according to the training set to obtain the geochemical data reconstruction model.

[0094] Embodiment 4 As an embodiment of the present invention, this embodiment provides a computer electronic device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the geochemical data reconstruction method based on deep learning described in Embodiment 1.

[0095] Embodiment 5 As another embodiment of the present invention, this embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the geochemical data reconstruction method based on deep learning described in Embodiment 1.

[0096] The geochemical data reconstruction method and system provided by the present invention break through the traditional single-element analysis mode that only relies on geochemistry, integrate multi-dimensional spatial information of geochemical data, digital elevation model, fault data, and lithology data, construct a geological-topographic-structure collaborative analysis framework, realize implicit feature association through deep learning, effectively capture the complex non-linear relationships between element migration and enrichment, topographic undulation, fault zone distribution, lithology combination, etc., realize the effective processing of non-stationary spatial data and multi-variable collaborative inversion, reduce prediction errors, improve the reconstruction accuracy of geochemical data, and have important application value in the fields of mineral resource prediction, environmental geochemical evaluation, etc.

[0097] The above is only a preferred embodiment of the present invention, and it is not a limitation of the present invention in any other form. Any modification or equivalent change made according to the technical essence of the present invention still belongs to the scope of protection required by the present invention.

Claims

1. A geochemical data reconstruction method based on deep learning, characterized in that, The method includes: Obtaining a spatial data set of a target area and preprocessing the spatial data set of the target area, where the spatial data set includes geochemical data and at least one of digital elevation model data, fault data, and lithology data; Training a deep learning model based on the preprocessed spatial data set of the target area to obtain a geochemical data reconstruction model; Using the geochemical data reconstruction model to reconstruct the geochemical data of the target area.

2. The geochemical data reconstruction method based on deep learning according to claim 1, wherein The spatial data set includes geochemical data, digital elevation model data, fault data, and lithology data.

3. The geochemical data reconstruction method based on deep learning according to claim 1, characterized in that, The obtaining of the spatial data set of the target area and the preprocessing of the spatial data set of the target area include converting the geochemical data and at least one of digital elevation model data, fault data, and lithology data into raster data available for the deep learning model.

4. The geochemical data reconstruction method based on deep learning according to claim 3, wherein The converting of the geochemical data and at least one of digital elevation model data, fault data, and lithology data into raster data available for the deep learning model includes: Leveling and removing steps from the geochemical data, and interpolating the leveled geochemical data to obtain geochemical raster data with a set raster accuracy; Performing resampling on the digital elevation model data to obtain digital elevation model raster data with a set raster accuracy; Dividing the fault data into buffers according to distance using the circular buffer method, and representing different buffers with different values to obtain fault raster data with a set raster accuracy; Representing different lithologies in the lithology data with different values to obtain lithology raster data with a set raster accuracy; Wherein, the set raster accuracies of the geochemical raster data, digital elevation model raster data, fault raster data, and lithology raster data are the same.

5. The geochemical data reconstruction method based on deep learning according to claim 1, wherein The training of the deep learning model based on the preprocessed spatial data set of the target area to obtain a geochemical data reconstruction model includes training an RFR deep learning model based on the preprocessed spatial data set of the target area to obtain a geochemical data reconstruction model.

6. The geochemical data reconstruction method based on deep learning according to claim 5, wherein, The training of the RFR deep learning model based on the preprocessed spatial data set of the target area to obtain a geochemical data reconstruction model includes: Setting a data missing area in the target area, performing a large-area continuous missing operation on the data in the spatial data set of the data missing area of the target area including geochemical data and at least one of digital elevation model data, fault data, and lithology data according to a set ratio, using the spatial data set after the large-area continuous missing operation to make samples, and then dividing the samples into a training set and a test set according to a set ratio; Inputting the training set into the RFR deep learning model for training to obtain the geochemical data reconstruction model.

7. A geochemical data reconstruction system based on deep learning, characterized in that, The system includes a spatial data set acquisition and preprocessing module, a geochemical data reconstruction model construction module, and a geochemical data reconstruction module; Among them, the spatial data set acquisition and preprocessing module is used to acquire the spatial data set of the target area and preprocess the spatial data set of the target area. Among them, the spatial data set includes geochemical data and at least one of digital elevation model data, fault data, and lithology data; The geochemical data reconstruction model construction module is used to train a deep learning model based on the preprocessed spatial data set of the target area to obtain a geochemical data reconstruction model; The geochemical data reconstruction module is used to reconstruct the geochemical data of the target area by using the geochemical data reconstruction model.

8. The geochemical data reconstruction system based on deep learning according to claim 7, characterized in that, The spatial data set acquisition and preprocessing module includes a raster data conversion unit for converting the acquired spatial data set into raster data usable by a deep learning model.

9. A computer electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1-6 are implemented.