High-precision provenance identification method and system based on multi-mineral tracing and machine learning

By combining multi-mineral tracing with machine learning, the problem of multiple solutions in source area identification under complex geological conditions is solved. This method achieves high-precision quantification of source area contribution ratios and efficient data processing, and is applicable to oil and gas exploration and mineral resource surveys.

CN122131420APending Publication Date: 2026-06-02SHANDONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG UNIV OF SCI & TECH
Filing Date
2026-02-04
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies suffer from severe ambiguity when identifying source areas in complex geological environments, making it difficult to accurately distinguish and quantify the contribution ratio of different source areas. Furthermore, they are inefficient and costly in processing large-scale, multi-dimensional data.

Method used

A method combining multi-mineral tracing and machine learning was adopted. Multidimensional features were extracted from zircon U-Pb age and geochemical data. The random forest algorithm was used to train the model, optimize the feature set, and quantify the contribution ratio of the source region.

Benefits of technology

It achieves high-precision source area identification, reduces manual intervention, improves data processing efficiency, reduces costs, and can accurately quantify the contribution ratio of source areas, making it suitable for oil and gas exploration and mineral resource surveys in complex geological environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122131420A_ABST
    Figure CN122131420A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of geological exploration and resource exploration technology, and discloses a high-precision source area identification method and system based on multi-mineral tracing and machine learning. This invention collects and preprocesses zircon U-Pb age and geochemical data to construct a training set; extracts age distribution features through kernel density estimation and combines them with geochemical features to construct a multi-dimensional feature vector; eliminates redundant features through correlation analysis to construct a sample-feature matrix; trains a classification model using a random forest algorithm and optimizes hyperparameters through cross-validation and grid search; inputs the features of the sample to be tested into the optimization model, and outputs the probability of each source area as a contribution ratio, achieving high-precision quantitative identification of complex sources. This invention employs machine learning technology, which can effectively handle high-dimensional and complex geological datasets, overcoming the limitations of traditional methods in handling large-scale, multi-dimensional data. This invention addresses the challenges of high-dimensional data analysis through feature selection and model optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of geological exploration and resource exploration technology, and in particular relates to a high-precision source area identification method and system based on multi-mineral tracing and machine learning. Background Technology

[0002] Provenance zone identification is a crucial step in sedimentary basin analysis and oil and gas exploration. Existing technologies largely focus on single mineral tracing techniques (such as zircon U-Pb dating) or simple geochemical index analysis, but they lack sufficient expertise in identifying mixed sources in complex geological environments and quantifying their contributions with high precision. When existing methods directly apply single-index analysis, they neglect the comprehensive information from multiple sources (such as heavy minerals, light minerals, and geochemical data) and complex geological backgrounds, leading to overlapping provenance zone identification results and significant ambiguity, making it difficult to accurately distinguish and quantify different provenance zones. Existing methods lack the ability to dynamically evaluate large-scale, multi-dimensional data. Traditional models are often based on manual operations or simple statistical tools, without incorporating machine learning algorithms for feature extraction and model optimization. Therefore, they cannot accurately quantify the contribution ratio of different provenance zones to sediments and fail to meet the demands of efficient analysis. Acquiring and processing provenance analysis data for large-scale oil and gas exploration tasks faces multiple challenges, such as high data dimensionality and large processing volumes. Existing technologies often rely on manual operation for data analysis and use a single parameter (such as the concentration of a single element) to simplify calculations, resulting in a cumbersome and costly analysis process that is particularly difficult to meet the needs of speed and low cost.

[0003] Therefore, this solution solves and implements an efficient source analysis method that can integrate multi-source information, automatically extract comprehensive features, and intelligently quantify contribution ratios.

[0004] Based on the above analysis, the problems and shortcomings of the existing technology are as follows: (1) Foreign countries applied machine learning + geology earlier, but mostly focused on seismic facies identification or well logging curve prediction. There is very little application in the sub-field of detrital zircon U-Pb dating + Hf isotope + machine learning.

[0005] (2) Traditional domestic methods can only provide qualitative conclusions about the main source areas, while this invention can output numerical contribution values. This leap from qualitative to quantitative analysis is what industrial applications urgently need and is also the key to its advancement. Summary of the Invention

[0006] To overcome the problems existing in related technologies, the present invention discloses a method and system for high-precision source region identification based on multi-mineral tracing and machine learning. By converting multi-mineral tracing data into multi-dimensional features and utilizing machine learning models for pattern recognition and contribution quantification, high-precision differentiation of complex source regions can be achieved. The technical solution is as follows: This invention is implemented as follows: a high-precision source region identification method based on multi-mineral tracing and machine learning, comprising the following steps: S1. Data Collection and Preprocessing: Collect sample data from the study area and potential provenance areas. The sample data should include at least zircon U-Pb age data and geochemical data. Clean and standardize the data, and assign provenance labels to the sample data from known provenance areas to form a training set. S2. Multidimensional feature extraction: For each sample in the training set, the age distribution characteristics are obtained by estimating the kernel density of zircon U-Pb age data and analyzing the age probability density curve. Geochemical features are extracted from the geochemical data and together they form the initial feature vector of the sample. S3. Feature selection and matrix construction: Calculate the correlation between each initial feature, remove redundant features, and obtain the optimized feature set; merge the optimized feature sets of all samples to construct the sample-feature matrix; S4. Machine Learning Model Training and Optimization: The sample-feature matrix is ​​randomly divided into training and test sets. The random forest algorithm is used to train the source region classification model using the training set. Cross-validation and grid search strategies are used to optimize the model hyperparameters to obtain the best trained random forest model. S5. Source Region Identification and Contribution Quantification: Input the feature data of the sample in the study area to be identified into the trained model, and output the probability of the sample belonging to each potential source region to characterize the contribution ratio of each source region.

[0007] In step S1, the data is standardized to a Z-score, expressed as: ; In the formula, The distance of the original dataset from the overall data mean The standard deviation multiple, For a single data point in the dataset, The population mean of the original dataset. This represents the overall standard deviation of the original dataset.

[0008] In step S2, the age distribution features extracted from the zircon U-Pb age data include at least one of the main age peaks, peak area ratios, and age spans.

[0009] In step S2, the age distribution characteristics are obtained by estimating the kernel density of zircon U-Pb age data and analyzing the age probability density curve, including: Discrete zircon U-Pb age data are converted into continuous probability density curves using a kernel density estimation function. The main age peaks are automatically identified and extracted from the probability density curves, and the integration interval is defined. The probability density function is numerically integrated within the integration interval to obtain the peak area ratio of the main age peaks.

[0010] For the zircon U-Pb age probability density curve, the algorithm first traverses every point on the curve. Ensure every point Does it exist and have meaning? Then each point... The points on both sides and Comparison, This is a reasonable numerical radius within the window value. If one of them... If a point is greater than all points within the window, then it is the maximum value; otherwise, it is the minimum value.

[0011] In geology, detrital zircon U-Pb age spectroscopy is one method to reflect source contribution. The more zircon grains within a given age range, the more source region was influenced by basin margins or tectonic movements during geological history, thus receiving a greater source flux. This invention proposes that a larger peak area ratio indicates a greater number of zircon grains, and the probability area of ​​the major age peak is calculated based on the number of zircon grains.

[0012] The major age peak corresponds to the largest number of zircon grains, meaning the probability area of ​​the major age peak is approximately equal to the relative contribution of the source region material in the sample. The probability area of ​​the major age peak is numerically equal to its area proportion, meaning the peak area proportion is approximately equal to the relative contribution of the source region material in the sample.

[0013] In step S2, the geochemical data includes zircon data. Values ​​and / or Values, geochemical characteristics including average value, Isotope dispersion and / or average value; in, Isotope dispersion The calculation formula is: ; In the formula, for Isotope dispersion For the first Zircon value, For this sample The arithmetic mean of the values This represents the number of zircon grains.

[0014] The full name is The standard deviation of isotopic age is calculated for the same sample in this invention. The degree of dispersion in isotopic crustal ages. Low Representing zircon The ages are very concentrated, with a very small range of variation, indicating that the population is relatively homogeneous. Representing zircon The age distribution is highly dispersed and varies widely, indicating heterogeneity. The main reason is... Isotopes are very stable in zircon; essentially, after crystallization, their... The value will not change with the span of age. If the zircon crystals originate from a relatively homogeneous source region, then it indicates... The dispersion of isotopic crustal ages is very low. In theory, it would be very low, and therefore based on Value calculated The value will be very low. Conversely, if the zircon crystals originate from a more dispersed source region, it indicates... The dispersion of isotopic crustal ages is very high. The value would be very high under theoretical conditions, and thus based on Value calculated The value will be very high.

[0015] In step S3, the correlation between each initial feature is calculated using the Pearson correlation coefficient matrix; a correlation threshold is set, a correlation heatmap is generated, and highly correlated features are removed.

[0016] In step S4, the machine learning model training and optimization includes: After optimization and filtering, the obtained sample-feature matrix is ​​spliced ​​into a new two-dimensional sample-feature matrix; the new two-dimensional sample-feature matrix is ​​then divided into a training subset and a test subset according to the proportion. On the training subset, a grid search method is used to traverse the preset combination of random forest hyperparameters; for each set of parameters, the average value of the evaluation index of model performance is calculated using cross-validation. Random forest algorithms possess strong capabilities for handling high-dimensional features. In geology, complex correlations and nonlinear relationships exist between data from different sources. Random forests, through a large number of decision trees, randomly select features. They can effectively integrate complex and nonlinear relationships between sources without relying on linear assumptions. Because random forests integrate the results of multiple decision trees, they typically achieve high prediction accuracy and perform exceptionally well in classification tasks. Geological processes are often mixed and gradual environments, so the large number of decision trees within random forests can output the required probabilities (contribution quantifications) with minimal error in such environments, exhibiting strong resistance to overfitting, making them particularly suitable for these conditions. Random forests incorporate a method for evaluating feature importance. In geology, features with high importance rankings are typically key geochemical indicators controlling the differences in lithogenesis and evolution of source regions.

[0017] Select the combination of hyperparameters that optimizes the average value of the evaluation metrics as the optimal hyperparameters; use the optimal hyperparameters and all training set data to retrain and obtain the final random forest model; After the model training is completed, the importance ranking of each optimized feature for the source region classification is extracted and output.

[0018] Furthermore, the hyperparameters of the random forest include at least the number of decision trees, the maximum number of splits, and the minimum number of leaf node samples, and the evaluation index is the Kappa coefficient.

[0019] In step S5, the probability of the output sample belonging to each potential source region is: The classification voting results of all decision trees are aggregated using the trained optimal random forest model, and the proportion of votes obtained by each source region is taken as the contribution proportion of the sample from that source region.

[0020] Another objective of this invention is to provide a high-precision source region identification system based on multi-mineral tracing and machine learning, the system comprising a memory and a processor; The memory is used to store computer programs and a source region classification model constructed according to the high-precision source region identification method based on multi-mineral tracing and machine learning. The processor is used to execute the computer program to achieve the functions of the following modules: The data preprocessing module is used to clean, standardize, and label the collected raw mineral and geochemical data to build a training set; The multidimensional feature extraction module is used to extract age distribution features and geochemical features from the preprocessed training set data and construct the initial feature vector. The feature engineering module is used to perform correlation analysis and redundant feature screening on the initial feature vectors to construct an optimized sample-feature matrix. The model training and optimization module is used to train and optimize the source region classification model based on the random forest algorithm and the cross-validation grid search strategy. The source region prediction and output module is used to input the feature data of the sample to be tested into the optimized source region classification model and output the source region identification results and contribution ratio.

[0021] Combining all the above technical solutions, the beneficial effects of this invention are as follows: First, traditional methods rely on single minerals or a few geochemical features. This invention, by combining multi-mineral tracers (such as zircon U-Pb dating) with geochemical data (such as Hf isotopes and Th / U ratios) and machine learning algorithms, effectively reduces the ambiguity caused by single minerals or geochemical indicators, resulting in more accurate source region identification. It clearly distinguishes substances from different sources and their contribution ratios, significantly improving the reliability and accuracy of source region identification. This invention, through data preprocessing, feature extraction, and optimization using a machine learning-based random forest model, effectively integrates multiple features such as zircon age, Hf isotopes, and Th / U ratios, eliminating highly correlated features and optimizing the feature set. Experimental data shows that the machine learning model is more accurate in identifying source regions and quantifying the contribution ratios of different sources than traditional methods.

[0022] This invention automates the data processing and analysis process by introducing machine learning algorithms, significantly improving data processing efficiency, reducing manual intervention, and thus lowering experimental costs. Traditional methods rely on extensive manual operations or simple statistical analysis, resulting in low efficiency and high costs when processing large-scale samples. Machine learning models, on the other hand, can automatically perform feature extraction, data standardization, correlation analysis, and model training, making the entire analysis process more efficient, automated, and repeatable, capable of processing large amounts of sample data in a shorter time. This invention uses various functions in the Matlab platform (such as ksdensity, corrcoef, heatmap, etc.) for data standardization and feature extraction, while optimizing the processing and training of the feature matrix through the random forest algorithm. Experimental results show that when using machine learning algorithms for analysis, this invention is significantly faster than traditional methods, and the automated process greatly reduces manual costs.

[0023] Secondly, by employing the random forest algorithm in machine learning, this invention can more accurately quantify the contribution ratio of each source region to sediments. Traditional methods typically provide a coarse quantification of source region contributions and struggle to handle multi-source, complex data. This invention, however, achieves more refined and accurate source region contribution quantification through the fusion of multiple features (such as zircon age, Hf isotope values, Th / U ratio, etc.) and the optimization and training of these features using a machine learning model. In the implementation steps of this invention, feature correlation matrix filtering, high-dimensional data fusion, and random forest algorithm optimization enable precise quantification of the source region's contribution.

[0024] This invention employs machine learning techniques to effectively process high-dimensional and complex geological datasets, overcoming the limitations of traditional methods in handling large-scale, multi-dimensional data. Geological exploration often involves a large number of samples and complex mineral and geochemical characteristics, while traditional methods are prone to overfitting or excessive computation when dealing with high-dimensional data. This invention addresses the challenges of high-dimensional data analysis through feature selection and model optimization.

[0025] Third, by introducing machine learning algorithms, this invention automates the data preprocessing, feature extraction, and analysis processes, greatly reducing reliance on manual operations and tedious experimental work. This significantly lowers experimental and labor costs, meeting the stringent economic requirements of large-scale exploration tasks. Furthermore, computational learning algorithms greatly accelerate the progress of research and production projects, resulting in time savings. Traditional source analysis suffers from multiple solutions and overlapping identification results; the introduction of machine learning algorithms enables accurate identification of source areas and precise quantification of the contribution ratio of different sources. Simultaneously, due to the complexity of geological backgrounds, this invention can be applied to complex geological environments, making it more applicable and commercially valuable in fields involving complex geological environments such as oil and gas exploration and mineral resource surveys.

[0026] Existing methods often rely on single indicators (such as zircon U-Pb age or a few geochemical elements) for qualitative or semi-qualitative provenance identification, which can only roughly infer the source and cannot provide an objective and high-precision numerical characterization of the contribution ratio of different provenance areas. Furthermore, existing technologies generally use mineral tracing data and statistical analysis separately, lacking a complete technical solution that integrates multidimensional mineral-geochemical feature systems and machine learning modeling for provenance area identification. This invention innovatively establishes an intelligent identification model that integrates multidimensional features such as zircon U-Pb age, Hf isotopes, and Th / U ratio with a random forest algorithm. This overcomes the limitations of traditional statistical methods in processing high-dimensional, nonlinear geological data, achieving a leap from coarse qualitative tracing to precise quantitative contribution calculation. It effectively solves the technical problems of overlapping provenance area identification results in complex geological environments, difficulty in quantifying the contribution ratio of different provenance areas, and low efficiency and high cost in large-scale exploration tasks.

[0027] For a long time, provenance analysis has been limited by the linear thinking of traditional single-mineral tracing (such as zircon U-Pb dating) and simple statistical methods (such as principal component analysis), making it difficult to overcome bottlenecks such as multiple solutions, overlapping age spectra, and insufficient fusion of multidimensional data, thus preventing the transition from qualitative description to quantitative calculation. This invention, by introducing the random forest machine learning algorithm, creatively constructs a technical system for the deep fusion of multi-mineral tracing and geochemical data. By using kernel density estimation to extract features and performing correlation filtering, it not only significantly reduces manual costs and subjective errors, but also achieves an excellent test set accuracy of 85.7% in experiments. It successfully overcomes the core technical challenge of traditional methods being unable to effectively handle high-dimensional, nonlinear, and complex data to achieve accurate quantification of provenance contributions. Attached Figure Description

[0028] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure; Figure 1 This is a flowchart of a high-precision source region identification method based on multi-mineral tracing and machine learning provided in an embodiment of the present invention; Figure 2 This is a distribution map of the source area in the southwest of the tower provided in the embodiments of the present invention; wherein, (a) is an age probability density curve and (b) is an original age distribution map; Figure 3 This is a distribution map of the West Kunlun source region provided in the embodiments of the present invention; wherein, (a) is an age probability density curve and (b) is an original age distribution map; Figure 4 This is a distribution map of the Tianshan source area provided in the embodiments of the present invention; wherein, (a) is an age probability density curve and (b) is an original age distribution map; Figure 5 This is a feature correlation matrix heatmap provided in an embodiment of the present invention; Figure 6 This is a feature importance ranking diagram of the random forest model provided in this embodiment of the invention; Figure 7 This is the evaluation graph of the random forest optimization model provided in the embodiments of the present invention; Figure 8 This is a training set confusion matrix analysis diagram provided in an embodiment of the present invention; Figure 9 This is a test set confusion matrix analysis diagram provided in an embodiment of the present invention. Figure 10 This is a graph showing the quantification results of the material source contribution provided in an embodiment of the present invention. Detailed Implementation

[0029] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of the present invention. However, the present invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0030] The key technological innovations of this invention include: (1) the determination logic of the analysis scheme for systematically integrating traditional and non-traditional indicators; (2) Definition of indicators; (3) High-precision source area contribution quantification identification; Compared with the prior art, the key point of this invention is high-precision source area contribution quantification identification, and the protection point is the technology of high-precision source area contribution quantification by combining multiple indicators with machine learning.

[0031] This invention introduces an optimized bandwidth selection strategy when generating age probability distribution curves using Matlab's ksdensity function. While the default parameters can handle general cases, in geological scenarios with sparse or complex data distributions, cross-validation can be used to automatically find the optimal bandwidth that minimizes the mean square error of integral (MISE). By adaptively adjusting the bandwidth, it is possible to effectively avoid overly smoothed curves (loss of key age peak details) due to excessive bandwidth or overfitting (introducing noise) due to insufficient bandwidth, thereby extracting more robust age spectrum features.

[0032] Although the Random Forest algorithm is considered a preferred option due to its insensitivity to high-dimensional data, strong resistance to overfitting, and natural support for probabilistic output, the technical concept of this invention is also applicable to other machine learning models with high classification performance and probabilistic output capabilities. As alternatives, Gradient Boosting Decision Tree (GBDT) algorithms, such as XGBoost, LightGBM, or CatBoost, can be used. As long as the model can output probability values ​​or confidence scores belonging to each source region, and normalization can be used to quantify the source contribution ratio, these should all be considered equivalent implementations of this invention.

[0033] In the feature preprocessing stage, the Pearson correlation coefficient threshold used to remove redundant features (such as |r|>0.8 in the above embodiment) is not a fixed value. In practical applications, this threshold can be flexibly adjusted according to the dispersion of the data in the specific study area and the sample size, for example, set to |r|>0.75 or |r|>0.85, to balance the number of features retained with the model's need to remove noise. In addition, besides the correlation coefficient-based filtering strategy, this invention can also employ other feature screening mechanisms as alternatives or supplements.

[0034] The differences between this invention and existing technologies are: ① The fusion of multi-mineral tracer and geochemical data is used to improve the accuracy of source region identification. ② The application and optimization of machine learning models are used to achieve high-precision quantification of source region contributions.

[0035] Step 1: Identify the study area and provenance area. Collect zircon U-Pb, geochemical, and other data based on the literature. Remove data that are clearly outside the reasonable range. Standardize all geochemical concentration and age data using Z-score Z=(X-μ) / σ. Label the training set data according to their known provenance areas.

[0036] Step 2: Using the `ksdensity` function in Matlab 2016b, kernel density estimation is performed on the zircon age data of each sample to generate a continuous age probability distribution curve. The main age peaks, peak area ratios, and age spans are extracted as features from the age probability density curves. For geochemical data, representative zircon grains from each sample are obtained through literature review. Values, and based on the multiple values ​​obtained Value, calculate the average for each sample Value and Hf isotope dispersion ( ); The Th / U values ​​of zircon grains with determined ages in each sample were obtained through literature review.

[0037] Step 3: Use the corrcoef function in Matlab 2016b to calculate the Pearson correlation coefficient matrix between different features, set the correlation threshold (|r|>0.8), remove highly correlated features, and use the heatmap function to generate a correlation heatmap.

[0038] Step 4: Based on the correlation heatmap generated in Step 3, remove highly correlated features. Use Matlab 2016b to merge feature data from different sources, concatenating all feature values ​​of each sample into a one-dimensional feature vector, and stacking the feature vectors of all samples to form a standard sample-feature matrix. Use the cvpartition function of Matlab 2016b to randomly partition the feature matrix into training and test sets in an 8:2 ratio.

[0039] Step 5: Use the training set data to train the model using the random forest algorithm; and optimize the parameters of the random forest by combining network search with 5-fold cross-validation.

[0040] Step 6: Input the "sample-feature" matrix of the unknown samples in the study area into the trained final model to predict the contribution ratio of the source region for each sample.

[0041] Example 1, as Figure 1 As shown, the high-precision source region identification method based on multi-mineral tracing and machine learning provided in this embodiment of the invention includes the following steps: S1. Data Collection and Preprocessing: Collect sample data from the study area and potential provenance areas. The sample data should include at least zircon U-Pb age data and geochemical data. Clean and standardize the data, and assign provenance labels to the sample data from known provenance areas to form a training set. This invention collected 100 sandstone samples from the research area (basin center) and provenance areas (Tianshan, southwestern Tarim, and western Kunlun) of the Tarim Basin. Through systematic review of relevant domestic and international geological literature and publicly available geological databases, zircon U-Pb age data from 30 measuring points in each sample were obtained. The data includes zircon U-Pb ages and Th / U values. These three types of data constitute a complete source region analysis system from the perspectives of time, source region nature, crustal evolution, and zircon genesis. The cross-verification of source information from these three aspects—time, space, and material origin—is a key element in solving the problem. All data are derived from published academic papers or geological survey reports. For all numerical data, including zircon U-Pb ages, Th / U values, and Th / U values, the analysis is conducted in a comprehensive manner. The Th / U values ​​were Z-score standardized using the formula Z=(X-μ) / σ to eliminate dimensional differences between different features. Based on the clearly defined geological background in the literature, the provenance areas of 80 samples were accurately delineated. These 80 samples were used as the training set and labeled as southwestern Tarim, western Kunlun, and Tianshan. The remaining 20 samples were considered as unknown samples to be predicted.

[0042] S2. Multidimensional feature extraction: For each sample in the training set, the age distribution characteristics are obtained by estimating the kernel density of zircon U-Pb age data and analyzing the age probability density curve. Geochemical features are extracted from the geochemical data and together they form the initial feature vector of the sample. a. Age feature extraction: Using Matlab2016b software, the ksdensity function was called to estimate the kernel density of zircon U-Pb age data for each sample, generating a continuous age probability distribution curve.

[0043] The age probability density curves and original age distribution curves of the southwestern Tarim Basin, the western Kunlun Mountains, and the Tianshan Mountains are respectively as follows: Figure 2 , Figure 3 and Figure 4 As shown.

[0044] in, Figure 2 , Figure 3 and Figure 4 Here is an example of the age probability distribution curve for one sample from each source region. For each sample, the following features are extracted from its age probability distribution curve: ① Major age peaks: Identify and extract the age values ​​(Ma) corresponding to the most significant age peaks.

[0045] Zircon U-Pb age datasets collected from literature were imported into Matlab 2016b software. Without requiring any manual parameter settings, the ksdensity function was automatically invoked via input commands, ultimately generating a continuous age probability distribution curve. The curves for each provenance region are shown below. Figure 2 , Figure 3 and Figure 4 As shown, on the generated continuous probability distribution curve, the algorithm automatically searches for and locates the maximum point of the probability density (vertical axis). The horizontal axis value corresponding to this maximum point is the age value corresponding to the most significant age peak of the sample, and the value is recorded in an Excel spreadsheet as a feature value.

[0046] ② Peak area ratio: Calculate the percentage of area occupied by the main age peak.

[0047] The continuous probability density curve of the age distribution of a single sample, generated by the `ksdensity` function in Matlab 2016b, is used to automatically search and locate the maximum point of the probability density (vertical axis) based on the algorithm described above. From this peak point, a search is performed in both directions towards older and younger ages to find the nearest local minima. The horizontal axis values ​​corresponding to these two minima, i.e., the age values, define the integration interval of the main age peak. Within this defined integration interval, the probability density function is numerically integrated using the `trapz` function in Matlab 2016b. The calculated result is the probability area of ​​the main age peak. Since the total area of ​​the entire probability density curve is 1, the probability area of ​​the main age peak is numerically equal to its area ratio. This value is recorded in an Excel spreadsheet as a feature value.

[0048] ③ Age span: Calculate the range of age distribution from the youngest to the oldest.

[0049] By collecting zircon U-Pb age datasets from literature, the ages of different measurement points for each sample were found. According to the data processing step, 30 zircon ages were collected for each sample in this invention. The MAX and MIN functions were used in an Excel spreadsheet to search for the maximum and minimum ages, and the maximum age was subtracted from the minimum age to obtain the age range of the sample.

[0050] b. Geochemical Feature Extraction: Through literature review, geochemical features were extracted from each sample. Values. Based on these values, the average for each sample is calculated. Value and Isotope dispersion The expression is: ; In the formula, for Isotope dispersion For the first Zircon value, For this sample The arithmetic mean of the values This represents the number of zircon grains.

[0051] This formula is the mean square deviation formula in probability theory. In this invention, it is given a geological meaning. The high and low values ​​represent whether the source area is a single source area or a mixed source area, respectively, thus realizing a quantitative characterization of the complexity of the geological history of the source area.

[0052] Through literature review, the Th / U values ​​of zircon grains with determined ages in each sample were obtained, and their average value was calculated as the Th / U characteristic of the sample.

[0053] S3. Feature selection and matrix construction: Calculate the correlation between each initial feature, remove redundant features, and obtain the optimized feature set; merge the optimized feature sets of all samples to construct the sample-feature matrix; The `corrcoef` function in Matlab 2016b was used to calculate the Pearson correlation coefficient matrix among all features extracted in step S2 (including age features and geochemical features). A feature selection and optimization mechanism based on intelligent diagnosis was proposed and implemented to eliminate redundant features, thereby reducing overfitting during model training and improving generalization ability to ensure high-accuracy identification. Specifically, the original features of all samples extracted in step S2 (including main age peaks, peak area ratios, age spans, average values, etc.) were processed. Hf isotope dispersion The samples (including Th / U mean, etc.) are integrated in the Matlab 2016b environment to form an initial M×N two-dimensional "sample-feature" matrix, where M is the total number of samples (100) and N is the number of extracted original features (6). This "sample-feature" matrix is ​​used as the dataset for intelligent diagnosis. The `corrcoef` function in Matlab 2016b is used to perform a comprehensive correlation diagnosis on this initial feature matrix. This function generates a 6×6 Pearson correlation coefficient matrix, with a strict redundancy threshold |r|>0.8. The algorithm diagnoses this Pearson correlation coefficient matrix to identify all features exceeding this threshold. The diagnostic results are analyzed as follows: Figure 5 As shown, The correlation coefficient between the mean and the Th / U mean is as high as r≈0.95. Therefore, to ensure high accuracy, the mean with more comprehensive information is selected. The mean Th / U features are precisely removed. A heatmap function is then used to generate a feature correlation heatmap of the diagnostic results as visual evidence.

[0054] S4. Machine Learning Model Training and Optimization: The sample-feature matrix is ​​randomly divided into training and test sets. The random forest algorithm is used to train the source region classification model using the training set. Cross-validation and grid search strategies are used to optimize the model hyperparameters to obtain the best trained random forest model. The original 100×6 two-dimensional "sample-feature" matrix from step three, after optimization and filtering (removing the Th / U mean features), is reassembled into a new 100×5 two-dimensional "sample-feature" matrix. Using the `cvpartition` function in Matlab 2016b, the two-dimensional "sample-feature" matrix is ​​randomly divided into a training set (80 samples) and a test set (20 samples) at an 8:2 ratio.

[0055] This invention selects the random forest algorithm as the classifier. This algorithm can effectively handle high-dimensional feature data and is insensitive to multicollinearity. The model itself has good anti-overfitting ability and can output feature importance ranking, which is helpful for geological interpretation.

[0056] In Matlab 2016b, the TreeBagger function from the Statistics and Machine Learning Toolbox is used to build a random forest model. The specific process is as follows: Step 1: Input the training set feature matrix X_train: an 80×5 matrix; where 80 is the number of training samples and 5 is the total number of filtered features. Input the training set label vector Y_train: a vector containing 80 elements, each element corresponding to the source region label of a sample. Step 2: Initially set a baseline model, setting the number of decision trees NumTrees to 100 and the maximum depth to 3. Then, build a baseline random forest classifier and enable out-of-bag (OOB) error estimation. Step 3: The TreeBagger function trains the random forest model based on the input feature matrix and target labels. During training, each tree learns from different subsets of the training data, thereby improving the model's generalization ability.

[0057] To avoid the blindness of subjective parameter setting, this invention employs a strategy combining grid search and 5-fold cross-validation to automatically find the optimal hyperparameter combination. The optimized evaluation metric is the average Kappa coefficient of the cross-validation.

[0058] The search ranges were defined for three key hyperparameters of the random forest: number of decision trees (NumTrees): [20, 30, 50, 100], maximum number of splits (MaxNumSplits): [4, 6, 10], and minimum number of leaves (MinLeafSize): [5, 8, 12]. A Matlab script was then written to systematically iterate through all possible combinations of these three parameters. For each parameter combination: a. Use the cvpartition function to create a 5-fold cross-validation partition on 80 training samples.

[0059] b. In one loop, the model is trained using 4 folds of data, and validated using the remaining 1 fold of data.

[0060] c. Calculate the Kappa coefficient for this validation. After repeating the validation 5 times, calculate the average Kappa coefficient for this parameter combination.

[0061] d. Record the parameter combination and its corresponding average Kappa coefficient.

[0062] The system automatically compares the average Kappa coefficients obtained from multiple rounds of verification and determines the parameter combination that maximizes the average Kappa coefficient as the final optimal hyperparameter combination. This invention determines the optimal hyperparameter combination as: NumTrees=30, NumPredictorsToSample=6, MinLeafSize=8.

[0063] Using the optimal hyperparameters determined in the previous step (NumTrees=30, NumPredictorsToSample=6, MinLeafSize=8), retrain a final random forest model on all 80 training set samples.

[0064] After training, feature importance is extracted from the final model, these features are sorted from highest to lowest importance, and plotted as a bar chart, such as... Figure 6 As shown in the figure. It can be clearly seen from the image that... Main peak age and These three characteristics (>70%) are the most discriminative features for distinguishing the three source regions.

[0065] Using the trained final model, provenance regions were predicted for 20 independent test set samples from the previously defined basin center. The prediction results were compared with the true labels to generate a random forest model evaluation map. Figure 7 ) and confusion matrix analysis chart ( Figure 8 , Figure 9 ).pass Figure 7 , Figure 8 and Figure 9 As can be seen, the accuracy of the model reached 95.7% for 80 samples from the three source regions in the training set, and 85.7% for 20 samples from the basin center in the test set. The out-of-bag (OOB) error estimation reached 89.1%. This means that the model has high accuracy on the training, out-of-bag, and test datasets, and the accuracy difference between the training and test sets is not significant, indicating that the model has good fitting effect and strong generalization ability, and can be effectively applied to the source region identification task.

[0066] S5. Source Region Identification and Contribution Quantification: Input the feature data of the sample in the study area to be identified into the trained model, and output the probability of the sample belonging to each potential source region to characterize the contribution ratio of each source region.

[0067] The "sample-feature" matrix of 20 samples from the unknown study area was input into a pre-trained optimal random forest model. The model output the contribution proportion of each of these 20 samples from the three source regions of southwestern Tarim, western Kunlun, and Tianshan. Figure 10 As shown.

[0068] Example 2: This embodiment of the invention provides a high-precision source region identification system based on multi-mineral tracing and machine learning. The system includes a memory and a processor. The memory is used to store computer programs and a source region classification model constructed according to the high-precision source region identification method based on multi-mineral tracing and machine learning. The processor is used to execute the computer program to achieve the functions of the following modules: The data preprocessing module is used to clean, standardize, and label the collected raw mineral and geochemical data to build a training set; The multidimensional feature extraction module is used to extract age distribution features and geochemical features from the preprocessed training set data and construct the initial feature vector. The feature engineering module is used to perform correlation analysis and redundant feature screening on the initial feature vectors to construct an optimized sample-feature matrix. The model training and optimization module is used to train and optimize the source region classification model based on the random forest algorithm and the cross-validation grid search strategy. The source region prediction and output module is used to input the feature data of the sample to be tested into the optimized source region classification model and output the source region identification results and contribution ratio.

[0069] The embodiments of this invention have achieved significant beneficial effects during research and development and use, as verified by a practical application case in the Tarim Basin, demonstrating substantial advantages over existing technologies. The following detailed description of the positive effects, in conjunction with specific experimental procedures, data, and charts, further illustrates these advantages: 1. This invention demonstrates extremely high recognition accuracy and generalization ability. It was validated on 100 sandstone samples from the Tarim Basin research area (basin center) and its provenance areas (Tianshan, southwestern Tarim, and western Kunlun). Experimental data show that the random forest model constructed using this invention (optimal parameters: NumTrees=30, MinLeafSize=8) achieves an accuracy of 95.7% on the training set (80 samples) and 85.7% on the independent test set (20 basin center samples), with an out-of-bag (OOB) error estimate of 89.1%.

[0070] refer to Figure 7 (Evaluation diagram of random forest optimization model) and Figure 8 and Figure 9(Confusion matrix analysis) shows that the model maintains high accuracy across the training, out-of-bag, and test datasets, with a small difference in accuracy between the training and test sets. This indicates that the present invention not only fits known data well but, more importantly, possesses strong generalization ability, effectively overcoming the overfitting problem that often occurs in traditional machine learning methods. Compared to traditional single-indicator methods, which often suffer from misjudgments due to overlapping age spectra, the multi-indicator fusion strategy of the present invention significantly improves the reliability of source region identification.

[0071] 2. It achieves precise quantification of source contribution. This invention can not only identify the source region, but also output the contribution ratio of each sample from different source regions.

[0072] like Figure 10 (The graph showing the quantitative results of sediment contribution) illustrates that for 20 samples from an unknown study area, the model outputs the specific contribution proportions from three source regions: southwestern Tarim, western Kunlun, and Tianshan. This overcomes the limitations of traditional techniques, which typically only provide qualitative descriptions (e.g., mainly from area A, with a small amount mixed from area B) or extremely rough estimates. The quantitative results provided by this invention accurately reflect the mixing degree of sediments, offering a quantitative basis for decision-making in sedimentary migration path analysis during oil and gas exploration.

[0073] 3. Automated Feature Engineering and Effective Identification of Key Factors: This invention automatically extracts and optimizes key features through a standardized data processing workflow. Data Analysis: In the feature extraction stage, kernel density estimation (KDE) is used to transform discrete age data into continuous probability distribution curves, extracting features such as major age peaks and peak area ratios. In the feature selection stage, the Pearson correlation coefficient matrix (…) is used… Figure 5 Analysis revealed and eliminated highly redundant Th / U mean features (r≈0.95). Finally, the importance of the features extracted by the model was ranked. Figure 6 )show, The contribution of the mean and the main age peak exceeds 70%. This result proves that the method of the present invention can accurately capture the key geological indicators (such as Hf isotopes and main peak age) that play a decisive role in complex geological data, eliminate interference factors, and thus improve the interpretability and geophysical significance of the model, which is superior to the current situation where traditional methods are difficult to clarify the main controlling factors.

[0074] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention and within the spirit and principles of the present invention should be covered within the scope of protection of the present invention.

Claims

1. A high-precision source region identification method based on multi-mineral tracing and machine learning, characterized in that, The method includes the following steps: S1. Data Collection and Preprocessing: Collect sample data from the study area and potential provenance areas. The sample data should include at least zircon U-Pb age data and geochemical data. Clean and standardize the data, and assign provenance labels to the sample data from known provenance areas to form a training set. S2. Multidimensional feature extraction: For each sample in the training set, the age distribution characteristics are obtained by estimating the kernel density of zircon U-Pb age data and analyzing the age probability density curve. Geochemical features are extracted from the geochemical data and together they form the initial feature vector of the sample. S3. Feature selection and matrix construction: Calculate the correlation between each initial feature, remove redundant features, and obtain the optimized feature set; merge the optimized feature sets of all samples to construct the sample-feature matrix; S4. Machine Learning Model Training and Optimization: The sample-feature matrix is ​​randomly divided into training and test sets. The random forest algorithm is used to train the source region classification model using the training set. Cross-validation and grid search strategies are used to optimize the model hyperparameters to obtain the best trained random forest model. S5. Source Region Identification and Contribution Quantification: Input the feature data of the sample in the study area to be identified into the trained model, and output the probability of the sample belonging to each potential source region to characterize the contribution ratio of each source region.

2. The high-precision source region identification method based on multi-mineral tracing and machine learning according to claim 1, characterized in that, In step S1, the data is standardized to a Z-score, expressed as: ; In the formula, The distance of the original dataset from the overall data mean The standard deviation multiple, For a single data point in the dataset, The population mean of the original dataset. This represents the overall standard deviation of the original dataset.

3. The high-precision source region identification method based on multi-mineral tracing and machine learning according to claim 1, characterized in that, In step S2, the age distribution features extracted from the zircon U-Pb age data include at least one of the main age peaks, peak area ratios, and age spans.

4. The high-precision source region identification method based on multi-mineral tracing and machine learning according to claim 3, characterized in that, In step S2, the age distribution characteristics are obtained by estimating the kernel density of zircon U-Pb age data and analyzing the age probability density curve, including: Discrete zircon U-Pb age data are converted into continuous probability density curves using a kernel density estimation function. The main age peaks are automatically identified and extracted from the probability density curves, and the integration interval is defined. The probability density function is numerically integrated within the integration interval to obtain the peak area ratio of the main age peaks.

5. The high-precision source region identification method based on multi-mineral tracing and machine learning according to claim 4, characterized in that, In step S2, the geochemical data includes zircon data. Values ​​and / or Values, geochemical characteristics including average value, Isotope dispersion and / or average value; in, Isotope dispersion The calculation formula is: ; In the formula, for Isotope dispersion For the first Zircon value, For this sample The arithmetic mean of the values This represents the number of zircon grains.

6. The high-precision source region identification method based on multi-mineral tracing and machine learning according to claim 1, characterized in that, In step S3, the correlation between each initial feature is calculated using the Pearson correlation coefficient matrix; a correlation threshold is set, a correlation heatmap is generated, and highly correlated features are removed.

7. The high-precision source region identification method based on multi-mineral tracing and machine learning according to claim 1, characterized in that, In step S4, the machine learning model training and optimization includes: After optimization and filtering, the obtained sample-feature matrix is ​​spliced ​​into a new two-dimensional sample-feature matrix; the new two-dimensional sample-feature matrix is ​​then divided into a training subset and a test subset according to the proportion. On the training subset, a grid search method is used to traverse the preset combination of random forest hyperparameters; for each set of parameters, the average value of the evaluation index of model performance is calculated using cross-validation. Select the combination of hyperparameters that optimizes the average value of the evaluation metrics as the optimal hyperparameters; use the optimal hyperparameters and all training set data to retrain and obtain the final random forest model; After the model training is completed, the importance ranking of each optimized feature for the source region classification is extracted and output.

8. The high-precision source region identification method based on multi-mineral tracing and machine learning according to claim 7, characterized in that, The hyperparameters of the random forest include at least the number of decision trees, the maximum number of splits, and the minimum number of leaf node samples, and the evaluation index is the Kappa coefficient.

9. The high-precision source region identification method based on multi-mineral tracing and machine learning according to claim 1, characterized in that, In step S5, the probability of the output sample belonging to each potential source region is: The classification voting results of all decision trees are aggregated using the trained optimal random forest model, and the proportion of votes obtained by each source region is taken as the contribution proportion of the sample from that source region.

10. A high-precision source region identification system based on multi-mineral tracing and machine learning, characterized in that, The system includes memory and a processor; The memory is used to store computer programs and a source region classification model constructed according to any one of claims 1-9 based on the high-precision source region identification method of multi-mineral tracing and machine learning. The processor is used to execute the computer program to achieve the functions of the following modules: The data preprocessing module is used to clean, standardize, and label the collected raw mineral and geochemical data to build a training set; The multidimensional feature extraction module is used to extract age distribution features and geochemical features from the preprocessed training set data and construct the initial feature vector. The feature engineering module is used to perform correlation analysis and redundant feature screening on the initial feature vectors to construct an optimized sample-feature matrix. The model training and optimization module is used to train and optimize the source region classification model based on the random forest algorithm and the cross-validation grid search strategy. The source region prediction and output module is used to input the feature data of the sample to be tested into the optimized source region classification model and output the source region identification results and contribution ratio.