A method for calculating water content in zircon using machine learning modeling

By using machine learning modeling and LA-ICP-MS technology, a model for calculating the water content of zircon was established, which solved the problems of high cost and limited applicability of traditional methods. This enabled low-cost and efficient calculation of the water content of zircon, thus promoting geological research.

CN122337397APending Publication Date: 2026-07-03CHINA GOLD INNER MONGOLIA MINING +3

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA GOLD INNER MONGOLIA MINING
Filing Date
2025-12-17
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing methods for measuring the water content of zircon are expensive and not widely adopted. Traditional methods such as FTIR and SIMS require expensive equipment and complex sample preparation processes, and only a few institutions have the testing capabilities, which limits the acquisition and application of zircon water content data.

Method used

Machine learning modeling was employed, utilizing published SIMS zircon water content data. A zircon water content calculation model was established using a machine learning regression algorithm. Zircon trace element data was obtained by combining LA-ICP-MS, and a random forest model was used for calculation, reducing costs and improving efficiency.

Benefits of technology

It enables efficient and low-cost calculation of zircon water content with an accuracy of R²=0.957, lowering the technical threshold for researchers, promoting more geological research, and providing more geological information, especially for the study of magma evolution and mineralization mechanisms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122337397A_ABST
    Figure CN122337397A_ABST
Patent Text Reader

Abstract

The application provides a method for calculating water content in zircon by machine learning modeling, construction of a zircon database, cataloged training data set, classification model construction and selection, identification of important features and model hyperparameter optimization, application of the zircon water content calculation model, determination of the proportion of zircon OH ‑ content to determine available zircon water content data. Through the technical solution of the application, the patent uses big data machine learning analysis technology, and on the basis of part of the published SIMS zircon water content data, a regression algorithm of machine learning is used to establish an algorithm model for calculating the water content in zircon. This technology only needs conventional laser ablation inductively coupled plasma mass spectrometry zircon trace element data, and can accurately calculate the water content by combining the optimized random forest model. The model innovatively introduces multidimensional feature engineering, screens out 8 core parameters such as Age, Nb and Ce through recursive feature elimination, and establishes a granite type specific correction coefficient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of mineralogy research and mineral exploration, and more particularly to a method for calculating the water content in zircon using machine learning modeling. Background Technology

[0002] Zircon (ZrSiO4) has a stable crystal structure and is widely distributed in silicate rocks. It preserves rich information about the original magma, such as geological age, temperature, oxygen fugacity, and magma source region. Many recent studies have also found that the trace elements in zircon can reflect the water content of the parent magma. Ge et al. (2023) proposed a method for calculating magma water content based on zircon oxygen fugacity and whole-rock composition. However, considering the mixing and assimilation of melts within the magma system, the host rock composition may not represent the melt during zircon crystallization, especially for trapped zircons and zircons that have undergone multi-stage growth. Another index for assessing magma water content has been proposed. 4 ×(Eu / Eu*) / Yb N The analysis relies solely on trace elements in zircon. It is based on the elemental distribution among plagioclase, amphibole, and zircon during the crystallization of water-rich magma. However, due to the low zircon water content, the correlation between zircon and magmatic water content, and the controlling factors of zircon water content in different types of granites, remain unclear.

[0003] Zircon, as a "nominal anhydrous mineral," was once thought to be only capable of holding water, specifically metamorphic zircon. However, it has been discovered that crystalline zircon can also hold water in the form of OH radicals. - While water diffuses more slowly in zircon than in other magmatic minerals such as olivine, garnet, and orthopyroxene, a single zircon crystal can hold 100–1000 ppm of water. Therefore, zircon water content is used to study magma evolution, the degree of metal enrichment in ore deposits, and to differentiate granite types.

[0004] Currently, there are two main methods for obtaining the water content of zircon in crystalline zircon. 1. Measuring the integral area of ​​the OH stretching band using Fourier transform infrared spectroscopy (FTIR) (Trail et al., 2011) and calculating the water content. However, this method is strongly dependent on the crystallization orientation of the zircon and requires zircon grains to be >80×80μm in size. This method, based on spectral peak area, yields a semi-quantitative water content, and the zircon grains are typically small, with most zircon grain sizes failing to meet the testing requirements. 2. The newly improved secondary ion mass spectrometry (SIMS) can directly measure the OH- in zircon and calculate the water content. This is currently the most advanced and reliable method for quantitatively testing the water content of zircon. However, this requires pre-attaching the zircon grains to a special tin-bismuth alloy target and taking cathodoluminescence images of the zircon sample to confirm the test sites. Before formal testing, the sample also needs to be baked for a long time (~5-7 days) to eliminate interference from adsorbed water on the sample surface. Sample preparation and testing are both very expensive. The alloy target costs around 3000 yuan, and the price of cathodoluminescence imaging depends on the number of images, generally around 30-50 yuan per image, totaling approximately 5000-10000 yuan for all samples. SIMS testing also costs 25000 yuan per day. Currently, only the Guangzhou Institute of Geochemistry, Chinese Academy of Sciences, has a mature and reliable SIMS instrument for zircon water content testing, and in most cases, this instrument is used to test mineral samples from lunar regolith collected by the Chinese government. Therefore, the high price and scarcity of the instrument limit the widespread application of this method, and only a small amount of zircon water content data has been tested and published by this institute and a few collaborating institutions.

[0005] Machine learning is an artificial intelligence technique that can automatically learn patterns from massive amounts of data and use these patterns for prediction. Zircon composition data is widely published and readily available, and is primarily used in machine learning to distinguish host rocks and deposit types. Furthermore, machine learning can help calculate the content of certain elements in minerals, establish mineral thermobarometers, or other crystallization conditions. For example, multivariate polynomial regression analysis was used to determine the Li content based on the major element content in lepidolite; a machine learning method for calculating the P content in zircon was proposed using support vector machines; and least squares analysis and extreme random tree algorithms were used to estimate the temperature and pressure of the solution based on the composition of clinopyroxene. Therefore, by employing appropriate machine learning methods to learn the relationship between zircon water content and trace elements obtained from SIMS testing, it is possible to construct an algorithmic model that calculates zircon water content based solely on zircon trace elements.

[0006] Zircon is a crucial indicator mineral in the study of geological processes and metal mineralization mechanisms. Extensive research both domestically and internationally has confirmed that zircon can indicate conditions such as magma temperature, oxygen fugacity, water content, and source material properties. Regarding water content, the only current method for quantitatively determining the water content in zircon is the secondary ion mass spectrometry (SIMS) instrument used at the Guangzhou Institute of Geochemistry, Chinese Academy of Sciences. Specifically, the CMECA IMS 1280-HR, manufactured in France and improved by China, requires rigorous sample preparation, and both sample preparation and testing are costly. This reality makes it difficult for many geological researchers to analyze the water content in zircon, and consequently, to investigate the water content in the parent magma of zircon. Summary of the Invention

[0007] To overcome the shortcomings of existing technologies, this invention provides a method for calculating the water content in zircon using machine learning modeling. This patent employs big data machine learning analysis technology, building upon some published SIMS zircon water content data, and uses a regression algorithm to establish an algorithmic model for calculating the water content in zircon. One of the characteristics of this method is its multidisciplinary approach, integrating mineralogy, geology, and computer science, reflecting the trend of data-driven earth science research. At the operational level of zircon water content calculation, this method allows geologists to calculate the water content of zircon simply by testing its trace elements, achieving a high accuracy (R0). 2 The value can reach 0.957. The zircon trace elements are obtained using laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS), a relatively mature and widely used in-situ micro-area mineral testing technique. This technique is significantly cheaper than SIMS for determining zircon water content, and offers faster testing speeds and lower sample preparation requirements. More significantly, it can calculate the corresponding zircon water content from a large number of published zircon trace element data, providing more geologically valuable information. Therefore, the zircon water content algorithm model mentioned in this patent objectively reduces costs and increases efficiency, encouraging more researchers to study and interpret zircon water content. Water content is also a crucial factor limiting metal enrichment in rock masses, thus providing a new opportunity to study the relationship between metal enrichment and water content occurrence.

[0008] This invention is achieved through the following technical solution: a method for calculating the water content in zircon using machine learning modeling, specifically including the following steps: Step S1: Construction of the Zircon Database: Two datasets are compiled: a training dataset and a global dataset. The training dataset is used to train the zircon H2O content calculation model, while the global dataset is used to explore the variation pattern of zircon water content from a larger scale of data. The training dataset contains the zircon H2O content obtained from test analysis as the target value and 29 features, while the global dataset only contains 29 features. The target value of H2O content needs to be obtained through the calculation model. Step S2: The compiled training dataset is preprocessed by Python code. The data preprocessing process for AIS type granite zircon is the same, which includes outlier cleaning, missing value imputation, and data standardization. Step S3, Classification Model Construction and Selection: Perform regression analysis on the processed training dataset, using H2O as the target value and 27 elements and calculated values ​​as features; the coefficient of determination (R²) is used. 2 R is used to measure the performance of the evaluation model (Equation 7); 2 The closer the value is to 1, the better the model performs; R 2 = (Formula 7) y i This is the actual value. is the mean of the actual values, and n is the total number of data points.

[0009] Nine machine learning strategies were compared: linear regression (0.47), Lasso (0.46), support vector machine (-0.06), K-nearest neighbors (0.53), artificial neural network (0.12), random forest (0.63), gradient boosting decision tree (GBDT) (0.60), XGBoost (0.56) and LightGBM (0.54). Step S4: Identification of Important Features and Model Hyperparameter Optimization: Based on the selection of Random Forest as the computational model, RFE (Reference-Free Evaluator) is used to screen the feature importance of the model; for the zircon H2O content calculation model, Age, Nb, Ce, Sm, Lu, 10 4 ×(Eu / Eu*) / Yb N Dose (×10) 15 Yb / Gd were selected as the 8 most important features by RFE screening; The random forest model is simplified using eight key features and employs Bayesian optimization to find the optimal combination of hyperparameters. Thus, based on the training dataset, a random forest calculation model for zircon water content with 8 features was obtained. The model's R... 2 =0.66; H2O RFThe values ​​were then fitted again to the actual water content H2O to obtain the final predicted values ​​of zircon water content H2O* and R. 2 To avoid being affected by large deviations in H2O ML Value interference, ignore |H2O- H2O RF Discrete values ​​of | / H2O≥0.3; H2O* = 0.9993 H2O RF 1.0396 (R 2 = 0.957) (Formula 11) Step S5: Application of the Zircon Water Content Calculation Model The compiled global dataset containing 8 features is imported into the random forest model and Formula 11 built based on the training dataset to calculate the zircon water content in the global dataset; the model calculates the zircon water content based on Dose (×10). 15 The LREE-I filter removes data that is disturbed by inclusions and uses random forest and power function fitting strategies to give the water content of zircon. For the global dataset, a correction factor α is needed to obtain a more accurate predicted value of zircon water content H2O*. α = 1 + (Median of H2O in the training dataset - H2O of each granite class in the global dataset) # (median) / (H2O per type of granite in the global dataset) # (The median) (Formula 12) H2O* = α×H2O # (Formula 13) Step S6, combining zircon OH - The available zircon water content data can be used to determine the content ratio. Water in zircon exists in three forms: molecular water (H₂O), structural water (H₂O), and water in zircon. + OH - H + With (Y+REE) 3+ Coupling replaces Zr 4+ or Si 4+ It becomes part of the zircon structure (Formula 14). OH - Zircon is substituted with hydrogarnet-type ions (Formulas 15 and 16). Zircon contains only OH. - It can accurately record the water activity of the host magma; H+ + (REE+Y) 3+ = Zr 4+ (Formula 14) H + + Al 3+ = Si 4+ (Formula 15) (Formula 16). As a preferred option, step S1 specifically includes the following steps: Step S1-1: The training dataset contains a total of 1418 zircon analyses, of which 399 are from the Himalayan orogenic belt, 674 are from the North China Craton, 218 are from South China, 20 are from the Czech Republic, and 107 are from the New England Orogen in Australia; including 129, 985, and 304 zircon data from A-type, I-type, and S-type granites, respectively. The target value H2O in the training dataset was measured using SIMS; the training dataset contains 29 features, namely: zircon. 206 Pb / 238 U-age, labeled as Age; Trace elements in 18 zircon samples: Hf, U, Th, Y, P, Ce, Nb, Pr, Sm, Eu, Gd, Tb, Dy, Ho, Er, Tm, Yb, Lu; 10 calculated values: Eu / Eu*, Ce / Ce*, 10 4 ×(Eu / Eu*) / Yb N , 10 4 ×Ce / Nd / Y, U / Yb, Dose (×10 15 ), LREE-I, Yb / Gd, Ce / Sm, Th / U; where 206 Pb / 238 U-age and trace elements were obtained simultaneously by LA-ICP-MS, after water testing, and at the same location as the water testing points; The calculated values ​​selected indicators that can determine zircon crystal structure and indicate magma water content include LREE-I and Dose (×10). 15 ), 10 4 ×Ce / Nd / Y, 10 4 ×(Eu / Eu*) / Yb N and U / Yb; LREE-I and Dose (×10) 15 It is also an indicator for determining whether zircon has undergone hydrothermal alteration and metamorphism; 10 4 ×Ce / Nd / Y, 10 4 ×(Eu / Eu*) / Yb N U / Yb are indicators reflecting the content of magmatic water; Steps S1-2: The global dataset contains a total of 4196 zircon data points, with 29 features identical to those in the training dataset; 1028 data points are A-type granites, 2409 are I-type granites, and 759 are S-type granites; all I-type granites are from arc-magmatic environments, including adakitic granites (labeled as rich granites) and low-mineral calc-alkaline granites (labeled as low-mineral granites), with 1472 and 937 zircon trace element data points respectively.

[0010] Furthermore, the LREE-I value in step S1-1 refers to the light rare earth element index, representing the relative abundance of light rare earth elements to medium rare earth elements. The specific formula is as follows: LREE-I = (Dy / Nd) + (Dy / Sm) (Formula 1) Dose (×10 15 The value is based on the α-decay dose of radioactive elements in zircon and indicates the degree of self-radiation damage within the zircon. The specific formula is as follows: D = 8 N 1[exp(a1t) - l] + 7 N 2[exp(a2t) - l]+ 6 N 3[exp(a3t) - l] (Formula 2) In the formula, D is the α-decay radiation dose, and the unit is event / mg. N 1, N 2 and N 3 represents the current quantities of 238U, 235U, and 232Th, respectively, and a1, a2, and a3 represent... 238 U、 235 U and 232 The decay constant of Th, where t is the age of zircon; The Eu / Eu* value is calculated using the following formula: Eu / Eu*=(Eu / 0.0563) / [(Sm / 0.148)×(Gd / 0.199)] 1 / 2 (Formula 3) The Ce / Ce* value is calculated using the following formula: Ce / Ce*=(Ce / 0.613) / [(Nd / 0.457) 2 / (Sm / 0.148)] (Formula 4) 10 4 ×(Eu / Eu*) / Yb N The value, using zircon composition to reflect the water content level of the melt, is calculated using the following formula: 10 4 ×(Eu / Eu*) / Yb N= 10 4 ×(Eu / Eu*) / (Yb / 0.161) (Formula 5).

[0011] As a preferred option, step S2 specifically includes the following steps: Step S2-1: Outlier cleanup; La and Pr elements are removed from zircon, P>10 4 Analysis of ppm was omitted, LREE-I <10 and Dose (×10) 15 Zircon samples with a value less than 3 were excluded from analysis. Step S2-2: The missing values ​​are filled with the mean; Step S2-3: Data Standardization; Data standardization uses z-score, which calculates the difference between each data point and the overall data mean, and measures it in units of standard deviation. Z = (X−μ) / σ (Formula 6) In the formula, Z is the standardized value Z-score, X is the original data value, μ is the mean of the dataset, and σ is the standard deviation of the dataset.

[0012] As a preferred option, the linear regression model formula in step S3 is as follows, where w is the weight coefficient vector: y = w0 + w1x1 + w2x2 + ... + w n x n + ϵ (Formula 8) Lasso is a model that adds an L1 regularization term to the linear regression loss function; L1 regularization refers to the sum of the absolute values ​​of the weight coefficient vector w; the loss function formula is as follows, where ||w||1 = ∑|w j | represents the L1 norm, and λ is a hyperparameter that controls the severity of the penalty; (Formula 9) Support Vector Machines use kernel tricks to map low-dimensional linearly inseparable data to a high-dimensional space, making it linearly separable in the high-dimensional space, thereby handling complex nonlinear decision boundaries. When making predictions, K-Nearest Neighbors calculates the distance between a new sample and all samples in the dataset, and then finds the K nearest neighbors. For regression tasks, it takes the average of the target values ​​of the K neighbors. The formula is as follows: Nk(x) is the set of indices of the K nearest neighbors of x, and mode is the mode.

[0013] y = mode ({ yi | i ∈ N k ( x )}) (Formula 10) Random forests work by constructing multiple decision trees and combining their predictions; during training, each tree not only performs random sampling with replacement from the original dataset, but also randomly selects a subset of features from all features for optimal splitting at each node.

[0014] By employing the above technical solutions, this invention has the following beneficial effects compared to existing technologies: This patent achieves highly efficient calculation of water content in zircon through machine learning modeling, offering significant advantages over existing technologies. Firstly, it completely overcomes the technical bottlenecks of traditional methods: Fourier transform infrared spectroscopy (FTIR) testing relies on the crystallization direction of zircon and requires a particle size >80μm×80μm, while secondary ion mass spectrometry (SIMS) requires a special tin-bismuth alloy target (costing 3000 RMB), 5-7 days of baking treatment, and cathodic emission imaging (costing 5000-10000 RMB), resulting in a daily testing cost as high as 25000 RMB, and is only available to a few institutions. This technology only requires conventional laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS) to obtain trace element data in zircon, combined with an optimized random forest model (R²=0.957) to accurately calculate water content. This lowers the technical threshold, time, and financial costs for researchers to obtain water content data in zircon.

[0015] Furthermore, the model innovatively introduces multi-dimensional feature engineering, using recursive feature elimination (RFE) to filter out eight core parameters, including Age, Nb, and Ce, and establishes granite type-specific correction coefficients (1.62 for A-, 1.28 for I-, and 0.35 for S-type). This is particularly relevant to addressing the low OH⁻ / H⁺ ratio in S-type granites. Figure 2 (C,D) The analysis of structural water occurrence mechanisms ensures the geological significance of the calculation results. More significantly, this method enables the reuse of massive historical data—4196 published zircon trace element data points worldwide can be directly used to generate zircon water content using this model, opening up new avenues for the study of magma evolution and mineralization mechanisms.

[0016] Additional aspects and advantages of the invention will become apparent in the following description or may be learned by practice of the invention. Attached Figure Description

[0017] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1Box plots of trace elements and calculated values ​​for zircon in the training dataset are provided. The data distribution ranges for the target value H₂O and 35 candidate features are listed. Candidate features include age, 21 trace elements, and 13 calculated values, with the 31 marked in blue being the features used in the model. 1-Eu / Eu*, 2-10⁴ × (Eu / Eu*) / YbN, 3-Ce / (Ui×Ti)^0.5, 4-△FMQ, 5-Ce / Ce*, 6-1000×(Ce / Nd)N / Y, 7-T(K), 8-Dose (×10¹⁵), 9-LREE-I, 10-U / Yb, 11-Yb / Gd, 12-Ce / Sm, 13-Th / U.

[0018] Figure 2 Relationship between zircon OH-, H+ and total H2O in A-, I- and S-type granites In the figure, (A) is the OH- vs H2O projection of the training dataset; (B) is the OH- vs H2O projection of the global dataset; (C) is the range of OH- / H+ in the training dataset; and (D) is the range of OH- / H+ in the global dataset. H2Ototal (apfu×1000)=(H2Oppm×2 / 18) / 5, H+(apfu×1000)=(REE + Y3+)(apfu×1000)- P5+(apfu×1000) (Trail et al., 2011), OH-(apfu×1000)=H2O(apfu×1000)-H+(apfu×1000). The positions of the median values ​​are marked by stars.

[0019] Figure 3 Key features of the regression model for calculating zircon water content and the projection of zircon and key features in AIS-type granite. In the figure, (a) is the ranking of feature importance; (b) H2O* (ppm) vs. Dose (×10¹⁵); (c) H2O* (ppm) vs. Ce / (Ui×Ti)^0.5; (d) H2O* (ppm) vs. Yb / Gd; (e) H2O* (ppm) vs. 10⁴×(Eu / Eu*) / YbN. Detailed Implementation

[0020] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0021] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0022] The following is combined Figures 1 to 3 The method of using machine learning modeling to calculate the water content in zircon is described in detail in the embodiments of the present invention.

[0023] This invention proposes a method for calculating the water content in zircon using machine learning modeling, specifically including the following steps: Step S1: Construction of the Zircon Database: Two datasets are compiled: a training dataset and a global dataset. The training dataset is used to train the zircon H2O content calculation model, while the global dataset is used to explore the variation pattern of zircon water content from a larger-scale dataset. The training dataset contains the zircon H2O content obtained from test analysis as the target value and 29 features, while the global dataset only contains 29 features. The target H2O content needs to be obtained through the calculation model. Specifically, the following steps are included: Step S1-1: The training dataset contains a total of 1418 zircon analyses, all crystallized from granitic rocks rather than captured or inherited. 399 analyses are from the Himalayan orogenic belt, 674 from the North China Craton, 218 from South China, 20 from the Czech Republic, and 107 from the New England Orogen in Australia; including 129, 985, and 304 zircon data points from A-type, I-type, and S-type granites, respectively. Figure 1 ); The target value H2O in the training dataset was measured using SIMS; the training dataset contains 29 features, namely: zircon. 206 Pb / 238 U-age, labeled as Age; Trace elements in 18 zircon samples: Hf, U, Th, Y, P, Ce, Nb, Pr, Sm, Eu, Gd, Tb, Dy, Ho, Er, Tm, Yb, Lu; 10 calculated values: Eu / Eu*, Ce / Ce*, 10 4 ×(Eu / Eu*) / Yb N , 10 4 ×Ce / Nd / Y, U / Yb, Dose (×10 15 ), LREE-I, Yb / Gd, Ce / Sm, Th / U; where 206 Pb / 238 U-age and trace elements were obtained simultaneously by LA-ICP-MS, after water testing, and at the same location as the water testing points; The calculated values ​​are based on mineralogical and ore deposit research experience, selecting indicators that can determine zircon crystal structure and indicate magma water content, including LREE-I and Dose (×10). 15 ), 10 4 ×Ce / Nd / Y, 10 4 ×(Eu / Eu*) / Yb N and U / Yb; LREE-I and Dose (×10) 15 It is also an indicator for determining whether zircon has undergone hydrothermal alteration and metamorphism; 10 4 ×Ce / Nd / Y, 10 4 ×(Eu / Eu*) / Yb N U / Yb is considered an indicator of magma water content; the LREE-I value, referencing the light rare earth element index, represents the relative abundance of light rare earth elements to medium rare earth elements, and the specific formula is as follows: LREE-I = (Dy / Nd) + (Dy / Sm) (Formula 1) Dose (×10 15 The value is based on the α-decay dose of radioactive elements in zircon and indicates the degree of self-radiation damage within the zircon. The specific formula is as follows: D = 8 N 1[exp(a1t) - l] + 7 N 2[exp(a2t) - l] + 6 N 3[exp(a3t) - l] (Formula 2) In the formula, D is the α-decay radiation dose, and the unit is event / mg. N 1, N 2 and N 3 represents the current quantities of 238U, 235U, and 232Th, respectively, and a1, a2, and a3 represent... 238 U、 235 U and 232 The decay constant of Th, where t is the age of zircon; The Eu / Eu* value is calculated using the following formula: Eu / Eu*=(Eu / 0.0563) / [(Sm / 0.148)×(Gd / 0.199)] 1 / 2 (Formula 3) The Ce / Ce* value is calculated using the following formula: Ce / Ce*=(Ce / 0.613) / [(Nd / 0.457) 2 / (Sm / 0.148)] (Formula 4) 10 4 ×(Eu / Eu*) / YbN The value, using zircon composition to reflect the water content level of the melt, is calculated using the following formula: 10 4 ×(Eu / Eu*) / Yb N = 10 4 ×(Eu / Eu*) / (Yb / 0.161) (Formula 5) It should be noted that, due to the limited data on zircon Ti content (61% missing), the calculated values ​​related to Ti content were not listed as features.

[0024] Steps S1-2: List the zircon data sources from different granites worldwide in the global dataset, used to calculate zircon H2O content and for case analysis. The global dataset contains a total of 4196 zircon data points, with 29 features identical to those in the training dataset; 1028 data points are from A-type granites, 2409 from I-type granites, and 759 from S-type granites; all I-type granites originate from arc-magmatic environments, primarily occurring in the Circum-Pacific and Tethys-Himalayan orogenic belts. It includes adakitic granites (labeled as high-grade granites) spatially closely related to copper deposits and low-grade calc-alkaline granites (labeled as low-grade granites), with 1472 and 937 zircon trace element data points respectively.

[0025] Step S2: The compiled training dataset is preprocessed by Python code to facilitate machine learning. It should be noted that the data preprocessing process for AIS-type granite zircon is the same, including outlier removal, missing value imputation, and data standardization; specifically, the following steps are included: Step S2-1: Outlier Cleanup; La and Pr elements are removed from zircon because they are present in low concentrations and are often contaminated by tiny LREE-rich mineral inclusions. Furthermore, intense hydrothermal alteration and metamorphism can lead to anomalous water content in zircon. The P content in zircon is typically in the range of several hundred to several thousand, with P > 10. 4 Analysis of ppm was omitted because it may be contaminated by apatite inclusions, which are common in zircons of granite; LREE-I <10 and Dose (×10) 15 Zircon analyses with ages <3 were discarded to avoid interference from alteration and metamorphism; zircon data with discordant ages also needed to be manually deleted. Step S2-2: Missing values ​​are filled with the mean to ensure the integrity of the dataset; Step S2-3: Data Standardization; Data standardization uses the z-score, whose core principle is to transform the original data into a dataset with a mean of 0 and a standard deviation of 1 that follows a normal distribution (or is close to a normal distribution). By calculating the difference between each data point and the overall data mean, and measuring it in units of standard deviation, the z-score can eliminate differences in data dimensions and variability, allowing data of different specifications to be compared and analyzed on the same scale.

[0026] Z = (X−μ) / σ (Formula 6) In the formula, Z is the standardized value Z-score, X is the original data value, μ is the mean of the dataset, and σ is the standard deviation of the dataset.

[0027] Step S3, Classification Model Construction and Selection: Perform regression analysis on the processed training dataset, using H2O as the target value and 27 elements and calculated values ​​as features. Figure 1 ); coefficient of determination (R) 2 R is used to measure the performance of the evaluation model (Equation 7); 2 The closer the value is to 1, the better the model performs; R 2 = (Formula 7) y i This is the actual value. is the mean of the actual values, and n is the total number of data points.

[0028] Nine machine learning strategies were compared: linear regression (0.47), Lasso (0.46), support vector machine (-0.06), K-nearest neighbors (0.53), artificial neural network (0.12), random forest (0.63), gradient boosting decision tree (GBDT) (0.60), XGBoost (0.56) and LightGBM (0.54). Linear regression is a model that makes predictions by fitting a linear relationship between an independent variable (X) and a dependent variable (y). Its goal is to find a straight line (or hyperplane) that minimizes the sum of squared prediction errors (residuals) from all data points to that line. The model formula is as follows, where w is the weight coefficient vector: y = w0 + w1x1 + w2x2 + ... + w n x n + ϵ (Formula 8) Lasso (Least Absolute Shrinkage and Selection Operator) is a model that adds an L1 regularization term to the linear regression loss function. L1 regularization refers to the sum of the absolute values ​​of the weight coefficient vector w; it compresses the coefficients of some unimportant features to exactly zero, thereby achieving feature selection and effectively preventing overfitting. The loss function formula is as follows, where ||w||1=∑|w j | represents the L1 norm, and λ is a hyperparameter that controls the severity of the penalty; (Formula 9) The core idea of ​​Support Vector Machines (SVM) is to find an optimal separating hyperplane that maximizes the "margin" between the two classes of data points. This hyperplane is determined by a few key support vectors (the data points closest to the hyperplane). SVM uses kernel tricks to map low-dimensional linearly inseparable data to a high-dimensional space, making it linearly separable in the high-dimensional space, thus handling complex nonlinear decision boundaries. K-Nearest Neighbors (K-NN) is an instance-based lazy learning algorithm. It does not have an explicit training process. During prediction, for a new sample, the distance (e.g., Euclidean distance) to all samples in the dataset is calculated, and then the K nearest "neighbors" are found; for regression tasks, the average of the K neighbor target values ​​is taken; the formula is as follows, where Nk(x) is the set of indices of x's K nearest neighbors, and mode is the mode.

[0029] y = mode ({ yi | i ∈ N k ( x )}) (Formula 10) Random forests are a representative of Bagging ensemble learning. They work by constructing multiple decision trees and combining their predictions. During training, each tree is not only randomly sampled with replacement from the original dataset (bootstrap), but also a subset of features is randomly selected from all features for optimal splitting at each node. This dual randomness effectively reduces the model's variance, prevents overfitting, and enhances generalization ability.

[0030] Gradient Boosting Decision Tree (GBDT) is a representative of Boosting ensemble learning. Its core idea is to sequentially build a series of weak decision trees (usually shallow trees), with each tree dedicated to correcting the prediction residuals of the previous tree; it uses the idea of ​​gradient descent, taking the negative gradient direction of the loss function as the target that the current model needs to improve. XGBoost (eXtreme Gradient Boosting) is an efficient implementation and extension of GBDT. It explicitly adds a regularization term to the loss function of GBDT to further prevent overfitting. LightGBM shares the same core principle as XGBoost and is another efficient implementation based on GBDT.

[0031] Step S4: Identification of Important Features and Optimization of Model Hyperparameters: Based on the selection of Random Forest as the computational model, RFE (Recursive Feature Elimination) is used to filter the importance of features in the model. RFE is a wrapper-style feature selection method based on model weights. Its core principle is to recursively eliminate the least important features, thereby ranking the features by importance. This process begins by training a baseline model using all features, and the model assigns a weight or importance score to each feature; RFE eliminates the lowest-ranked subset of features based on this score; then, the above process is repeated with the remaining features: the model is retrained, and the least important features are removed again; this process is repeated recursively until the preset number of features is met; finally, the reverse order in which the features are removed is their importance ranking; for the zircon H2O content calculation model, Age, Nb, Ce, Sm, Lu, 10 4 ×(Eu / Eu*) / Yb N Dose (×10) 15 Yb / Gd were selected as the 8 most important features by RFE screening; The Random Forest model simplifies the model using eight key features and employs Bayesian optimization to find the optimal combination of hyperparameters. Bayesian optimization is an efficient method for finding the global optimum of a black-box function (especially functions with extremely high evaluation costs). Its core principle is to construct the probability distribution of the objective function using a Gaussian process based on all existing evaluation results. Then, the sampling function balances sampling in regions of high uncertainty with sampling in regions of high predicted optimum value. This process continuously updates the understanding of the objective function, thereby approximating the global optimum with as few steps as possible.

[0032] Thus, based on the training dataset, a random forest calculation model for zircon water content with 8 features was obtained. The model's R... 2 =0.66; In order to further capture more complex nonlinear relationships, H2O RF The values ​​were then fitted again to the actual water content H2O to obtain the final predicted values ​​of zircon water content H2O* and R. 2 To avoid being affected by large deviations in H2O ML Value interference, ignore |H2O-H2O RF Discrete values ​​of | / H2O≥0.3; H2O* = 0.9993 H2ORF 1.0396 (R 2 = 0.957) (Formula 11) Step S5: Application of the Zircon Water Content Calculation Model The compiled global dataset containing 8 features is imported into the random forest model built based on the training dataset and Formula 11 to calculate the zircon water content in the global dataset; the model can automatically calculate the water content of zircon based on Dose (×10). 15 The LREE-I filter removes data interfered with by inclusions, but analyses with discordant ages need to be manually removed. It uses random forest and power function fitting strategies to provide zircon water content; A correction factor α is needed for the global dataset to obtain more accurate zircon water content predictions (H₂O*); this is because although the model performs well on the original dataset, it still needs to be adapted to the new dataset distribution. In this study, the correction factors for A-, I-, and S-type granites are 1.62, 1.28, and 0.35, respectively. This calculation does not change R0. 2 .

[0033] α = 1 + (Median of H2O in the training dataset - H2O of each granite class in the global dataset) # (median) / (H2O per type of granite in the global dataset) # (The median) (Formula 12) H2O* = α×H2O # (Formula 13) Step S6, combining zircon OH - The available zircon water content data can be used to determine the content ratio. Water in zircon exists in three forms: molecular water (H₂O), structural water (H₂O), and water in zircon. + OH - Molecular water typically resides in the pores, lattice defects, or inclusions of crystals, and is most commonly found in metamorphosed or altered zircon, exhibiting an anomalous water content reaching 16.6 wt.%. The calculation methods used in this study focus on water content in crystalline zircon expressed as H... + and OH - Water in form. H + With (Y+REE) 3+ Coupling replaces Zr 4+ or Si 4+ It becomes part of the zircon structure (Formula 14). OH - Zircon is substituted with hydrogarnet-type ions (Formulas 15 and 16). Zircon contains only OH. - It can accurately record the water activity of the host magma; H+ + (REE+Y) 3+ = Zr 4+(Formula 14) H + + Al 3+ = Si 4+ (Formula 15) (Formula 16) Type A granite has the highest OH content. - The median values ​​for H2O in the training and global datasets were 27 (apfu×1000) and 30 (apfu×1000), and 29 (apfu×1000) and 32 (apfu×1000), respectively. Figure 2 A, B). Type I granite has moderate OH content. - The median values ​​for H2O in the training and global datasets were 13 (apfu×1000) and 13 (apfu×1000), and 14 (apfu×1000) and 15 (apfu×1000), respectively. Figure 2 A, B). S-type granite has the lowest OH content. - The median values ​​for H2O in the training and global datasets were 1.4 (apfu×1000) and 2.7 (apfu×1000), and 2.0 (apfu×1000) and 3.0 (apfu×1000), respectively. Figure 2 A, B). In Type A and Type I granites, over 70% of zircon has an OH content of [missing information]. - / H + >4 ( Figure 2 C, D) indicate that hydrogarnet-type substitution is the main mechanism for H2O absorption, OH - It is the primary form in which water exists. And S-type granite has even lower OH content. - / H + In the training dataset and the global dataset, OH - / H + The values ​​are even lower, with median values ​​of 0.5 and 0.11 respectively, and cannot effectively reflect the H2O content in zircon.

[0034] In the description of this invention, the term "a plurality of" refers to two or more. Unless otherwise explicitly defined, the terms "upper," "lower," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. The terms "connection," "installation," "fixing," etc., should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a direct connection or an indirect connection through an intermediate medium. For those skilled in the art, the specific meaning of the above terms in this invention can be understood according to the specific circumstances.

[0035] In the description of this specification, the terms "one embodiment," "some embodiments," "specific embodiment," etc., refer to a specific feature, structure, material, or characteristic described in connection with that embodiment or example, which is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0036] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for calculating the water content in zircon using machine learning modeling, characterized in that, Specifically, the following steps are included: Step S1: Construction of the Zircon Database: Two datasets are compiled: a training dataset and a global dataset. The training dataset is used to train the zircon H2O content calculation model, while the global dataset is used to explore the variation pattern of zircon water content from a larger scale of data. The training dataset contains the zircon H2O content obtained from test analysis as the target value and 29 features, while the global dataset only contains 29 features. The target value of H2O content needs to be obtained through the calculation model. Step S2: The compiled training dataset is preprocessed by Python code. The data preprocessing process for AIS type granite zircon is the same, which includes outlier cleaning, missing value imputation, and data standardization. Step S3, Classification Model Construction and Selection: Perform regression analysis on the processed training dataset, using H2O as the target value and 27 elements and calculated values ​​as features; the coefficient of determination (R²) is used. 2 R is used to measure the performance of the evaluation model (Equation 7); 2 The closer the value is to 1, the better the model performs; R 2 = (Formula 7) y i This is the actual value. is the mean of the actual values, and n is the total number of data points; Nine machine learning strategies were compared: linear regression (0.47), Lasso (0.46), support vector machine (-0.06), K-nearest neighbors (0.53), artificial neural network (0.12), random forest (0.63), gradient boosting decision tree (GBDT) (0.60), XGBoost (0.56) and LightGBM (0.54). Step S4: Identification of Important Features and Model Hyperparameter Optimization: Based on the selection of Random Forest as the computational model, RFE (Reference-Free Evaluator) is used to screen the feature importance of the model; for the zircon H2O content calculation model, Age, Nb, Ce, Sm, Lu, 10 4 ×(Eu / Eu*) / Yb N Dose (×10) 15 Yb / Gd were selected as the 8 most important features by RFE screening; The random forest model is simplified using eight key features and employs Bayesian optimization to find the optimal combination of hyperparameters. Thus, based on the training dataset, a random forest calculation model for zircon water content with 8 features was obtained. The model's R... 2 =0.66; H2O RF The values ​​were then fitted again to the actual water content H2O to obtain the final predicted values ​​of zircon water content H2O* and R. 2 To avoid being affected by large deviations in H2O ML Value interference, ignore |H2O- H2O RF Discrete values ​​of | / H2O≥0.3; H2O* = 0.9993 H2O RF 1.0396 (R 2 = 0.957) (Official 11) Step S5: Application of the Zircon Water Content Calculation Model The compiled global dataset containing 8 features is imported into the random forest model and Formula 11 built based on the training dataset to calculate the zircon water content in the global dataset; the model calculates the zircon water content based on Dose (×10). 15 The LREE-I filter removes data that is disturbed by inclusions and uses random forest and power function fitting strategies to give the water content of zircon. For the global dataset, a correction factor α is needed to obtain a more accurate predicted value of zircon water content H2O*. α = 1 + (Median of H2O in the training dataset - H2O of each granite class in the global dataset) # (median) / (H2O per type of granite in the global dataset) # (The median) (Formula 12) H2O* = a x H2O # (Formula 13) Step S6, combine zircon OH - Determination of available zircon water content data Water in zircon exists in three forms: molecular water (H₂O), structural water (H₂O), and water in zircon. + OH - H + With (Y+REE) 3+ Coupling replaces Zr 4+ or Si 4 + It becomes part of the zircon structure (Formula 14), OH - The zircon is replaced by a hydrogarnet-type ion (Formulas 15 and 16), and only OH is present in the zircon. - It can accurately record the water activity of the host magma; H+ + (REE+Y) 3+ = Zr 4+ (Equation 14) H + + Al 3+ = Si 4+ (Formula 15) (Official 16).

2. The method for calculating the water content in zircon using machine learning modeling according to claim 1, characterized in that, Step S1 specifically includes the following steps: Step S1-1: The training dataset contains a total of 1418 zircon analyses, of which 399 are from the Himalayan orogenic belt, 674 are from the North China Craton, 218 are from South China, 20 are from the Czech Republic, and 107 are from the New England Orogen in Australia; including 129, 985, and 304 zircon data from A-type, I-type, and S-type granites, respectively. The target value H2O in the training dataset was measured using SIMS; the training dataset contains 29 features, namely: zircon. 206 Pb / 238 U-age, labeled as Age; Trace elements in 18 zircon samples: Hf, U, Th, Y, P, Ce, Nb, Pr, Sm, Eu, Gd, Tb, Dy, Ho, Er, Tm, Yb, Lu; 10 calculated values: Eu / Eu*, Ce / Ce*, 10 4 ×(Eu / Eu*) / Yb N , 10 4 ×Ce / Nd / Y, U / Yb, Dose (×10 15 ), LREE-I, Yb / Gd, Ce / Sm, Th / U; where 206 Pb / 238 U-age and trace elements were obtained simultaneously by LA-ICP-MS, after water testing, and at the same location as the water testing points; The calculated values ​​selected indicators that can determine zircon crystal structure and indicate magma water content include LREE-I and Dose (×10). 15 ), 10 4 ×Ce / Nd / Y, 10 4 ×(Eu / Eu*) / Yb N and U / Yb; LREE-I and Dose (×10) 15 It is also an indicator for determining whether zircon has undergone hydrothermal alteration and metamorphism; 10 4 ×Ce / Nd / Y, 10 4 ×(Eu / Eu*) / Yb N U / Yb are indicators reflecting the content of magmatic water; Steps S1-2: The global dataset contains a total of 4196 zircon data points, with 29 features identical to those in the training dataset; 1028 data points are A-type granites, 2409 are I-type granites, and 759 are S-type granites; all I-type granites are from arc-magmatic environments, including adakitic granites (labeled as rich granites) and low-mineral calc-alkaline granites (labeled as low-mineral granites), with 1472 and 937 zircon trace element data points respectively.

3. The method for calculating the water content in zircon using machine learning modeling according to claim 2, characterized in that, The LREE-I value in step S1-1 refers to the light rare earth element index, which represents the relative abundance of light rare earth elements to medium rare earth elements. The specific formula is as follows: LREE-I = (Dy / Nd) + (Dy / Sm) (Formula 1) Dose (x 10 15 ) values refer to the alpha-decay dose of radioactive elements in zircon, indicating the degree of self-irradiation damage in zircon. The specific formula is as follows: D = 8 N 1[exp(a1t) - l] + 7 N 2[exp(a2t) - l] + 6 N 3[exp(a3t) - l] (Formula 2) In the formula, D is the α-decay radiation dose, expressed in events / mg. N 1, N 2 and N 3 represents the current quantities of 238U, 235U, and 232Th, respectively, and a1, a2, and a3 represent... 238 U、 235 U and 232 The decay constant of Th, where t is the age of zircon; The Eu / Eu* value is calculated using the following formula: Eu / Eu* = (Eu / 0.0563) / [(Sm / 0.148) x (Gd / 0.199)] 1 / 2 (Equation 3) The Ce / Ce* value is calculated using the following formula: Ce / Ce* = (Ce / 0.613) / [(Nd / 0.457) 2 / (Sm / 0.148)] (Equation 4) 10 4 x(Eu / Eu*) / Yb N values, which reflect the melt water content level with zircon composition, as follows: 10 4 ×(Eu / Eu*) / Yb N = 10 4 ×(Eu / Eu*) / (Yb / 0.161) (Formula 5).

4. The method for calculating the water content in zircon using machine learning modeling according to claim 1, characterized in that, Step S2 specifically includes the following steps: Step S2-1: Outlier cleanup; La and Pr elements are removed from zircon, P>10 4 Analysis of ppm was omitted, LREE-I <10 and Dose (×10) 15 Zircon analysis results for values ​​less than 3 were rejected. Step S2-2: The missing values ​​are filled with the mean; Step S2-3: Data Standardization; Data standardization uses z-score, which calculates the difference between each data point and the overall data mean, and measures it in units of standard deviation. Z = (X−μ) / σ (Formula 6) In the formula, Z is the standardized value Z-score, X is the original data value, μ is the mean of the dataset, and σ is the standard deviation of the dataset.

5. The method for calculating the water content in zircon using machine learning modeling according to claim 1, characterized in that, The linear regression model formula in step S3 is as follows, where w is the weight coefficient vector: y = w0 + w1x1 + w2x2 + ... + w n x n + ϵ (Official 8) Lasso is a model that adds an L1 regularization term to the linear regression loss function; L1 regularization refers to the sum of the absolute values ​​of the weight coefficient vector w; the loss function formula is as follows, where ||w||1 = ∑|w j | represents the L1 norm, and λ is a hyperparameter that controls the severity of the penalty; (Official 9) Support Vector Machines use kernel tricks to map low-dimensional linearly inseparable data to a high-dimensional space, making it linearly separable in the high-dimensional space, thereby handling complex nonlinear decision boundaries. When making predictions using K-nearest neighbors, for a new sample, the distance to all samples in the dataset is calculated, and then the K nearest neighbors are found. For regression tasks, the average of the target values ​​of the K neighbors is taken. The formula is as follows: Nk(x) is the set of indices of the K nearest neighbors of x, and mode is the mode. y = mode ({ yi | i ∈ N k ( x )}) (Official 10) Random forests work by constructing multiple decision trees and combining their predictions; during training, each tree not only performs random sampling with replacement from the original dataset, but also randomly selects a subset of features from all features for optimal splitting at each node.