Method for reconstructing paleo-elevation of mountain-building zone based on paleontological data and geochemical data
Through machine learning methods combining paleontological data and geochemical data in the orogenic zone, the problems of insufficient sample size and inference uncertainty are solved, and efficient and accurate paleoelevation reconstruction is achieved.
Patent Information
- Application Number
- CN202510436355.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The prior art has problems of insufficient sample size, inference uncertainty and inefficiency in paleoelevation reconstruction in orogenic zones.
Using machine learning methods based on paleobiological data and geochemical data, efficient and accurate paleoelevation reconstruction is achieved through data preprocessing, feature extraction, model construction and optimization.
It improves the accuracy and efficiency of paleoelevation prediction, reduces errors in the speculation process, has good generalization capabilities, and is suitable for different geographical regions and geological periods.
Smart Images

Figure CN119942019A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of orogenic belt elevation reconstruction, and in particular to a method for reconstructing orogenic belt regional paleo-elevation based on paleontological data and geochemical data. Background Art
[0002] As an important indicator of the evolution of the earth's surface, paleoelevation is of great significance for studying the formation and development process of orogenic belts. For example, by analyzing the chemical composition of igneous rocks and the depth of the Moho surface, the crust thickness and terrain height of the ancient orogenic belt can be inferred.
[0003] At present, the inference of paleo-elevation mainly relies on geochemical analysis of geological samples, and its commonly used methods include: 1) Whole-rock Moho: This method is based on the relationship between the chemical composition (major elements and trace elements) of igneous rocks and the thickness of the crust to estimate the thickness of the crust. For example, by analyzing the composition of igneous rocks in the Bend Mountain structure, a quantitative model that can be applied to ancient orogenic belts is calibrated. There are mainly Moho meters based on the Ce / Y ratio and Moho meters based on the Sr / Y and La / Yb ratios. 2) Zircon Moho: This method is based on the trace element and isotope ratios of zircon to estimate the crustal thickness. Zircon is a common accessory mineral whose chemical composition can reflect the source and evolution of magma. The zircon Moho based on the La / Yb ratio was used to reconstruct the crustal thickness evolution of the Gangdise Mountains in the southern Qinghai-Tibet Plateau. The results are consistent with the regional geology, but the La concentration in zircon is low and the measurement error may be large. The zircon Moho based on Eu anomaly was used to reconstruct the crustal thickness evolution of the Gangdise Mountains in the southern Qinghai-Tibet Plateau. The concentrations of elements such as Eu, Sm and Gd in zircon are high, and the measurement accuracy is high, making this method more reliable than the La / Yb ratio in some cases.
[0004] 3) Radioactive isotope method: The radioactive isotope method is a method of estimating crustal thickness based on the Sr and Nd isotope ratios of igneous rocks. This method assumes that the isotopic composition of igneous rocks reflects the mixing process of the crust and mantle. The Nd isotope method is used to reconstruct the crust thickness evolution in the southern Qinghai-Tibet Plateau. The results are highly accurate. The radioactive isotope method can provide independent information about crustal thickness, especially when the chemical composition changes are not obvious.
[0005] Although these methods have been widely used in the field of geology, they still have the following problems: 1) Insufficient sample size: Due to the particularity of orogenic belts, some geochemical proxy indicators are missing in orogenic belts. For example, basic igneous rocks are relatively rare in thick continental arcs and cannot be inferred using the Ce / Y-based Moho meter. 2) Uncertainty of inference: The inference of paleo-elevation is related to many factors. In orogenic belts, due to the large amplitude of depth variation, it is uncertain whether certain chemical data are positively or negatively correlated with the data to be inferred; 3) Inefficiency: Traditional experimental analysis methods are time-consuming and costly, and cannot meet the needs of large-scale regional reconstruction. Summary of the invention
[0006] In view of the problems existing in the prior art, the present invention provides a method for reconstructing paleoelevation in an orogenic belt region based on paleontological data and geochemical data. On the basis of ensuring a geochemical Moho meter of sufficient accuracy, paleoelevation data are revised by introducing paleontological data to improve the accuracy of paleoelevation prediction, thereby achieving efficient, accurate and generalizable paleoelevation reconstruction.
[0007] The present invention provides a method for reconstructing the paleo-elevation of an orogenic belt region based on paleontological data and geochemical data, comprising: Data preparation: Collect known geochemical data, elevation data and paleontological data in the orogenic belt and pre-process the collected data, wherein the geochemical data include trace element composition and isotopes in rock, sediment and fossil samples; the elevation data include modern elevation data directly measured by modern geological measurement technology and ancient elevation data previously recorded or inferred by geological methods; the paleontological data include species information, age information and location information of paleontological organisms; Feature extraction: Analyze the correlation between different geochemical data and elevation changes, and select geochemical feature data from geochemical data according to the degree of correlation between different geochemical data and elevation changes; Elevation annotation: Revise the elevation data of the orogenic belt area based on paleontological data; Model construction: Input the screened geochemical characteristic data and the revised elevation data into the constructed machine learning prediction model to learn the nonlinear mapping relationship between geochemical data and elevation; Model evaluation and optimization: Evaluate the accuracy and performance of the elevation prediction model and further optimize the model based on the evaluation results; Application prediction: Input the geochemical sample data of the orogenic belt area to be predicted into the trained elevation prediction model, output the paleo-elevation value of the orogenic belt area to be predicted and display it visually.
[0008] Optionally, the preprocessing of the collected data includes: Select interpolation methods based on data type to fill in missing geochemical characteristic data; Use statistical methods to identify and remove outliers in geochemical signature data and elevation data; The Min-Max method was used to scale the cleaned data to a preset range, and the Z-Score method was used to convert the data into a distribution with a mean of 0 and a standard deviation of 1.
[0009] The analysis of the correlation between different geochemical data and elevation changes, and screening of geochemical characteristic data from the geochemical data according to the degree of correlation between different geochemical data and elevation changes, includes: The theoretical correlation between different geochemical data and elevation changes is analyzed through literature, and the quantitative indicators of the correlation between different geochemical data and elevation changes are obtained through statistical analysis. The geochemical data that has both theoretical correlation with elevation and quantitative indicators reaching the preset threshold are selected as geochemical characteristic data; Ratio features are constructed by calculating the ratios between chemical elements. For those geochemical data that have a significant nonlinear relationship with elevation, polynomial transformation is applied to expand them into quadratic or cubic terms, and principal component analysis is used to extract the most representative principal components of the geochemical data.
[0010] Optionally, the revising of elevation data of the orogenic belt region according to paleontological data includes: Select paleontological data that can represent paleo-elevation information based on the age, distribution, and altitude range of paleontological fossils; According to the age information and location information in the selected paleontological data combined with the location of paleontological fossils, the specific paleo-elevations at different locations in the orogenic belt area in a specific period are estimated; The specific paleoelevations of the orogenic belt area are used to revise the collected paleoelevation data of the orogenic belt area.
[0011] Optionally, evaluating the accuracy and performance of the elevation prediction model and further optimizing the model according to the evaluation results include: The R² value, root mean square error RMSE and mean absolute error MAE are used to comprehensively judge the accuracy and performance of the elevation prediction model; If the elevation prediction model does not achieve the set accuracy and performance, a variety of optimization strategies are implemented: introducing penalty terms through regularization methods to control model complexity, generating samples with small perturbations through data enhancement technology, expanding the training set, and introducing the prediction results of multiple models through ensemble learning to improve the overall prediction performance.
[0012] After adopting the above technical solution, the present invention has at least the following beneficial effects: 1. Accuracy and efficiency The present invention utilizes the precise records of the altitude survival range of ancient organisms in specific periods in paleontological data, combined with the inference capability of geochemical data, to accurately infer the paleoelevation of a certain geographical location in the past. Traditional methods often rely on a single data source or rough assumptions, while the present invention provides a more accurate and efficient paleoelevation inference method by integrating multiple data sources.
[0013] 2. Data diversity and wide applicability The annotation method and model building method adopted by the present invention can not only process large-scale paleontological data, but also effectively combine a variety of geochemical data. It has strong adaptability and versatility. Whether in different geographical regions or in different geological periods, the method of the present invention can ensure the wide applicability of the inference results and avoid regional and temporal limitations.
[0014] 3. Avoid the errors of traditional inference methods Traditional ancient elevation inference often relies on indirect inference or hypothesis, which has large errors and uncertainties. The present invention greatly reduces the errors in the inference process and improves the accuracy and credibility of the inference results by directly utilizing the actual data of the living altitude range of ancient organisms in a specific period.
[0015] 4. Efficient data processing and automation The present invention combines machine learning algorithms and automated annotation tools to efficiently process and analyze large-scale paleontological and geochemical data. In the process of data preprocessing, feature selection, model training and optimization, automated tools can greatly improve data processing efficiency, save a lot of manual labor costs, and ensure the consistency and accuracy of the processing process.
[0016] 5. Provide accurate basis for paleoenvironmental research By reconstructing the ancient elevation of the orogenic belt area in the historical period, the present invention can provide a more accurate basis for paleoenvironmental reconstruction, climate change research and geological history analysis. Researchers can reconstruct the past geographical and climatic environment based on accurate paleoelevation data, and then infer important information such as paleoecological changes and species distribution, providing strong data support for research in related fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0018] Figure 1A schematic diagram of a process for reconstructing the paleoelevation of an orogenic belt region based on paleontological data and geochemical data provided in an embodiment of the present disclosure; Figure 2 This is a scatter plot comparing the model's predicted paleoelevation values with the actual values. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] At present, the inference of paleo-elevation mainly relies on geochemical analysis of geological samples, and its commonly used methods include: Stable isotope analysis: e.g. δ 18 O and δD are used to reconstruct the source areas of ancient temperature or precipitation. In orogenic belts, the systematic changes in oxygen isotope composition can reveal the spatial variation process of orogenic belts and plateau uplift. Cluster isotope analysis: The carbonate cluster isotopes in soil (such as ) can be used to calculate the temperature at which calcite minerals are formed, and then combined with the δ 18 Paleoelevation is calculated using O values and climate models; Analysis of the chemical composition of sediments and rocks: such as the changes in element ratios (Sr / Ca, Mg / Ca) at different elevations. By sampling and experimentally analyzing soil carbonates and sediments, carbonate rocks and biofossils, the Sr / Ca and Mg / Ca ratios are combined with other climate indicators to restore ancient elevations.
[0021] In addition, the thickness of the crust at different ages can be estimated through geochemical and isotopic data of igneous rocks, which is the so-called "chemical Moho meter". This method identifies the correlation between certain chemical parameters and crust thickness and uses these parameters as proxy indicators of crust thickness. This method is particularly useful in studying ancient orogenic belts because direct seismic or geophysical data are not available in these areas. So far, a variety of chemical proxy indicators have been identified and calibrated as indicators of crust thickness, such as Ce / Y ratio, Sr / Y and La / Yb ratio, etc. These proxy indicators come from the chemical composition of volcanic arc magma. The depth of magma generation is closely related to the depth of the Moho surface, and the depth is related to the interaction process between the crust and the mantle.
[0022] In addition, sedimentary zircons can also be used to estimate crustal thickness. Zircons are particularly important for reconstructing the evolution of continental crust thickness because they are very well preserved in sedimentary rocks and can carry information about the original magma source. Studies have shown that zircon-based Moho meters can provide new perspectives on the evolution of crust in ancient orogenic belts. However, this method also has certain limitations, such as the difficulty in distinguishing whether the source of zircons comes from different tectonic environments, and that there may be certain errors in the measurement of trace elements in zircons.
[0023] Based on the above existing research status, the defects of different methods are summarized as follows: Whole Rock Chemical Moho Meter: 1. Sample size limitations: The whole-rock chemical Moho method relies on large-scale geochemical data sets, but in some areas, especially ancient or eroded orogenic belts, it may be very difficult to obtain sufficient samples, resulting in scarce or insufficient data. These limitations may affect the accuracy of crustal thickness estimates.
[0024] 2. Regional adaptability differences: The relationship between magma composition and crustal thickness may be different in different geological environments (such as subduction zones, collision zones, etc.). Therefore, some methods may work well in some areas but may not be applicable in other areas. For example, whole-rock geochemical parameters based on igneous rocks (such as Sr / Y ratio) have limited adaptability to different tectonic environments and may lead to inaccurate crustal thickness estimates in some environments.
[0025] Zircon-based Moho: 1. Limitations of zircon sample sources: Although zircon is well preserved in sedimentary rocks, its source tends to be in tectonic environments such as orogenic belts and collision zones. Analysis of zircons originating from severely eroded orogenic belts or low-altitude areas may lead to errors. Therefore, the zircon-based Moho meter may over-estimate the thickness of the ancient crust in some areas.
[0026] 2. Challenges of analytical technology: Trace element and isotope measurements of zircon usually require high-precision instruments, such as laser profile mass spectrometry (LA-ICP-MS). These technologies are not only expensive, but may also have systematic errors between different instruments and laboratories. The quality of the data directly affects the accuracy of the estimate of the crust thickness.
[0027] 3. The correlation between zircon chemistry and magma source is complex: The chemical composition of zircon may be affected by many factors, such as different magma sources, local crustal contamination, etc., which makes it complicated to directly correlate the chemical composition of zircon with crustal thickness.
[0028] The inventors analyzed the shortcomings of the current method and proposed the following improvement directions after further research: 1. A paleoelevation prediction method based on paleontological information annotation data is proposed, which provides a new idea for the current paleoelevation prediction method. This method uses data-driven modeling and the actual information of paleontological life strata at different time points to assist in building a prediction model, which can automatically capture the nonlinear relationship between complex geochemical data and elevation.
[0029] 2. Integrating the life information of paleontology with the existing chemical Moho meter improves the credibility of the prediction model, further improves and expands the accuracy and applicability of the prediction model, and provides a data-driven and efficient paleoelevation prediction method. This method has the advantages of high efficiency and automation, can process geological data from multiple regions and eras, and is widely applicable to different types of geological research.
[0030] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0031] like Figure 1 As shown, the embodiment of the present disclosure provides a method for reconstructing the paleo-elevation of an orogenic belt region based on paleontological data and geochemical data, comprising: 1. Data preparation In order to ensure the accuracy and reliability of the model, the data preparation stage is crucial. High-quality input data sets are the basis for subsequent feature extraction and model training. The following are detailed data sources and preprocessing steps: 1.1 Data Source Geochemical data is a core component and mainly comes from the following categories: Stable isotope data (δ 18 O, δD, etc.): These two indicators reflect the source of water, precipitation and climate conditions and are widely used in geological and climatological research. 18 O data reflect the precipitation characteristics in different geographical regions and under different climatic conditions, while δD can provide temperature information of water vapor source areas.
[0032] Cluster isotopes ( The isotope data can be used to estimate the formation temperature of rocks or sediments and indirectly reflect the elevation changes during geological periods. The data can be used to infer the climate and environmental conditions of the samples in different eras, and then calculate the ancient elevation when they were formed.
[0033] Sediment and rock chemistry (Sr / Ca and Mg / Ca ratios, etc.): These ratios can provide information on the chemical composition of rocks or sediments, further enriching the data dimension and thus increasing the accuracy of elevation inference.
[0034] Igneous rock chemistry (Ce / Y ratio, Sr / Y ratio, etc.): Identify correlations with crustal thickness, especially in areas where direct seismic data are unavailable, and use these chemical parameters as proxies for crustal thickness.
[0035] Chemical composition of sedimentary zircon: The thickness of the earth's crust is estimated through trace element and isotope data in zircon, which is particularly suitable for reconstructing the evolution of the thickness of the continental crust.
[0036] Elevation data is used for comparison and correlation with paleontological data and comes mainly from the following categories: Modern elevation data: In some known areas, the elevation data obtained through modern measurement technology provides an accurate reference for the model. Modern elevation data is used as supervisory information in the model to help train a more accurate regression model.
[0037] Ancient elevation data: Since ancient elevation is difficult to obtain through direct measurement, it is often necessary to rely on literature, geological surveys and inference methods for correction. The ancient elevation data reported in the literature and the ancient elevation records calculated by geological methods can provide support for the training data set.
[0038] Select to download the paleontological data of the corresponding area in The Paleobiology Database. The data includes key information such as the type, age, location (including longitude and latitude and corresponding altitude range information) of the paleontology, and perform preliminary processing to prepare for the subsequent manual labeling steps.
[0039] 1.2 Data Preprocessing Step 1: Missing value handling Select interpolation methods according to data types to fill in missing geochemical characteristic data, such as mean interpolation, Lagrange interpolation, etc., to ensure data integrity; Step 2: Data cleaning Use statistical methods to identify and remove outliers in geochemical characteristic data and elevation data to ensure the quality of the data set. For example, in elevation data, abnormal elevation changes caused by earthquakes or human activities may appear and need to be corrected; Step 3: Data standardization and normalization In order to adapt to the input requirements of the machine learning model, the Min-Max method is used to scale the cleaned data to a preset range to eliminate the impact of different feature dimensions, which helps prevent certain features from having an excessive impact on the results due to their large numerical range during model training. The Z-Score method is used to convert the data into a distribution with a mean of 0 and a standard deviation of 1, which helps to process data with different distribution characteristics and improve the convergence speed and stability of the model.
[0040] 2. Feature extraction 2.1 Feature Selection Feature selection evaluates the correlation between geochemical data and elevation changes through literature analysis and correlation analysis to ensure that the selected features have both theoretical basis and statistically quantifiable correlation with paleoelevation. Geochemical data that have both theoretical correlation with elevation and quantitative indicators that reach the preset threshold are selected as geochemical feature data, such as δ 18 O and As an indicator substance known to be closely related to temperature and climate conditions, it can effectively assist in the inference of paleo-elevation, further eliminate redundant geochemical data and geochemical data with low correlation with elevation, such as recursive feature elimination (RFE) and Lasso regression, reduce data dimensions, and significantly improve the efficiency of model training and the accuracy of prediction; 2.2 Feature Transformation Ratio features are constructed by calculating the ratios between chemical elements (such as Sr / Ca and Mg / Ca) to extract more geochemical information. These ratio features can capture the relative relationship between different chemical elements and help to better understand the impact of mineral composition on elevation changes. For geochemical data that have a significant nonlinear relationship with elevation, polynomial transformation is applied to expand them into quadratic or cubic terms to facilitate the capture of complex patterns, thereby enhancing the model's ability to model nonlinear relationships. Principal component analysis is used to extract the most representative principal components of geochemical data to reduce redundancy and multicollinearity problems between features, effectively alleviate the curse of dimensionality, and improve the efficiency and stability of model training.
[0041] 3. Elevation marking Based on geological data and paleontological analysis, combined with the regression analysis results of the machine learning model, the elevation annotation is accurately corrected. By combining the positioning characteristics of paleontological data (such as species distribution, fossil remains layer, etc.) with geochemical data (such as sediment composition, isotope ratio, etc.), the predicted paleo-elevation values are fine-tuned using manual annotation methods to ensure that the inferred paleo-elevation values are more consistent with the actual geological background.
[0042] 3.1 Labeling strategy With paleontological data as the core, paleontological data that can represent paleoelevation information are selected according to the age, distribution and survival altitude range of paleontological fossils. The specific paleoelevations of different locations in the orogenic belt area in a specific period are estimated based on the age information and location information in the selected paleontological data combined with the location where the paleontological fossils were found. The specific paleoelevations of the orogenic belt area are used to revise the collected paleoelevation data of the orogenic belt area. For example, a certain paleontological species only lives within a specific altitude range. Combined with the location where its fossils were found, the paleoelevation of the area in that period can be determined. This method can accurately reflect the specific elevation information of a certain location in the past period. With the help of these data, the paleoelevation data can be revised more accurately and quantitatively. This annotation strategy is widely representative and can cover the diversity of different geographical regions and geological periods, avoiding annotation deviations. In addition, this annotation strategy is operational, ensuring that the annotation process is clear and easy to understand, convenient for practical application, and improving the consistency and accuracy of annotation.
[0043] 3.2 Annotation Execution The annotation execution phase is led by geological experts, who infer the paleo-elevation of the sample site based on the altitude survival range of paleontology and geological history, combined with scientific speculation and models, and combine it with the paleo-elevation data of the orogenic belt area collected by automated annotation tools.
[0044] 3.3 Labeling quality control In order to ensure the accuracy and consistency of the labeled data, strict quality control methods are adopted, mainly including the following two means: 1. Multiple annotations: Samples in the same data set will be independently annotated by multiple experts, and finally the consistency of annotations will be improved by comparing different annotation results. Multiple annotations can effectively reduce the bias of a single expert's judgment and improve the reliability of data labels.
[0045] 2. Cross-validation: By comparing the annotation results with known real data, cross-validating the scientificity and accuracy of the annotations, and comparing them with existing elevation data and geological research results, it ensures that the label of each sample has a scientific basis, thereby improving the prediction performance of the model.
[0046] 4. Model Construction During the model building phase, a machine learning algorithm is selected to train the model to learn the nonlinear mapping relationship between geochemical data and elevation.
[0047] 4.1 Algorithm Selection Advanced machine learning algorithms, such as LightGBM and random forest algorithms, are used to construct regression models between geochemical data and paleoelevation based on paleontological data. For large-scale data sets and complex regression problems, the LightGBM algorithm has outstanding performance advantages. Its excellent performance and strong interpretability can provide faster training speed and higher prediction accuracy than traditional algorithms. For small and medium-sized data sets, the random forest algorithm is selected because of its ability to efficiently process multidimensional data and its feature importance analysis function, which helps to identify and select the most effective feature combinations.
[0048] 4.2 Model Training During the model training process, the efficiency and high accuracy of the model were ensured by rationally selecting algorithms, scientifically dividing data sets, accurately evaluating errors, and optimizing hyperparameters. To prevent overfitting, the data set was divided into training and test sets in a ratio of 70:30 to ensure that the model can maintain good generalization capabilities on unseen data; in order to effectively evaluate the prediction accuracy of the model, mean square error (MSE) and root mean square error (RMSE) were selected as loss functions. These indicators can accurately reflect the error between the prediction results and the actual values; in addition, in order to further improve the performance of the model, hyperparameter optimization methods were used, including grid search, random search, and Bayesian optimization. These methods systematically adjust model parameters to find the optimal parameter combination, thereby significantly improving the prediction ability and stability of the model.
[0049] 5. Model evaluation and optimization After the model is built, comprehensive evaluation and optimization are key steps to ensure that the model has good predictive ability and stability. The model evaluation and optimization stage mainly includes the selection and application of evaluation indicators and the implementation of various optimization strategies.
[0050] 5.1 Selection and application of evaluation indicators In order to comprehensively measure the performance of the model, a variety of evaluation indicators are used to evaluate the fitting ability and prediction accuracy of the model from different angles. The evaluation indicators include R² value, RMSE (root mean square error) and MAE (mean absolute error). The R² value is used to measure the fitting ability of the model, and its value range is 0 to 1. The closer the R² value is to 1, the stronger the model's ability to explain the target variable and the more accurate the prediction result. RMSE and MAE are also commonly used indicators for evaluating the prediction accuracy and stability of the model. RMSE pays more attention to the penalty of larger errors, while MAE gives the same weight to all errors. By calculating the absolute difference between the predicted value and the actual value, it provides an average level of error. Combining these indicators, the performance of the model at different error scales can be comprehensively judged to ensure that the model has both high accuracy and stability in the prediction task.
[0051] 5.2 Model Optimization After the initial evaluation of the model performance, in order to further improve the model's predictive ability and generalization performance, a variety of optimization strategies are implemented: the penalty term is introduced through the regularization method to control the model complexity, thereby preventing the model from overfitting. Common regularization methods include L1 regularization and L2 regularization. L1 regularization helps feature selection and promotes the generation of sparse models by introducing a penalty for the absolute value of the weight, that is, reducing the influence of irrelevant features. L2 regularization reduces excessive fluctuations in model parameters by introducing a penalty for the square of the weight, avoiding overfitting of a few abnormal data points. Appropriate regularization can help reduce the interference of noise in the training data on the model's predictive ability and improve the model's generalization ability; generate specific regularization data through data enhancement technology. There are samples with small disturbances, and the training set is expanded to simulate more diverse geological conditions and input data distributions. Data augmentation can not only increase the amount of training data, but also improve the model's adaptability to different types of data, especially when the training data samples are relatively limited. In this way, the model can learn a wider range of features, thereby improving its prediction performance on unknown data; through integrated learning, the prediction results of multiple models are introduced to improve the overall prediction performance, including voting method, weighted average method and stacking method. For example, combining multiple algorithms such as random forest and LightGBM can give full play to the advantages of different models and reduce the bias and overfitting problems of a single model, so as to obtain more accurate and robust prediction results.
[0052] 6. Application prediction After completing the model construction and optimization, the geochemical sample data of the orogenic belt area to be predicted is input into the trained elevation prediction model, and the paleo-elevation value of the orogenic belt area to be predicted is output and visualized. The specific operation steps are as follows: Step 1: Enter geochemical data The geochemical sample data of the orogenic belt to be predicted, including δ 18 O, δD and And other key indicators, input the trained elevation prediction model, start the model's calculation process, and begin to infer the ancient elevation value corresponding to the sample.
[0053] Step 2: Output prediction results After receiving the input geochemical sample data, the elevation prediction model infers the paleo-elevation value corresponding to the sample. The prediction result is output in numerical form, indicating the paleo-elevation of the target area in a specific geological period. The output paleo-elevation value provides important reference data for geological researchers, helping them to further analyze and reconstruct paleo-elevation change trends.
[0054] Step 3: Visualize the results In order to facilitate geological researchers to conduct in-depth analysis and interpretation of the prediction results, the predicted paleoelevation values will be displayed through various visualization tools to help researchers better understand the paleoelevation characteristics of different regions and conduct comprehensive analysis with other geological data, thereby providing strong support for geochemical research, paleoenvironmental reconstruction, etc.
[0055] The following specific examples are proposed in combination with the above embodiments. It can be understood that the following specific examples are only illustrative of the specific implementation of the above embodiments, and are not intended to limit the technical solutions of the above embodiments.
[0056] 1. Data Source and Sample Composition Integrates 1,195 sets of sample data, covering the following multi-source data sets: Geochemical data: including stable isotopes (δ 18 O, δD), cluster isotopes ( ), igneous rock element ratios (Sr / Y, La / Yb), etc., are derived from the geological survey database of the global orogenic belt region; Paleontological data: Extract fossil distribution information from ThePaleobiologyDatabase, including species survival time, geographic coordinates and corresponding altitude range, as the basis for elevation label correction; Modern elevation data: SRTM digital elevation model (DEM) is used as benchmark verification data; 2. Data Preprocessing 2.1 Data cleaning and missing value processing Missing value filling: interpolation methods are used to fill in missing geochemical features. For the processing of missing values, interpolation methods (such as mean interpolation, Lagrange interpolation, etc.) are selected according to the data type to ensure data integrity; Outlier removal: Identify and remove outliers in geochemical parameters (such as extreme isotope ratios) based on the 3σ principle; 2.2 Data Standardization and Coding Numerical characteristics: Min-Max normalization is used to scale geochemical parameters to the [0,1] interval to eliminate dimensional differences; 3. Feature Engineering Feature concatenation: concatenates numerical features (such as δ 18 O, ancient longitude and latitude, , Sr / Ca and Mg / Ca), the label-encoded features are stacked into a joint feature matrix (total dimension: 215) as the model input; 4. Model construction and training 4.1 Algorithm selection and parameter setting LightGBM regressor: suitable for large-scale data, set the learning rate to 0.05, the tree depth to 8, the number of leaf nodes to 64, and the early stopping rounds to 50; Random Forest Regressor: used for comparative experiments, with the number of trees set to 200, the maximum depth to 10, and the minimum number of leaf samples to 5; 4.2 Dataset Division and Training The training set (837 samples) and the test set (358 samples) were divided into a ratio of 7:3, and 5-fold cross validation was used to optimize the hyperparameters; 5. Model Evaluation and Results 5.1 Performance Indicators Training set R²(train): 0.966.
[0057] Test set R²(test): 0.816.
[0058] Root mean square error (RMSE): 21.3m.
[0059] R² (coefficient of determination) measures the explanatory power of the model. The closer the value is to 1, the stronger the predictive power of the model. The root mean square error (RMSE) measures the degree of difference between the model's predicted value and the true value.
[0060] 5.2 Results Visualization The results are as follows Figure 2 As shown in the figure, the horizontal axis is the model prediction value, the vertical axis is the actual elevation value, the solid line is the 1:1 ideal fitting line, the triangulated points are the training set samples, and the circular points are the test set samples. Both the triangulated points and the circular points are densely distributed along the 1:1 ideal fitting line, indicating that the model has good generalization ability in both the training set and the test set.
[0061] Compared with traditional methods, the present invention has significant advantages in many aspects: first, the machine learning model based on paleontological and geochemical data can efficiently complete the paleoelevation reconstruction of large-scale orogenic belt areas, solving the problems of scarce samples and low experimental efficiency of traditional methods; second, machine learning can automatically perform nonlinear fitting, overcoming the difficulty that traditional empirical formulas are difficult to fully describe the coupling relationship of complex factors; in addition, the model has strong generalization ability, can adapt to the application needs of different regions and geological eras, and significantly improves the prediction accuracy and reliability; finally, in traditional methods, many annotation and inference processes often rely on manual experience and have a certain degree of subjectivity, while paleontological data can accurately reflect the paleoelevation information of a certain location in a certain period in the past, and with the help of these data, the elevation data can be revised more accurately and quantitatively. By combining paleontological data and machine learning technology, it is possible to minimize human bias, improve the objectivity of the annotation and inference process, and thus improve the accuracy and reliability of the paleoelevation inference results.
[0062] Embodiments of the present invention So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A method for reconstructing the paleoelevation of an orogenic belt based on paleontological data and geochemical data, characterized in that: include: Data preparation: Collect known geochemical data, elevation data and paleontological data in the orogenic belt and pre-process the collected data, wherein the geochemical data include trace element composition and isotopes in rock, sediment and fossil samples; the elevation data include modern elevation data directly measured by modern geological measurement technology and ancient elevation data previously recorded or inferred by geological methods; the paleontological data include species information, age information and location information of paleontological organisms; Feature extraction: Analyze the correlation between different geochemical data and elevation changes, and select geochemical feature data from geochemical data according to the degree of correlation between different geochemical data and elevation changes; Elevation annotation: Revise the elevation data of the orogenic belt area based on paleontological data; Model construction: Input the screened geochemical characteristic data and revised elevation data into the constructed machine learning prediction model to learn the nonlinear mapping relationship between geochemical data and elevation; Model evaluation and optimization: Evaluate the accuracy and performance of the elevation prediction model and further optimize the model based on the evaluation results; Application prediction: Input the geochemical sample data of the orogenic belt area to be predicted into the trained elevation prediction model, output the paleo-elevation value of the orogenic belt area to be predicted and display it visually.
2. The method for reconstructing the paleoelevation of an orogenic belt region based on paleontological data and geochemical data according to claim 1, characterized in that: The preprocessing of the collected data includes: Select interpolation methods based on data type to fill in missing geochemical characteristic data; Use statistical methods to identify and remove outliers in geochemical signature data and elevation data; The Min-Max method was used to scale the cleaned data to a preset range, and the Z-Score method was used to convert the data into a distribution with a mean of 0 and a standard deviation of 1.
3. The method for reconstructing the paleo-elevation of an orogenic belt region based on paleontological data and geochemical data according to claim 1, characterized in that: The analysis of the correlation between different geochemical data and elevation changes, and screening of geochemical characteristic data from the geochemical data according to the degree of correlation between different geochemical data and elevation changes, includes: The theoretical correlation between different geochemical data and elevation changes is analyzed through literature, and the quantitative indicators of the correlation between different geochemical data and elevation changes are obtained through statistical analysis. The geochemical data that has both theoretical correlation with elevation and quantitative indicators reaching the preset threshold are selected as geochemical characteristic data; Ratio features are constructed by calculating the ratios between chemical elements. For those geochemical data that have a significant nonlinear relationship with elevation, polynomial transformation is applied to expand them into quadratic or cubic terms, and principal component analysis is used to extract the most representative principal components of the geochemical data.
4. The method for reconstructing the paleo-elevation of an orogenic belt based on paleontological data and geochemical data according to claim 1, characterized in that: The above-mentioned revision of elevation data of the orogenic belt region based on paleontological data includes: Select paleontological data that can represent paleo-elevation information based on the age, distribution, and altitude range of paleontological fossils; According to the age information and location information in the selected paleontological data combined with the location of paleontological fossils, the specific paleo-elevations at different locations in the orogenic belt area in a specific period are estimated; The specific paleoelevations of the orogenic belt area are used to revise the collected paleoelevation data of the orogenic belt area.
5. The method for reconstructing the paleo-elevation of an orogenic belt region based on paleontological data and geochemical data according to claim 1, characterized in that: The method of evaluating the accuracy and performance of the elevation prediction model and further optimizing the model based on the evaluation results includes: The R² value, root mean square error RMSE and mean absolute error MAE are used to comprehensively judge the accuracy and performance of the elevation prediction model; If the elevation prediction model does not achieve the set accuracy and performance, a variety of optimization strategies are implemented: introducing penalty terms through regularization methods to control model complexity, generating samples with small perturbations through data enhancement technology, expanding the training set, and introducing the prediction results of multiple models through ensemble learning to improve the overall prediction performance.
Citation Information
Patent Citations
Model for predicting lymph node metastasis of breast cancer patient without including clinical pathological characteristics
CN117152054A
Mining original digital elevation model reconstruction method and system and medium
CN117635856A
Palaeowater depth speculation method based on SA-RBF neural network model
CN118861643A
Geological survey whole-process data aggregation system for ancient marine environment evolution
CN119201925A
FY4A AGRI surface temperature-based mountain air temperature direct reduction rate space-time distribution estimation method
CN119226955A
Cited By
Ocean-land transition zone paleotopography inference method based on geochemical data and machine learning
CN121457643A