A method for reconstructing the paleo-elevation of orogenic belt regions based on paleontological data and geochemical data

By combining paleobiological data and geochemical data, machine learning models are used to reconstruct paleoelevation in orogenic zones, solving the problems of insufficient sample size, uncertainty in inference and inefficiency, and achieving efficient and accurate paleoelevation reconstruction.

CN119942019BActive Publication Date: 2025-06-20OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510436355.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-06-20
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

The prior art has problems of insufficient sample size, inference uncertainty and inefficiency in paleoelevation reconstruction in orogenic zones.

Method used

The paleoelevation reconstruction method in orogenic zones based on paleobiological data and geochemical data is adopted to improve the accuracy and efficiency of paleoelevation prediction through data preprocessing, feature extraction, elevation annotation and machine learning model construction.

Benefits of technology

It has achieved efficient, accurate and good generalization ability to paleoelevation reconstruction, reduced errors and uncertainties in traditional methods, and is suitable for different geographical regions and geological periods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942019B_ABST
    Figure CN119942019B_ABST
Patent Text Reader

Abstract

The present invention provides a method for reconstructing the paleo-elevation of an orogenic belt region based on paleontological data and geochemical data, belonging to the technical field of orogenic belt elevation reconstruction, including: collecting the known geochemical data, elevation data, and paleontological data of the orogenic belt region and performing preprocessing; screening geochemical characteristic data from the geochemical data by analyzing the correlation between different geochemical data and elevation changes; revising the elevation data of the orogenic belt region according to the paleontological data; inputting the screened geochemical characteristic data and the revised elevation data into the constructed machine learning prediction model to learn the non-linear mapping relationship between geochemical data and elevation; evaluating the accuracy and performance of the elevation prediction model, and further optimizing the model according to the evaluation results; inputting the geochemical sample data of the orogenic belt region to be predicted into the trained elevation prediction model, outputting the paleo-elevation value of the orogenic belt region and performing visual display.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of orogenic belt elevation reconstruction, and particularly relates to a method for reconstructing the paleo-elevation of an orogenic belt region based on paleontological data and geochemical data. Background Art

[0002] Paleo-elevation, as an important indicator of the evolution of the Earth's surface, is of great significance for studying the formation and development process of orogenic belts. For example, by analyzing the chemical composition of magmatic rocks and the Moho depth, the crustal thickness and topographic height of ancient orogenic belts can be inferred.

[0003] Currently, the inference of paleo-elevation mainly relies on the geochemical analysis of geological samples, and its common methods include:

[0004] 1) Whole-rock Mohometers: This method is a method for estimating the crustal thickness based on the relationship between the chemical composition (major elements and trace elements) of magmatic rocks and the crustal thickness. For example, by analyzing the composition of magmatic rocks in the Wanshan structure, a quantitative model that can be applied to ancient orogenic belts is calibrated. There are mainly Mohometers based on the Ce / Y ratio and Mohometers based on the Sr / Y and La / Yb ratios;

[0005] 2) Zircon Mohometers: This method is a method for estimating the crustal thickness based on the trace elements and isotope ratios of zircon. Zircon is a common accessory mineral, and its chemical composition can reflect the origin and evolution process of magma. The zircon Mohometer based on the La / Yb ratio has been used to reconstruct the crustal thickness evolution of the Gangdise Mountains in southern Tibet, and the results are consistent with regional geology. However, the concentration of La in zircon is relatively low, and the measurement error may be large. The zircon Mohometer based on Eu anomaly has been used to reconstruct the crustal thickness evolution of the Gangdise Mountains in southern Tibet. The concentrations of elements such as Eu, Sm, and Gd in zircon are relatively high, and the measurement accuracy is high, making this method more reliable than the La / Yb ratio in some cases.

[0006] 3) Radioactive isotope methods: Radioactive isotope methods are methods for estimating the crustal thickness based on the Sr and Nd isotope ratios of magmatic rocks. This method assumes that the isotope composition of magmatic rocks reflects the mixing process of the crust and the mantle. The Nd isotope method has been used to reconstruct the crustal thickness evolution of southern Tibet, and the results have a high accuracy. Radioactive isotope methods can provide independent information about the crustal thickness, especially in cases where the chemical composition changes are not obvious.

[0007] Although these methods have been widely used in the field of geology, there are the following problems:

[0008] 1) Insufficient sample size: Due to the particularity of the orogenic belt area, some geochemical proxy indicators are missing in the orogenic belt area. For example, basic magmatic rocks are relatively rare in thick continental arcs, and it is impossible to make inferences using the Moho meter based on Ce / Y.

[0009] 2) Uncertainty in inference: The inference of paleo - elevation is related to multiple factors. In the orogenic belt area, due to the large range of depth changes, it is impossible to determine whether certain chemical data is positively or negatively correlated with the data to be inferred.

[0010] 3) Low efficiency: Traditional experimental analysis methods are time - consuming and costly, and it is difficult to meet the needs of large - scale regional reconstruction. Summary of the Invention

[0011] In view of the problems existing in the prior art, the present invention provides a method for reconstructing the paleo - elevation of the orogenic belt area based on paleontological data and geochemical data. On the basis of ensuring a geochemical Moho meter with sufficient accuracy, paleontological data is introduced to revise the paleo - elevation data, improve the prediction accuracy of paleo - elevation, and achieve efficient, accurate and well - generalized paleo - elevation reconstruction.

[0012] The present invention provides a method for reconstructing the paleo - elevation of the orogenic belt area based on paleontological data and geochemical data, including:

[0013] Data preparation: Collect known geochemical data, elevation data and paleontological data in the orogenic belt area and pre - process the collected data. Among them, the geochemical data includes trace element components and isotopes in rocks, sediments and fossil samples, the elevation data includes modern elevation data directly measured by modern geological survey techniques and paleo - elevation data recorded previously or deduced by geological methods, and the paleontological data includes species information, age information and location information of paleontological organisms.

[0014] Feature extraction: Analyze the correlation between different geochemical data and elevation changes, and screen geochemical feature data from geochemical data according to the degree of association between different geochemical data and elevation changes.

[0015] Elevation annotation: Revise the elevation data of the orogenic belt area according to the paleontological data.

[0016] Model construction: Input the screened geochemical feature data and the revised elevation data into the constructed machine - learning prediction model to learn the non - linear mapping relationship between geochemical data and elevation.

[0017] Model evaluation and optimization: Evaluate the accuracy and performance of the elevation prediction model, and further optimize the model according to the evaluation results.

[0018] Application prediction: Input the geochemical sample data of the orogenic belt area to be predicted into the trained elevation prediction model, output the paleo-elevation value of the orogenic belt area to be predicted and display it visually.

[0019] Optionally, the preprocessing of the collected data includes:

[0020] Select interpolation methods based on data type to fill in missing geochemical characteristic data;

[0021] Use statistical methods to identify and remove outliers in geochemical signature data and elevation data;

[0022] The Min-Max method was used to scale the cleaned data to a preset range, and the Z-Score method was used to convert the data into a distribution with a mean of 0 and a standard deviation of 1.

[0023] The analysis of the correlation between different geochemical data and elevation changes, and screening of geochemical characteristic data from the geochemical data according to the degree of correlation between different geochemical data and elevation changes, includes:

[0024] The theoretical correlation between different geochemical data and elevation changes is analyzed through literature, and the quantitative indicators of the correlation between different geochemical data and elevation changes are obtained through statistical analysis. The geochemical data that has both theoretical correlation with elevation and quantitative indicators reaching the preset threshold are selected as geochemical characteristic data;

[0025] Ratio features are constructed by calculating the ratios between chemical elements. For those geochemical data that have a significant nonlinear relationship with elevation, polynomial transformation is applied to expand them into quadratic or cubic terms, and principal component analysis is used to extract the most representative principal components of the geochemical data.

[0026] Optionally, the revising of elevation data of the orogenic belt region according to paleontological data includes:

[0027] Select paleontological data that can represent paleo-elevation information based on the age, distribution, and altitude range of paleontological fossils;

[0028] According to the age information and location information in the selected paleontological data combined with the location of paleontological fossils, the specific paleo-elevations at different locations in the orogenic belt area in a specific period are estimated;

[0029] The specific paleoelevations of the orogenic belt area are used to revise the collected paleoelevation data of the orogenic belt area.

[0030] Optionally, evaluating the accuracy and performance of the elevation prediction model and further optimizing the model according to the evaluation results include:

[0031] The R² value, root mean square error RMSE, and mean absolute error MAE are used to comprehensively judge the accuracy and performance of the elevation prediction model;

[0032] If the elevation prediction model does not meet the set accuracy and performance, multiple optimization strategies are implemented: introducing a penalty term through regularization methods to control the model complexity, generating samples with small perturbations through data augmentation techniques to expand the training set, and improving the overall prediction performance by introducing the prediction results of multiple models through ensemble learning.

[0033] After adopting the above technical solutions, the present invention has at least the following beneficial effects:

[0034] 1. Precision and efficiency

[0035] The present invention utilizes the precise records in paleontological data regarding the elevation survival range of paleontological organisms in specific periods, combined with the inferential ability of geochemical data, to accurately calculate the paleo-elevation of a certain geographical location in the past. Traditional methods often rely on a single data source or rough assumptions, while the present invention provides a more precise and efficient method for inferring paleo-elevation by integrating multiple data sources.

[0036] 2. Data diversity and wide applicability

[0037] The annotation method and model construction method adopted by the present invention can not only process large-scale paleontological data, but also effectively combine various geochemical data, with strong adaptability and generality. Whether in different geographical regions or different geological periods, the method of the present invention can ensure the wide applicability of the inference results, avoiding regional and temporal limitations.

[0038] 3. Avoiding errors in traditional speculation methods

[0039] Traditional paleo-elevation inference often relies on indirect speculation or assumptions, with large errors and uncertainties. The present invention directly utilizes the actual data of the elevation survival range of paleontological organisms in specific periods, greatly reducing the errors in the speculation process and improving the accuracy and credibility of the inference results.

[0040] 4. Efficient data processing and automation

[0041] The present invention combines machine learning algorithms and automated annotation tools to efficiently process and analyze large-scale paleontological data and geochemical data. In the processes of data preprocessing, feature selection, model training, and optimization, automated tools can greatly improve data processing efficiency, save a large amount of labor costs, and ensure the consistency and accuracy of the processing process.

[0042] 5. Providing accurate basis for paleo-environment research

[0043] By reconstructing the paleo - elevation of the orogenic belt region in historical periods, the present invention can provide a more accurate basis for paleo - environment reconstruction, climate change research, and geological history analysis. Researchers can reconstruct the past geographical and climate environments based on accurate paleo - elevation data, and then infer important information such as paleo - ecological changes and species distribution, providing strong data support for research in related fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0045] Figure 1 FIG. is a schematic flowchart of a method for reconstructing the paleo - elevation of an orogenic belt region based on paleontological data and geochemical data provided by an embodiment of the present disclosure;

[0046] Figure 2 FIG. is a scatter plot comparing the predicted value and the actual value of the model paleo - elevation. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0048] Currently, the inference of paleo - elevation mainly relies on the geochemical analysis of geological samples. Its common methods include:

[0049] Stable isotope analysis: For example, δ 18 O and δD are used to reconstruct ancient air temperatures or precipitation source areas. In orogenic belt regions, systematic changes in the oxygen isotope composition can reveal the spatial change process of orogenic belt and plateau uplift;

[0050] Cluster isotope analysis: Through the temperature - indicating effect of carbonate rock clusters in soil (such as ), the temperature when calcite minerals are formed can be calculated, and then combined with the δ 18 O value of atmospheric precipitation and climate models to infer paleo - elevation;

[0051] Analysis of sediment and rock chemical composition: For example, the variation of element ratios (Sr / Ca, Mg / Ca) at different elevations. By sampling and experimental analysis of soil carbonates, sediments, carbonate rocks, and biological fossils, the Sr / Ca and Mg / Ca ratios are combined with other climate indicators to conduct paleo-elevation restoration.

[0052] In addition, the crustal thickness of different ages can be estimated through the geochemistry and isotope data of igneous rocks, that is, the so-called "chemical Mohometer". This method identifies the correlation between certain chemical parameters and the crustal thickness and uses these parameters as proxy indicators of the crustal thickness. This method is particularly useful in studying ancient orogenic belts because direct seismic or geophysical data are unavailable in these areas. To date, multiple chemical proxy indicators have been identified and calibrated for use as indicators of the crustal thickness, such as the Ce / Y ratio, Sr / Y, and La / Yb ratios, etc. These proxy indicators are from the chemical composition of volcanic arc magmas, and the generation depth of the magmas is closely related to the depth of the Mohorovicic discontinuity, and the depth is related to the interaction process between the crust and the mantle.

[0053] In addition, sedimentary zircons can also be used to estimate the crustal thickness. Zircons are particularly important for reconstructing the evolution of the continental crustal thickness because they are well-preserved in sedimentary rocks and can carry information about the original magma source. Studies have shown that the zircon-based Mohometer can provide a new perspective on the crustal evolution of ancient orogenic belts. However, this method also has certain limitations, such as the difficulty in distinguishing whether the source of zircons comes from different tectonic environments, and there may be certain errors in the measurement of trace elements in zircons.

[0054] Based on the above existing research status, summarize the defects of different methods:

[0055] Whole-rock chemical Mohometer:

[0056] 1. Limitations in sample size: The whole-rock chemical Mohometer method relies on large-scale geochemical datasets, but in some areas, especially ancient or eroded orogenic belts, it may be very difficult to obtain sufficient samples, resulting in scarce or insufficient data. These limitations may affect the accuracy of crustal thickness estimation.

[0057] 2. Differences in regional adaptability: The relationship between magma composition and crustal thickness may vary in different geological environments (such as subduction zones, collision zones, etc.). Therefore, some methods may work well in certain areas while being inapplicable in other areas. For example, whole-rock geochemical parameters based on igneous rocks (such as the Sr / Y ratio) have limited adaptability to different tectonic environments and may lead to inaccurate crustal thickness estimations in certain environments.

[0058] Zircon-based Mohometer:

[0059] 1. Limitations in the source of zircon samples: Although zircon is well-preserved in sedimentary rocks, its source is biased towards tectonic environments such as orogenic belts and collision zones. Analyzing zircon from severely eroded orogenic belts or low-elevation areas may lead to errors. Therefore, the Mohorovičić discontinuity meter based on zircon may overestimate the paleo-crustal thickness in certain regions.

[0060] 2. Challenges in analysis techniques: Measuring trace elements and isotopes in zircon usually requires high-precision instruments such as laser ablation inductively coupled plasma mass spectrometry (LA-ICP-MS). These techniques are not only expensive but also may have systematic errors between different instruments and laboratories. The quality of the data directly affects the accuracy of estimating the crustal thickness.

[0061] 3. The complex relationship between zircon chemistry and magma source: The chemical composition of zircon may be affected by various factors such as different magma sources and local crustal contamination, which makes it complex to directly correlate the chemical composition of zircon with the crustal thickness.

[0062] The inventors analyzed the deficiencies of current methods and proposed directions for improvement through further research:

[0063] 1. A paleo-elevation prediction method based on paleobiological information-labeled data is proposed, providing a new idea for the current paleo-elevation prediction method. This method uses data-driven modeling and utilizes the actual information of the strata where paleobiota lived at different times to assist in constructing a prediction model, which can automatically capture the non-linear relationship between complex geochemical data and elevation.

[0064] 2. Integrating the living information of paleobiota with the existing chemical Mohorovičić discontinuity meter improves the credibility of the prediction model, further enhances and expands the accuracy and applicability of the prediction model, and provides a data-driven and efficient paleo-elevation prediction method. This method has the advantages of high efficiency and automation, can process geological data in multiple regions and multiple epochs, and is widely applicable to different types of geological research.

[0065] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0066] As Figure 1 shown, an embodiment of the present disclosure provides a method for reconstructing the paleo-elevation of an orogenic belt region based on paleobiological data and geochemical data, including:

[0067] 1. Data preparation

[0068] To ensure the accuracy and reliability of the model, the data preparation stage is crucial. A high-quality input data set is the basis for subsequent feature extraction and model training. The following are the detailed data sources and preprocessing steps:

[0069] 1.1 Data sources

[0070] Geochemical data is a core component and mainly comes from the following categories:

[0071] Stable isotope data (δ 18 O, δD, etc.): These two indicators reflect the source of water bodies, precipitation, and climate conditions, and are widely used in geological and climatological research. δ 18 O data reflects precipitation characteristics under different geographical regions and climate conditions, while δD can provide temperature information of the water vapor source area.

[0072] Cluster isotopes ( , etc.): This isotope data can be used to estimate the formation temperature of rocks or sediments, and indirectly reflect the elevation changes during the geological period. Through data, the climate and environmental conditions of samples in different ages can be inferred, and then the paleo-elevation at the time of their formation can be estimated.

[0073] Chemical compositions of sediments and rocks (such as Sr / Ca and Mg / Ca ratios): These ratios can provide information on the chemical composition of rocks or sediments, further enrich the data dimension, and then increase the accuracy of elevation inference.

[0074] Chemical compositions of igneous rocks (such as Ce / Y ratio and Sr / Y ratio): Identify the correlation with the crustal thickness, especially in areas without direct seismic data, and use these chemical parameters as proxy indicators for the crustal thickness.

[0075] Chemical composition of sedimentary zircon: Estimate the crustal thickness through trace element and isotope data information in zircon, which is especially suitable for reconstructing the evolution of the continental crust thickness.

[0076] Elevation data is used for comparison and correlation with paleontological data and mainly comes from the following categories:

[0077] Modern elevation data: In some known areas, elevation data obtained through modern measurement techniques provides an accurate reference for the model. Modern elevation data serves as supervised information in the model, which helps to train a more accurate regression model.

[0078] Paleo-elevation data: Since it is difficult to obtain paleo-elevation through direct measurement, it often relies on literature, geological surveys, and inference methods for calibration. Reported paleo-elevation data in the literature and paleo-elevation records obtained through geological methods can provide support for the training dataset.

[0079] Select and download the paleobiological data of the corresponding area from The Paleobiology Database. The data includes key information such as the types, ages, and locations (including longitude, latitude, and the corresponding elevation range information) of paleobiological organisms, and perform preliminary processing to prepare for the subsequent manual annotation step.

[0080] 1.2 Data preprocessing

[0081] Step 1: Missing value handling

[0082] Select an interpolation method according to the data type to fill in the missing geochemical characteristic data, such as mean interpolation, Lagrange interpolation, etc., to ensure data integrity;

[0083] Step 2: Data cleaning

[0084] Use statistical methods to identify and remove outliers in geochemical characteristic data and elevation data to ensure the quality of the dataset. For example, in elevation data, there may be abnormal elevation changes caused by earthquakes or human activities that need to be corrected;

[0085] Step 3: Data standardization and normalization processing

[0086] To meet the input requirements of machine learning models, use the Min - Max method to scale the cleaned data to a preset range to eliminate the influence of different feature dimensions, which helps prevent some features from having an excessive impact on the results due to large numerical ranges during model training; use the Z - Score method to convert the data into a distribution with a mean of 0 and a standard deviation of 1, which helps process data with different distribution characteristics and improves the convergence speed and stability of the model.

[0087] 2. Feature extraction

[0088] 2.1 Feature selection

[0089] Feature selection evaluates the correlation between geochemical data and elevation changes through literature analysis and correlation analysis to ensure that the selected features have both theoretical basis and the correlation degree with paleo - elevation can be quantified statistically. Screen the geochemical data that has both theoretical association with elevation and the quantification index reaches the preset threshold as geochemical characteristic data. For example, δ 18 O and , as indicator substances known to be closely related to temperature and climate conditions, can effectively assist in the speculation of paleo - elevation. Further remove redundant geochemical data and geochemical data with low correlation with elevation, such as recursive feature elimination (RFE), Lasso regression, etc., to reduce the data dimension and significantly improve the efficiency of model training and the accuracy of prediction;

[0090] 2.2 Feature transformation

[0091] Ratio features are constructed by calculating the ratios between chemical elements (such as Sr / Ca and Mg / Ca) to extract more geochemical information. These ratio features can capture the relative relationships between different chemical elements and help better understand the influence of mineral composition on elevation changes. For geochemical data with a significant non-linear relationship with elevation, polynomial transformation is applied to expand it into quadratic or cubic terms to facilitate the capture of complex patterns, thereby enhancing the model's ability to model non-linear relationships. Principal component analysis is used to extract the most representative principal components of geochemical data to reduce redundancy and multicollinearity problems between features, effectively mitigate the curse of dimensionality, and improve the efficiency and stability of model training.

[0092] 3. Elevation annotation

[0093] Based on geological data and paleontological analysis, combined with the regression analysis results of machine learning models, the elevation annotation is accurately corrected. By combining the positioning features of paleontological data (such as species distribution, fossil remains horizons, etc.) with geochemical data (such as sediment composition, isotope ratios, etc.), the predicted paleo-elevation values are fine-tuned using manual annotation methods to ensure that the inferred paleo-elevation values are more consistent with the actual geological background.

[0094] 3.1 Annotation strategy

[0095] Taking paleontological data as the core, paleontological data that can represent paleo-elevation information are selected according to the characteristics of the age, distribution, and survival elevation range of paleontological fossils. Based on the age information and location information in the selected paleontological data and the fossil discovery locations of paleontology, the specific paleo-elevations at different locations in the orogenic belt region during a specific period are calculated. The paleo-elevation data of the orogenic belt region collected are revised using the specific paleo-elevations in the orogenic belt region. For example, a certain paleontological species only survives within a specific elevation range. Combining its fossil discovery location, the paleo-elevation of this region during this period can be determined. This method can accurately reflect the specific elevation information of a certain location in the past period. With the help of these data, the paleo-elevation data can be revised more accurately and quantitatively. This annotation strategy has broad representativeness, can cover the diversity of different geographical regions and geological periods, avoid annotation deviation. In addition, this annotation strategy is operable, ensuring that the annotation process is clear and easy to understand, facilitating practical application, and improving the consistency and accuracy of annotation.

[0096] 3.2 Annotation execution

[0097] The annotation execution stage is led by geological experts. Geological experts infer the paleo-elevation of the sample location based on the elevation survival range and geological history of paleontology, combined with scientific speculation and models, and combine the paleo-elevation data of the orogenic belt region collected by automated annotation tools.

[0098] 3.3 Annotation Quality Control

[0099] To ensure the accuracy and consistency of the annotated data, strict quality control methods are adopted, mainly including the following two means:

[0100] 1. Multiple annotation: Samples in the same dataset will be independently annotated by multiple experts, and finally, the consistency of the annotation is improved by comparing different annotation results. Through multiple annotation, the bias of a single expert's judgment can be effectively reduced, and the reliability of the data labels can be improved.

[0101] 2. Cross-validation: By comparing the annotation results with known real data, the scientificity and accuracy of the annotation are cross-validated. By comparing with existing elevation data and geological research results, it is ensured that the label of each sample has a scientific basis, thereby improving the prediction performance of the model.

[0102] 4. Model Construction

[0103] In the model construction stage, by selecting machine learning algorithms, the model is trained to learn the non-linear mapping relationship between geochemical data and elevation.

[0104] 4.1 Algorithm Selection

[0105] Advanced machine learning algorithms, such as LightGBM and random forest algorithms, are used to construct a regression model between geochemical data and paleo-elevation based on paleontological data. For large-scale datasets and complex regression problems, the LightGBM algorithm has excellent performance advantages. It has excellent performance and strong interpretability, and can provide faster training speed and higher prediction accuracy than traditional algorithms; for medium and small datasets, the random forest algorithm is selected because it can efficiently process multi-dimensional data and has the function of feature importance analysis, which helps to identify and select the most effective feature combination.

[0106] 4.2 Model Training

[0107] In the model training process, by reasonably selecting algorithms, scientifically dividing the dataset, accurately evaluating errors, and optimizing hyperparameters, the efficiency and high accuracy of the established model are ensured. To prevent overfitting, the dataset is divided into a training set and a test set according to a ratio of 70:30, ensuring that the model can maintain good generalization ability on unseen data; to effectively evaluate the prediction accuracy of the model, the mean squared error (MSE) and root mean squared error (RMSE) are selected as loss functions, and these metrics can accurately reflect the error between the prediction result and the actual value; in addition, to further improve the performance of the model, hyperparameter optimization methods are adopted, including grid search, random search, and Bayesian optimization. These methods systematically adjust the model parameters to find the optimal parameter combination, thereby significantly improving the prediction ability and stability of the model.

[0108] 5. Model Evaluation and Optimization

[0109] After the model is constructed, comprehensive evaluation and optimization are crucial steps to ensure that the model has good predictive ability and stability. The model evaluation and optimization stage mainly includes the selection and application of evaluation metrics, as well as the implementation of various optimization strategies.

[0110] 5.1 Selection and Application of Evaluation Metrics

[0111] To comprehensively measure the performance of the model, multiple evaluation metrics are adopted to evaluate the fitting ability and prediction accuracy of the model from different perspectives. The evaluation metrics include the R² value, RMSE (Root Mean Square Error), and MAE (Mean Absolute Error). The R² value is used to measure the fitting ability of the model, and its value ranges from 0 to 1. The closer the R² value is to 1, the stronger the explanatory ability of the model for the target variable and the more accurate the prediction results. RMSE and MAE are also commonly used metrics to evaluate the prediction accuracy and stability of the model. RMSE pays more attention to the penalty of larger errors, while MAE assigns the same weight to all errors and provides the average level of errors by calculating the absolute difference between the predicted value and the actual value. By combining these metrics, the performance of the model at different error scales can be comprehensively judged to ensure that the model has both high accuracy and stability in the prediction task.

[0112] 5.2 Model Optimization

[0113] After initially evaluating the model performance, to further enhance the model's prediction ability and generalization performance, multiple optimization strategies are implemented: By introducing penalty terms through regularization methods to control the model complexity, thereby preventing overfitting of the model. Common regularization methods include L1 regularization and L2 regularization. L1 regularization, by introducing penalties on the absolute values of the weights, helps with feature selection and promotes the generation of sparse models, that is, reducing the influence of irrelevant features. While L2 regularization, by introducing penalties on the squares of the weights, reduces excessive fluctuations in the model parameters and avoids overfitting to a few abnormal data points. Appropriate regularization can help reduce the interference of noise in the training data on the model's prediction ability and improve the model's generalization ability; Through data augmentation techniques, samples with slight perturbations are generated to expand the training set, thereby simulating more diverse geological conditions and input data distributions. Data augmentation can not only increase the quantity of training data but also enhance the model's adaptability to different types of data. Especially when the number of training data samples is limited, in this way, the model can learn a wider range of features and thus improve its prediction performance on unknown data; Through ensemble learning, the prediction results of multiple models are introduced to improve the overall prediction performance, including voting methods, weighted average methods, and stacking methods, etc. For example, combining multiple algorithms such as random forest and LightGBM can give full play to the advantages of different models, reduce the bias and overfitting problems of a single model, and thus obtain more accurate and robust prediction results.

[0114] 6. Application Prediction

[0115] After completing the model construction and optimization, the geochemical sample data of the orogenic belt area to be predicted is input into the trained elevation prediction model, and the paleo-elevation value of the orogenic belt area to be predicted is output and visually displayed. The specific operation steps are as follows:

[0116] Step 1: Input Geochemical Data

[0117] The geochemical sample data of the orogenic belt area to be predicted, including δ 18 O, δD and other key indicators, are input into the trained elevation prediction model, and the calculation process of the model is started to begin inferring the paleo-elevation value corresponding to the sample.

[0118] Step 2: Output Prediction Results

[0119] After receiving the input geochemical sample data, the elevation prediction model infers the paleo-elevation value corresponding to the sample. The prediction result is output in numerical form, representing the paleo-elevation situation of the target area during a specific geological period. The output paleo-elevation value provides important reference data for geological researchers to help them further analyze and reconstruct the trend of paleo-elevation changes.

[0120] Step 3: Result Visualization

[0121] To facilitate in-depth analysis and interpretation of the prediction results by geological researchers, the predicted paleo-elevation values will be presented through various visualization tools, helping researchers better understand the paleo-elevation characteristics of different regions and conduct comprehensive analysis with other geological data, thus providing strong support for geochemical research, paleo-environment reconstruction, etc.

[0122] Based on the above embodiments, the following specific examples are proposed. It can be understood that the following specific examples only exemplarily elaborate on the specific implementation of the above embodiments and do not limit the technical solutions of the above embodiments.

[0123] 1. Data Sources and Sample Composition

[0124] 1195 groups of sample data are integrated, covering the following multi-source datasets:

[0125] Geochemical data: including stable isotopes (δ 18 O, δD), cluster isotopes ( ), ratios of igneous rock elements (Sr / Y, La / Yb), etc., sourced from the geological survey database of global orogenic belts;

[0126] Paleontological data: Fossil distribution information is extracted from ThePaleobiologyDatabase, including the survival age, geographical coordinates, and corresponding elevation ranges of species, serving as the basis for correcting elevation labels;

[0127] Modern elevation data: The SRTM digital elevation model (DEM) is used as the reference verification data;

[0128] 2. Data Preprocessing

[0129] 2.1 Data Cleaning and Missing Value Handling

[0130] Missing value filling: Interpolation methods are used to fill in missing geochemical characteristics. For the handling of missing values, interpolation methods (such as mean interpolation, Lagrange interpolation, etc.) are selected according to the data type to ensure data integrity;

[0131] Outlier removal: Based on the 3σ principle, outliers (such as extreme isotope ratios) in geochemical parameters are identified and removed;

[0132] 2.2 Data Standardization and Encoding

[0133] Numeric features: Min-Max normalization is used to scale geochemical parameters to the [0, 1] interval, eliminating the dimension difference;

[0134] 3. Feature Engineering

[0135] Feature concatenation: Numeric features (such as δ 18O, ancient longitude and latitude, , Sr / Ca and Mg / Ca), and label encoding features are stacked into a combined feature matrix (total dimension: 215) as the model input;

[0136] 4. Model construction and training

[0137] 4.1 Algorithm selection and parameter setting

[0138] LightGBM regressor: suitable for large-scale data, set the learning rate to 0.05, tree depth to 8, number of leaf nodes to 64, and early stopping rounds to 50;

[0139] Random forest regressor: used for comparative experiments, set the number of trees to 200, maximum depth to 10, and minimum number of samples per leaf to 5;

[0140] 4.2 Dataset division and training

[0141] Divide the training set (837 samples) and test set (358 samples) in a 7:3 ratio, and use 5-fold cross-validation to optimize hyperparameters;

[0142] 5. Model evaluation and results

[0143] 5.1 Performance metrics

[0144] R² (train) of the training set: 0.966.

[0145] R² (test) of the test set: 0.816.

[0146] Root mean square error (RMSE): 21.3m.

[0147] R² (coefficient of determination) measures the explanatory power of the model. The closer the value is to 1, the stronger the prediction ability of the model. The root mean square error (RMSE) measures the degree of difference between the predicted value and the true value of the model.

[0148] 5.2 Result visualization

[0149] The results are as Figure 2 shown. The horizontal axis is the model predicted value, the vertical axis is the actual elevation value. The solid line is the 1:1 ideal fitting line. The triangular points are the training set samples, and the circular points are the test set samples. The triangular points and circular points are densely distributed along the 1:1 ideal fitting line, indicating that the model has good generalization ability on both the training set and the test set.

[0150] Compared with traditional methods, the present invention has significant advantages in multiple aspects: First, by using a machine learning model based on paleontological and geochemical data, it can efficiently complete the paleo - elevation reconstruction of large - scale orogenic belt regions, solving the problems of scarce samples and low experimental efficiency of traditional methods; Second, machine learning can automatically perform non - linear fitting, overcoming the problem that traditional empirical formulas are difficult to comprehensively describe the coupling relationship of complex factors; In addition, the model has strong generalization ability, can adapt to the application requirements of different regions and geological epochs, and significantly improves the prediction accuracy and reliability; Finally, in traditional methods, many annotation and inference processes often rely on artificial experience and there is a certain subjectivity. However, paleontological data can accurately reflect the paleo - elevation information of a certain location at a certain past time. By using these data, the elevation data can be revised more accurately and quantitatively. By combining paleontological data and machine learning technology, the human bias can be minimized, and the objectivity of the annotation and inference processes can be enhanced, thereby improving the accuracy and reliability of the paleo - elevation inference results.

[0151] So far in the embodiments of the present invention, the technical solutions of the present invention have been described in combination with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

Claims

1. A method for reconstructing the paleoelevation of an orogenic belt based on paleontological data and geochemical data, characterized in that: include: Data preparation: Collect known geochemical data, elevation data and paleontological data in the orogenic belt and pre-process the collected data, wherein the geochemical data include trace element composition and isotopes in rock, sediment and fossil samples; the elevation data include modern elevation data directly measured by modern geological measurement technology and ancient elevation data previously recorded or inferred by geological methods; the paleontological data include species information, age information and location information of paleontological organisms; Feature extraction: Analyze the correlation between different geochemical data and elevation changes, and select geochemical feature data from geochemical data according to the degree of correlation between different geochemical data and elevation changes; Elevation annotation: Revise the elevation data of the orogenic belt area based on paleontological data; Model construction: Input the screened geochemical characteristic data and revised elevation data into the constructed machine learning prediction model to learn the nonlinear mapping relationship between geochemical data and elevation; Model evaluation and optimization: Evaluate the accuracy and performance of the elevation prediction model and further optimize the model based on the evaluation results; Application prediction: Input the geochemical sample data of the orogenic belt area to be predicted into the trained elevation prediction model, output the paleo-elevation value of the orogenic belt area to be predicted and display it visually; The analysis of the correlation between different geochemical data and elevation changes, and screening of geochemical characteristic data from the geochemical data according to the degree of correlation between different geochemical data and elevation changes, includes: The theoretical correlation between different geochemical data and elevation changes is analyzed through literature, and the quantitative indicators of the correlation between different geochemical data and elevation changes are obtained through statistical analysis. The geochemical data that has both theoretical correlation with elevation and quantitative indicators reaching the preset threshold are selected as geochemical characteristic data; Ratio features are constructed by calculating the ratios between chemical elements. For those geochemical data that have a significant nonlinear relationship with elevation, polynomial transformation is applied to expand them into quadratic or cubic terms, and principal component analysis is used to extract the most representative principal components of the geochemical data.

2. The method for reconstructing the paleoelevation of an orogenic belt region based on paleontological data and geochemical data according to claim 1, characterized in that: The preprocessing of the collected data includes: Select interpolation methods based on data type to fill in missing geochemical characteristic data; Use statistical methods to identify and remove outliers in geochemical signature data and elevation data; The Min-Max method was used to scale the cleaned data to a preset range, and the Z-Score method was used to convert the data into a distribution with a mean of 0 and a standard deviation of 1.

3. The method for reconstructing the paleo-elevation of an orogenic belt region based on paleontological data and geochemical data according to claim 1, characterized in that: The above-mentioned revision of elevation data of the orogenic belt region based on paleontological data includes: Select paleontological data that can represent paleo-elevation information based on the age, distribution, and altitude range of paleontological fossils; According to the age information and location information in the selected paleontological data combined with the location of paleontological fossils, the specific paleo-elevations at different locations in the orogenic belt area in a specific period are estimated; The specific paleoelevations of the orogenic belt area are used to revise the collected paleoelevation data of the orogenic belt area.

4. The method for reconstructing the paleo-elevation of an orogenic belt based on paleontological data and geochemical data according to claim 1, characterized in that: The method of evaluating the accuracy and performance of the elevation prediction model and further optimizing the model based on the evaluation results includes: The R² value, root mean square error RMSE and mean absolute error MAE are used to comprehensively judge the accuracy and performance of the elevation prediction model; If the elevation prediction model does not achieve the set accuracy and performance, a variety of optimization strategies are implemented: introducing penalty terms through regularization methods to control model complexity, generating samples with small perturbations through data enhancement technology, expanding the training set, and introducing the prediction results of multiple models through ensemble learning to improve the overall prediction performance.

Citation Information

Patent Citations

  • Geological survey whole-process data aggregation system for ancient marine environment evolution

    CN119201925A

  • Computer-implemented method and system for predicting geohazard risk or pipeline strain in relation to a pipeline system using machine learning

    WO2024123170A1