Ocean-land transition zone paleotopography inference method based on geochemical data and machine learning

By combining multi-source data such as stratigraphic lithology, paleolatitude and longitude, paleoage, and geochemical indicators, and using the LightGBM algorithm to construct nonlinear mapping relationships, the uncertainties and inefficiencies of traditional paleotopographic inference methods are solved, and efficient and accurate paleotopographic reconstruction and inference are achieved.

CN121457643APending Publication Date: 2026-02-03OCEAN UNIV OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610007162.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-06
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Existing paleotopographic inference methods are affected by the nonlinear coupling of multiple factors, resulting in high uncertainty and uncertainty in the inference process. Traditional methods are time-consuming and costly, making it difficult to meet the needs of large-scale regional paleotopographic reconstruction.

Method used

By collecting multi-source geological data on stratigraphy, lithology, paleolatitude and longitude, paleoage, and geochemical indicators, and combining the LightGBM algorithm, a nonlinear mapping relationship between geochemical characteristics and elevation is constructed. Feature engineering and model training are then carried out, and the model is optimized by K-fold cross-validation and hyperparameter tuning to achieve automated prediction of paleotopography.

Benefits of technology

It improves the accuracy and reliability of paleotopographic inference, enables precise reconstruction across media, environments, and eras, enhances work efficiency, and possesses good generalization ability and interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121457643A_ABST
    Figure CN121457643A_ABST
Patent Text Reader

Abstract

The invention discloses an ocean-land transition zone paleotopography inference method based on geochemical data and machine learning, and belongs to the technical field of combination of geoscience and machine learning. The method comprises the following steps: acquiring multi-source geological data such as stratum lithology, ancient longitude and latitude, ancient age and geochemical indexes; carrying out numeralization conversion on the category-type features, and screening features related to elevation through correlation analysis; performing label distribution on the samples based on geological expert judgment; training the training set by adopting a LightGBM algorithm, and constructing a paleotopography inference model; a K-fold cross validation method is adopted to evaluate the generalization ability of the model, and the performance of the model is optimized through hyper-parameter adjustment and regularization processing; and inputting feature data of an unknown sample into the optimized model for prediction and visual display. Through the multi-source feature fusion and machine learning technology, the linear limitation of a traditional empirical formula is broken through, and efficient and accurate paleotopography reconstruction with good generalization ability is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of combining geoscience and machine learning, and particularly relates to an ocean-land transition zone paleotopography inference method based on geochemical data and machine learning. BACKGROUND

[0002] With the continuous progress of observation and experimental means, geology has entered an era of highly rich data. Remote sensing images, field measurements, laboratory tests, and various geophysical and geochemical detection means jointly provide multi-level data from microscopic mineral composition to macroscopic crustal structure. These data span from microns to kilometers in spatial scale and from seconds of tectonic activity to million years of landform evolution in time dimension, making the geological science research objects exhibit typical characteristics of nonlinearity, multi-scale, and strong coupling. Continental margins, as important source areas adjacent to continents, are one of the main sedimentary areas on Earth. Many folded uplift belts that have become inland mountains were formed in the early stage of paleocontinental margins. Research on continental margins not only helps to reveal the early orogenic process of the Earth, but also provides an important clue for understanding crustal evolution. Continental margins also contain rich oil and gas resources and various mineral resources, and the study of paleotopography is an important direction in this field.

[0003] Paleotopography, as an important indicator of the evolution of the Earth's surface, is of great significance for studying climate change, tectonic movement, crustal uplift, and weathering and erosion. Current paleotopography inference mainly relies on geochemical analysis of geological samples, and common methods include stable isotope analysis, cluster isotope analysis, and sediment and rock chemical composition analysis. Stable isotope analysis mainly uses oxygen and hydrogen isotopes to reflect changes in ancient air temperature or precipitation source area, and calculates paleoelevation by analyzing isotopic composition in precipitated carbonates or higher plant leaf lipids. Element ratio analysis uses strontium-calcium ratio and magnesium-calcium ratio to reflect the chemical composition of rocks or sediments. Elevation changes often accompany environmental changes in temperature and precipitation, significantly affecting rock weathering intensity and producing differentiation effects on element deposition and migration, thereby changing element chemical ratios.

[0004] However, existing paleotopography inference methods have several technical limitations. The relationship between geochemical characteristics and paleotopography is influenced by the coupling of multiple factors such as climate, diagenesis, tectonic setting, paleolatitude, and paleoage. This complex nonlinear relationship is difficult to fully describe through traditional single empirical formulas, resulting in high uncertainty in the inference process. High-quality geological samples in many key areas are sparse in spatial distribution and have large age spans, and the inference results relying on single methods or limited data may lack reliability and representativeness. Traditional experimental analysis methods are time-consuming and costly, and the manual integration process of multi-source data is tedious, making it difficult to meet the needs of large-scale regional paleotopography reconstruction and inefficient.

[0005] In view of the above technical problems, there is an urgent need to develop a new paleotopographic inference method that can handle multi-factor nonlinear coupling relationships, still has good prediction performance and generalization ability under limited sample conditions, and can realize efficient and automated prediction. Existing research shows that stratigraphic lithology, paleolatitude and paleoage information also play an important role in the reconstruction of paleotopography. Combining these geological constraints with geochemical indicators, and establishing a mapping relationship between multi-source features and elevation through data-driven machine learning methods, it is expected to break through the limitations of traditional methods and provide a new solution for paleotopographic research. SUMMARY

[0006] To solve the problems in the background art, the present application provides an ocean-land transition zone paleotopographic inference method based on geochemical data and machine learning, comprising the following steps:

[0007] S1, data preparation: collect multi-source geological data including stratigraphic lithology, paleolatitude, paleoage and geochemical indicators, and corresponding elevation data, and perform data cleaning, data normalization processing and missing value filling processing on the collected data to construct a multi-source geological data set;

[0008] S2, feature engineering: numerical conversion is performed on the category type features in the multi-source geological data set, the spatiotemporal features are constructed, and the features related to the elevation are selected through correlation analysis to construct a model input feature set;

[0009] S3, manual annotation: based on the judgment of geological experts and the paleodepositional environment information of samples, the samples in the data set are labeled to generate a labeled training data set;

[0010] S4, model construction: the labeled training data set is divided into a training set and a test set, the training set is trained using the LightGBM algorithm, the prediction error is evaluated using a loss function, and a paleotopographic inference model is constructed;

[0011] S5, model evaluation and optimization: the generalization ability of the model is evaluated using the K-fold cross-validation method, the model performance is optimized through hyperparameter adjustment and regularization processing, and the optimized paleotopographic inference model is output;

[0012] S6, application prediction: after the multi-source geological feature data of unknown samples are preprocessed and feature converted in the same way as the training data, the optimized paleotopographic inference model is input, and the paleotopographic prediction result is output and visualized.

[0013] Further, the data preparation in step S1 specifically includes the following substeps:

[0014] S11, data collection: collect stratigraphic lithology data, paleolatitude data, paleoage data, and geochemical index data including δ 18 O values, δD values, Sr / Ca ratios, and Mg / Ca ratios of rock, sediment, or fossil samples, and collect modern elevation data or geological record elevation data corresponding to the samples;

[0015] S12, data cleaning: statistics of the distribution of each feature, calculation of the mean and standard deviation of each feature value, identification of data points deviating from the mean by more than 3 times the standard deviation as outliers and removal, and obtaining of the cleaned data set;

[0016] S13, data normalization: for each numerical feature in the cleaned data set, Min-Max normalization is performed according to the following formula: ;

[0017] In the formula, X is the normalized feature value, with a value range of 0 to 1; X is the original feature value; Xmin is the minimum value of the feature in the data set; Xmax is the maximum value of the feature in the data set;

[0018] S14, missing value filling: for numerical features with missing values in the data set, the mean of all non-missing values of the feature is used for filling, and a complete multi-source geological data set is obtained.

[0019] Further, the feature engineering in step S2 specifically includes the following sub-steps:

[0020] S21, numerical conversion of categorical features: for stratigraphic lithology types, sedimentary facies types and other categorical features, a feature hashing method is used to convert them into numerical features, and the specific operation is: calculate the hash value of each category value, map the hash value to a predetermined numerical interval, and generate the corresponding numerical feature value;

[0021] S22, spatiotemporal feature construction: for paleolatitude data, calculate the spherical distance of each sample point to the paleoequator as a new spatial feature; for paleoage data, segment coding is performed according to eras or epochs to generate time segment features;

[0022] S23, correlation analysis and feature selection: calculate the Pearson correlation coefficient between each feature and the elevation according to the following formula: ;

[0023] In the formula, r is the Pearson correlation coefficient, with a value range of -1 to 1, and the larger the absolute value, the stronger the correlation; Xi is the feature value of the i th sample; is the mean of the feature value of all samples; Yi is the elevation value (m) of the i th sample; The mean (in meters) of the elevation values ​​of all samples is used; features with absolute values ​​of Pearson correlation coefficients greater than a preset threshold are selected to construct the model input feature set.

[0024] Furthermore, the manual annotation in step S3 specifically includes the following sub-steps:

[0025] S31. Standardization of Labelling: Based on the judgment of geological experts, the standard was determined to be δ. 18 O value, δD value, Sr / Ca ratio, and Mg / Ca ratio serve as markers reflecting paleodepositional environments;

[0026] S32. Sample Screening: Select samples from the dataset that have complete geochemical indicators and reliable paleolatitude and paleoage information to ensure that the selected samples are representative in terms of spatial distribution and time span.

[0027] S33. Label assignment: Geological experts assign the corresponding elevation values ​​of each sample as label values ​​based on the paleosedimentary environment information and geological background of each sample.

[0028] S34. Labeling quality verification: Using a multi-expert independent labeling method, at least two experts assign labels to the same batch of samples, compare the consistency of the labeling results of each expert, and review and confirm the samples with inconsistent labels to generate a labeled training dataset.

[0029] Furthermore, the model construction in step S4 specifically includes the following sub-steps:

[0030] S41. Dataset partitioning: The labeled training dataset is randomly divided into a training set and a test set in a 7:3 ratio. The training set is used for model parameter learning, and the test set is used for model performance evaluation.

[0031] S42, LightGBM Model Training: The LightGBM algorithm is used to iteratively train the training set. In each iteration, the algorithm performs the following operations: First, the continuous values ​​of each feature are discretized into several intervals and a histogram is constructed. Then, the leaf-by-leaf growth strategy is adopted to select the leaf node with the largest splitting gain from all leaf nodes of the current decision tree for splitting, generating new child nodes. The iteration continues until the stopping condition is met.

[0032] S43. Loss Function Calculation: After each iteration, the root mean square error is used as the loss function to evaluate the prediction error of the current model on the training set, calculated according to the following formula: ;

[0033] In the formula, L is the root mean square error (meters); n is the number of samples in the training set; and Yi is the actual elevation value (meters) of the i-th sample. Predict the elevation value (in meters) for the i-th sample using the model;

[0034] S44. Model Output: When the loss function value converges or reaches the preset maximum number of iterations, stop training and output the ancient terrain inference model.

[0035] Furthermore, the model evaluation and optimization in step S5 specifically includes the following sub-steps:

[0036] S51, K-fold cross-validation: The labeled training dataset is randomly shuffled and divided into K disjoint subsets, where K is 5. K iterations of validation are performed. In each iteration, one subset is selected as the validation set, and the remaining K-1 subsets are combined as the training set. The model is trained on the training set, and the accuracy is calculated on the validation set. After K iterations, the average of the K validation accuracies is calculated as the model's cross-validation accuracy, using the following formula: ;

[0037] In the formula, Acv is the cross-validation accuracy; K is the number of folds, with a value of 5; Ak is the accuracy on the k-th fold validation set;

[0038] S52. Model Performance Evaluation: Calculate the model's F1 score on the test set using the following formula: ;

[0039] In the formula, F1 is the F1 score, which ranges from 0 to 1; P is the precision, which represents the proportion of samples that the model predicted as positive but were actually positive; and R is the recall, which represents the proportion of samples that were actually positive but were correctly predicted as positive by the model.

[0040] S53. Hyperparameter tuning: Using a grid search method, within the preset range of hyperparameter values, all hyperparameter combinations are traversed, and the hyperparameter combination that maximizes the cross-validation accuracy is selected as the optimal hyperparameter.

[0041] S54. Regularization: During model training, an L2 regularization term is introduced to control model complexity, prevent overfitting, and output an optimized paleoterrain inference model.

[0042] Furthermore, the application prediction in step S6 specifically includes the following sub-steps:

[0043] S61. Preprocessing of data to be predicted: Obtain stratigraphic lithology, paleolatitude and longitude, paleoage and geochemical index data of unknown samples, and preprocess the data according to the same data cleaning and normalization methods as in step S1.

[0044] S62. Feature transformation: Following the same feature transformation method as in step S2, perform categorical feature numerical transformation and spatiotemporal feature construction on the preprocessed data to generate a feature vector to be predicted that is consistent with the format of the model input feature set.

[0045] S63. Model Prediction: Input the feature vector to be predicted into the optimized paleotopographic inference model. The model performs calculations through its internal decision tree ensemble and outputs the paleotopographic prediction values ​​for each unknown sample.

[0046] S64. Result Visualization: The predicted paleotopographic values ​​are plotted as contour maps, heat maps, or three-dimensional terrain models according to the spatial location information of the samples, thus completing the visualization of the paleotopographic inference results.

[0047] Furthermore, step S5 also includes a feature contribution analysis sub-step: using the SHAP method to analyze the contribution of each input feature to the model prediction result. Specifically, for each sample to be analyzed, the SHAP value of each feature is calculated. A positive SHAP value indicates that the feature has a positive contribution to the prediction result, and a negative SHAP value indicates that the feature has a negative contribution to the prediction result. The larger the absolute value of the SHAP value, the higher the contribution of the feature. The SHAP values ​​of all features of all samples are summarized, the average of the absolute values ​​of the SHAP values ​​of each feature is calculated, and the features are sorted from largest to smallest according to the average value to identify the key features that contribute the most to the paleotopographic prediction.

[0048] The beneficial effects achieved by this invention are as follows:

[0049] This invention constructs a multi-source composite feature system by integrating stratigraphic lithology, paleolatitude and longitude, paleoage, and various geochemical indicators. It then employs machine learning algorithms to establish a nonlinear mapping relationship between geochemical features and elevation, overcoming the limitations of traditional empirical formulas that can only handle linear relationships. This allows the model to automatically capture complex nonlinear correlations under multi-factor coupling conditions, significantly improving the accuracy and reliability of paleotopographic inference. Furthermore, this invention combines stable isotopic indicators with elemental ratio indicators, while introducing stratigraphic lithology, paleolatitude and longitude, and paleoage as spatial-temporal-environmental constraints. The model can effectively correct systematic deviations in geochemical indicators caused by climate gradients and geological ages, achieving accurate paleotopographic reconstruction across media, environments, and eras.

[0050] The LightGBM algorithm employed in this invention features a histogram-based decision tree construction method and a leaf-growth strategy, facilitating efficient processing of high-dimensional, multi-source, heterogeneous data. Compared to traditional experimental analysis methods, it significantly shortens data processing and prediction time, achieving full automation of the paleotopographic inference process, reducing errors caused by manual intervention and subjective judgment, and effectively improving the efficiency of large-scale regional paleotopographic reconstruction. This invention uses feature engineering to numerically transform categorical features, scientifically constructs spatiotemporal features, and filters features highly correlated with elevation through correlation analysis. This allows the model to focus on high-value features while filtering redundant information, further improving model training efficiency and prediction accuracy.

[0051] This invention evaluates the model's generalization ability using K-fold cross-validation and optimizes model performance through hyperparameter tuning and regularization, effectively preventing overfitting and ensuring highly consistent predictive performance across different data subsets. Furthermore, this invention employs the SHAP method to analyze the contribution of each input feature to the model's prediction results, identifying the key features that contribute most to paleotopographic predictions. This not only enhances the model's interpretability but also provides valuable scientific insights for geological research, contributing to a deeper understanding of the intrinsic relationship between geochemical indicators and paleotopography.

[0052] This invention's method possesses strong generalization ability and wide applicability, adaptable to the paleotopographic inference needs of different geological environments such as ocean-continental transition zones, plateaus, orogenic belts, and foreland basins. It can be extended to research in different regions and geological eras globally. This invention provides an efficient, accurate, and highly generalizable technical means for paleotopographic reconstruction, paleoclimate analysis, crustal evolution research, and geochemical analysis, promoting the interdisciplinary integration of geochemistry, stratigraphy, and artificial intelligence, and providing reliable technical support for regional geological evolution and tectonic movement research. Attached Figure Description

[0053] Figure 1 This is a scatter plot of the SHAP feature contribution of ancient latitude and longitude accuracy, showing the marginal impact of ancient latitude and longitude data accuracy on model prediction results.

[0054] Figure 2 This is a scatter plot of the SHAP feature contribution analysis of paleolar latitude, showing the impact of paleolar latitude on model predictions and the differences between the Northern and Southern Hemispheres.

[0055] Figure 3 This is a horizontal bar chart of the SHAP contribution values ​​of environmental features in the ocean-continent transition zone, sorted by average SHAP value to show the contribution intensity of each input feature.

[0056] Figure 4 It is a SHAP bee colony diagram of the ocean-continent transition zone environment, showing the heterogeneity and dispersion of the influence of each feature on the prediction.

[0057] Figure 5 This is a scatter plot of the SHAP interaction analysis of lithology and paleolatitude, showing the marginal impact of the synergistic effect of the two on the model predictions.

[0058] Figure 6 This is a technical roadmap for implementing the method of the present invention, showing the complete process from data preparation to application prediction. Detailed Implementation

[0059] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. In addition, the forms of the various structures described in the following embodiments are merely illustrative. The present invention is not limited to the structures described in the following embodiments. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0060] This invention provides a paleotopographic inference method for ocean-continental transition zones based on geochemical data and machine learning. This method addresses the problems of complex multi-factor coupling, sample scarcity, and low efficiency in traditional paleotopographic inference. By integrating multi-source geological data such as stratigraphy, lithology, paleolatitude and longitude, paleoage, and geochemical indicators, and combining this with the LightGBM machine learning algorithm, a nonlinear mapping relationship between geochemical features and elevation is constructed, thereby achieving efficient, accurate, and well-generalized paleotopographic reconstruction. (Refer to...) Figure 6 The technical roadmap shown indicates that this method includes six main steps: data preparation, feature engineering, manual annotation, model building, model evaluation and optimization, and application prediction.

[0061] Step S1 is the data preparation stage. The purpose of this stage is to collect multi-source geological data, including stratigraphic lithology, paleolatitude and longitude, paleoage, and geochemical indicators, as well as corresponding elevation data. The collected data undergoes data cleaning, normalization, and missing value imputation to ultimately construct a multi-source geological dataset. A high-quality input dataset is fundamental to ensuring the accuracy and reliability of the model; therefore, the data preparation stage is crucial.

[0062] S11. Data Acquisition: In step S11, multi-source geological data from rock, sediment, or fossil samples needs to be systematically collected. Stratigraphic lithological data includes rock type, lithological assemblage characteristics, sedimentary facies types, etc., and can be obtained through field geological surveys, well core observations, and thin-section identification. Stratigraphic lithological data is the core foundation of geological research. Different lithologies represent different sedimentary tectonic settings. For example, conglomerate often indicates high-energy, near-source paleohighlands, while mudstone or deep-water sediments indicate low-lying basin environments. Therefore, stratigraphic lithological data has important indicative significance for paleotopographic inference.

[0063] Paleolatitude and longitude data are obtained through paleomagnetic testing and plate tectonics reconstruction models, reflecting the spatial location of geological samples during the depositional period. Paleoage data are obtained through radiometric dating methods, such as U-Pb dating, Ar-Ar dating, and carbon-14 dating, to determine the depositional or formation age of the samples, providing a temporal constraint on topographic evolution. Introducing paleolatitude, longitude, and paleage data can effectively correct for systematic biases in geochemical indicators caused by climate gradients and geological ages, making paleotopographic inferences more accurate.

[0064] Geochemical index data include δ 18 O value, δD value, Sr / Ca ratio, and Mg / Ca ratio. δ 18 O and δD are stable isotope indicators that reflect the source of water bodies, precipitation, and climatic conditions. 18 O data reflects precipitation characteristics under different geographical regions and climatic conditions, while δD provides temperature information about water vapor sources. Sr / Ca and Mg / Ca ratios are important indicators reflecting the chemical composition of rocks or sediments. Increased elevation is often accompanied by environmental changes such as decreased temperature and increased precipitation, significantly enhancing rock weathering intensity and differentiating elemental deposition and migration, thereby altering elemental chemical ratios. Simultaneously, it is necessary to collect modern elevation data or geological record elevation data corresponding to the samples as supervisory information for model training.

[0065] S12. Data Cleaning: In step S12, the collected data needs to be cleaned. First, the distribution of each feature is statistically analyzed, and the mean and standard deviation of each feature value are calculated. Then, data points that deviate from the mean by more than three times the standard deviation are identified as outliers and removed. This outlier detection method based on the 3σ principle is based on the normal distribution assumption and can effectively identify outliers and noise points in the dataset. After removing data records that do not meet the standards, a cleaned dataset is obtained, ensuring that subsequent analysis is based on a high-quality data foundation.

[0066] S13. Data Normalization: In step S13, Min-Max normalization is performed on each numerical feature in the cleaned dataset. Since different features have significantly different dimensions and numerical ranges, without normalization, some features with large numerical ranges may have an excessive influence during model training. The Min-Max normalization method linearly scales each feature value to a uniform range of 0 to 1, eliminating the influence of different feature dimensions and helping to improve the convergence speed and stability of the model. The normalization process is performed according to the following formula: In the formula, These are the normalized eigenvalues, ranging from 0 to 1; These are the original eigenvalues; This is the minimum value of the feature in the dataset; This represents the maximum value of the feature in the dataset.

[0067] S14. Missing Value Imputation: In step S14, for numerical features in the dataset with missing values, the mean of all non-missing values ​​of that feature is used for imputation. Mean imputation is a commonly used missing value handling method that can ensure data integrity while maintaining the overall distribution characteristics of the dataset. After missing value imputation, a complete multi-source geological dataset is obtained, providing a data foundation for subsequent feature engineering and model training.

[0068] Step S2 is the feature engineering stage. The purpose of this stage is to numerically transform categorical features from the multi-source geological dataset, construct spatiotemporal features, and filter features related to elevation through correlation analysis, ultimately building the model's input feature set. Feature engineering is a crucial step in machine learning modeling; high-quality features can significantly improve the model's predictive performance.

[0069] S21. Categorical Feature Numerical Conversion: In step S21, categorical features such as stratigraphic lithology and sedimentary facies types are converted into numerical features using the feature hashing method. Feature hashing is an efficient categorical feature encoding method. Specifically, it calculates the hash value for each categorical value, maps the hash value to a preset numerical range, and generates the corresponding numerical feature value. Compared to traditional one-hot encoding, feature hashing effectively controls feature dimensionality, avoids the dimensionality explosion problem caused by high cardinality categorical features, and maintains the distinguishability between categories, making it easier for the model to capture the differences between different categories.

[0070] S22, Spatiotemporal Feature Construction; In step S22, spatial feature expansion is performed on the ancient latitude and longitude data, and the spherical distance from each sample point to the ancient equator is calculated as a new spatial feature. Since δ 18 Geochemical indicators such as O, δD, Sr / Ca, and Mg / Ca are all controlled by climate zonation, precipitation pathways, and latitudinal effects. Introducing the spatial feature of the sample's distance from the paleoequator can reflect the climate zone location of the sample, providing important spatial constraints for the model. Paleo-age data are segmented and coded according to geological periods or epochs to generate temporal segmentation features. Different geological periods exhibit significantly different climatic backgrounds, tectonic patterns, and weathering intensities. Temporal segmentation coding enables the model to identify characteristic patterns from different geological periods, avoiding systematic errors caused by cross-era mixing.

[0071] S23. Correlation Analysis and Feature Screening: In step S23, the Pearson correlation coefficient between each feature and elevation is calculated to quantitatively assess the strength of the linear association between each feature and the target variable. The Pearson correlation coefficient is a statistic that measures the degree of linear correlation between two variables, and its calculation formula is as follows: ;

[0072] In the formula, The Pearson correlation coefficient ranges from -1 to 1, with a larger absolute value indicating a stronger correlation. For the first Feature values ​​of each sample; This is the mean of the feature values ​​for all samples. For the first Elevation values ​​for each sample, in meters; This is the mean of all sample elevation values, in meters.

[0073] Based on the calculated Pearson correlation coefficient, features with absolute values ​​greater than a preset threshold are selected. Preferably, the threshold can be set between 0.1 and 0.3, and the specific value can be adjusted according to the characteristics of the dataset and the number of features. Filtering features highly correlated with elevation through correlation analysis helps to eliminate redundant and low-correlation features, reduce data dimensionality, improve model training efficiency and prediction accuracy, and ultimately construct the model input feature set.

[0074] Furthermore, step S2 may also include a polynomial feature construction sub-step. For features found to have a nonlinear relationship with elevation through correlation analysis, the feature value is transformed by a quadratic or cubic power to generate a polynomial feature. Let the original feature value be... The eigenvalue of the generated quadratic term is... The eigenvalue of the cubic term is Polynomial features can enhance the model's ability to model nonlinear relationships. By adding the generated polynomial features to the model's input feature set, the model can better fit the complex nonlinear relationship between features and elevation.

[0075] Step S3 is the manual annotation stage. The purpose of this stage is to assign labels to the samples in the dataset based on the judgment of geological experts and the paleosedimentary environment information of the samples, generating a labeled training dataset. In machine learning, the label value refers to the target variable that the model wants to predict, and the accuracy of the label directly affects the model's predictive performance.

[0076] S31, Standardization of Labelling; In step S31, based on the judgment of geological experts, the standard is determined to be δ. 18 O value, δD value, Sr / Ca ratio, and Mg / Ca ratio are used as markers to reflect paleodepositional environments. These geochemical indicators have been extensively studied and proven to be closely related to temperature and climatic conditions, effectively reflecting the characteristics of paleoelevation changes. Therefore, the selection of these indicators as markers has sufficient scientific basis.

[0077] S32, Sample Screening: In step S32, samples with complete geochemical indicators and reliable paleolatitude, longitude, and paleoage information are screened from the dataset. The screening process needs to ensure that the screened samples are representative in terms of spatial distribution and time span, covering the diversity of different geographical regions and geological periods, and avoiding annotation bias that would limit the generalization ability of the model.

[0078] S33, Label Assignment; In step S33, geological experts assign label values ​​to the corresponding elevation values ​​of each sample based on the paleosedimentary environment information and geological background. The geological experts will combine existing scientific methods, such as climate modeling and geological survey results, to meticulously label the samples, ensuring the scientific validity and accuracy of the label values.

[0079] S34. Labeling Quality Verification: In step S34, labeling quality verification is performed using a multi-expert independent labeling method. For the same batch of samples, at least two experts assign labels separately, and the consistency of the experts' labeling results is compared. For samples with inconsistent labels, verification is required, and consensus is reached through discussion or controversial samples are removed. This multi-labeling and cross-validation quality control measure effectively improves the consistency and accuracy of labeling, ultimately generating a labeled training dataset.

[0080] Step S4 is the model building stage. The purpose of this stage is to divide the labeled training dataset into training and test sets, train the model on the training set using the LightGBM algorithm, evaluate the prediction error using a loss function, and finally build the paleotopography inference model. LightGBM is an efficient machine learning algorithm based on gradient boosting decision trees, which has significant advantages in processing large-scale, high-dimensional data.

[0081] S41. Dataset Partitioning: In step S41, the labeled training dataset is randomly divided into a training set and a test set in a 7:3 ratio. The training set is used for model parameter learning, and the test set is used for model performance evaluation. This partitioning method ensures that there is sufficient data for model training, while reserving some untrained data for objective evaluation of the model's generalization ability on unseen data, effectively preventing model overfitting.

[0082] S42, LightGBM Model Training: In step S42, the LightGBM algorithm is used to iteratively train the training set. LightGBM is a gradient boosting framework that, compared with traditional gradient boosting decision trees, has advantages such as fast training speed, low memory consumption, and high prediction accuracy. The algorithm performs the following operations in each iteration: First, the continuous values ​​of each feature are discretized into several intervals and a histogram is constructed. This histogram-based decision tree construction method significantly reduces computational complexity and memory consumption. Then, a leaf-by-leaf growth strategy is used to construct the decision tree. The leaf node with the largest split gain is selected from all leaf nodes of the current decision tree and split to generate new child nodes. Compared with the traditional layer-by-layer growth strategy, the leaf-by-leaf growth strategy can reduce more errors and obtain better prediction accuracy with the same number of splits. To prevent overfitting, the algorithm sets a maximum depth parameter to limit the uncontrolled growth of the decision tree. Iteration continues until the stopping condition is met.

[0083] LightGBM also employs gradient-based one-sided sampling and mutually exclusive feature bundling techniques. Gradient-based one-sided sampling retains samples with large gradients and randomly samples those with small gradients, accelerating training while maintaining prediction accuracy. Mutually exclusive feature bundling combines mutually exclusive features, reducing the number of features and further improving training efficiency. These techniques enable LightGBM to efficiently learn key patterns from high-dimensional data and accurately capture the complex nonlinear features that determine paleotopographic changes.

[0084] S43. Loss Function Calculation; In step S43, after each iteration, the root mean square error (RMSE) is used as the loss function to evaluate the prediction error of the current model on the training set. RMSE is a commonly used metric to measure the deviation between predicted and actual values, accurately reflecting the model's prediction accuracy. Its calculation formula is as follows: ;

[0085] In the formula, The root mean square error is expressed in meters. The number of samples in the training set; For the first The actual elevation values ​​of each sample, in meters; For the first The model predicts the elevation value for each sample, in meters.

[0086] S44, Model Output: In step S44, training stops and the paleotopic inference model is output when the loss function value converges or reaches the preset maximum number of iterations. Loss function convergence means that after several consecutive iterations, the decrease in the loss function value is less than a preset threshold, indicating that the model has fully learned the patterns in the training data. The maximum number of iterations is a protection mechanism to prevent excessively long training times and can be set according to specific application scenarios.

[0087] Step S5 is the model evaluation and optimization stage. The purpose of this stage is to evaluate the model's generalization ability using K-fold cross-validation, optimize model performance through hyperparameter tuning and regularization, and finally output the optimized paleotopographic inference model. Model evaluation and optimization are crucial steps to ensure the model possesses good predictive ability and stability.

[0088] S51, K-fold cross-validation; In step S51, the labeled training dataset is randomly shuffled and then divided equally into... A set of mutually disjoint subsets The value is 5. K-fold cross-validation is a commonly used model evaluation method that can more comprehensively evaluate the model's generalization ability. The specific operation is as follows: Perform... In each iteration of the verification process, one subset is selected as the verification set, and the rest... A subset is used as the training set, the model is trained on the training set, and the accuracy is calculated on the validation set. After the iteration is completed, the calculation The average of the validation accuracies is used as the model's cross-validation accuracy, calculated using the following formula: In the formula, For cross-validation accuracy; The value is 5; For the first Accuracy on the validation set.

[0089] K-fold cross-validation, by splitting the dataset multiple times and performing training and validation separately, effectively reduces evaluation bias caused by the randomness of data partitioning, resulting in more stable and reliable model performance estimates. When the value is 5, 80% of the data is used for training and 20% for validation each time, which ensures both sufficient training data and a relatively accurate estimate of generalization performance.

[0090] S52. Model Performance Evaluation; In step S52, the F1 score of the model is calculated on the test set. The F1 score is the harmonic mean of precision and recall, which comprehensively reflects the classification performance of the model, and is especially suitable for class imbalance situations. Its calculation formula is as follows: ;

[0091] In the formula, The F1 score ranges from 0 to 1, with a higher value indicating better model performance. Precision rate represents the proportion of samples that the model predicts to be positive but are actually positive. Recall rate represents the proportion of samples that are actually positive but were correctly predicted as positive by the model.

[0092] The F1 score takes into account both precision and recall, avoiding the bias that can result from using only a single metric. A high F1 score indicates that the model performs well in both correctly identifying positive examples and avoiding false positives, demonstrating good overall performance.

[0093] S53. Hyperparameter Tuning; In step S53, a grid search method is used to tune the hyperparameters. Hyperparameters are parameters that need to be preset before model training, such as the learning rate, maximum tree depth, minimum number of samples per leaf node, etc. These parameters have a significant impact on model performance. The grid search method traverses all possible hyperparameter combinations within the preset hyperparameter value range, performs K-fold cross-validation evaluation on each set of hyperparameters, and selects the hyperparameter combination that maximizes the cross-validation accuracy as the optimal hyperparameters. Although grid search has a high computational cost, it can systematically search the hyperparameter space to find the optimal or near-optimal parameter configuration.

[0094] S54. Regularization Processing: In step S54, an L2 regularization term is introduced during model training to control model complexity and prevent overfitting. L2 regularization limits the range of model parameter values ​​by adding a penalty term to the sum of squared model parameters in the loss function, making the model tend to learn smoother decision boundaries. The regularization strength is controlled by the regularization coefficient; the larger the regularization coefficient, the stronger the penalty on model complexity. By reasonably setting the regularization coefficient, a balance can be achieved between the model's fitting ability and generalization ability, ultimately outputting an optimized paleoterrain inference model.

[0095] Furthermore, step S5 includes a feature contribution analysis sub-step. The SHAP method is used to analyze the contribution of each input feature to the model's prediction results. SHAP is a game theory-based model interpretation method that calculates the marginal contribution of each input feature to the prediction results. Specifically, for each sample to be analyzed, the SHAP value of each feature is calculated. A positive SHAP value indicates a positive contribution to the prediction results, while a negative SHAP value indicates a negative contribution. The larger the absolute value of the SHAP value, the higher the contribution of the feature. The SHAP values ​​of all features for all samples are summarized, and the average of the absolute values ​​of the SHAP values ​​of each feature is calculated. These features are then sorted from largest to smallest by average value to identify the key features that contribute the most to paleotopographic prediction. This feature importance analysis not only helps in understanding the model's prediction mechanism but also provides valuable scientific insights for geological research.

[0096] Step S6 is the application prediction stage. The purpose of this stage is to preprocess and transform the multi-source geological feature data of unknown samples in the same way as the training data, input them into the optimized paleotopographic inference model, output the paleotopographic prediction results and display them visually.

[0097] S61. Preprocessing of data to be predicted: In step S61, stratigraphic lithology, paleolatitude and longitude, paleoage, and geochemical index data of unknown samples are obtained, and the data are preprocessed using the same data cleaning and normalization methods as in step S1. Ensuring that the processing method of the data to be predicted is completely consistent with that of the training data is an important prerequisite for ensuring the accuracy of the prediction. Any inconsistency in the processing method may lead to deviations in the prediction results.

[0098] S62, Feature Transformation; In step S62, following the same feature transformation method as in step S2, the preprocessed data undergoes categorical feature numerical transformation and spatiotemporal feature construction to generate a feature vector to be predicted that is consistent with the format of the model input feature set. Consistency in feature transformation is also a key requirement for ensuring prediction accuracy.

[0099] S63. Model Prediction: In step S63, the feature vector to be predicted is input into the optimized paleotopographic inference model. The model performs calculations through its internal decision tree ensemble. Each decision tree sequentially splits nodes based on the input features, eventually reaching a leaf node and outputting a predicted value. The prediction results of all decision trees are weighted and averaged to obtain the final predicted value. The model outputs the paleotopographic prediction value for each unknown sample, and can also output the confidence interval of the prediction results, providing a reference for geological researchers.

[0100] S64. Result Visualization: In step S64, the predicted paleotopographic values ​​are visualized according to the spatial location information of the samples. Visualization methods include contour maps, heat maps, or 3D topographic models. Contour maps clearly show the undulations and elevation changes of the paleotope; heat maps intuitively reflect the elevation differences between different regions through color variations; and 3D topographic models can present the spatial morphology of the paleotope in three dimensions. Visualization facilitates in-depth analysis and interpretation of the prediction results by geological researchers, providing data support for regional geological evolution and tectonic movement research, and completing the visualization of the paleotopographic inference results.

[0101] Example 1: This example focuses on the sedimentary environment of the ocean-continent transition zone. The ocean-continent transition zone is a transitional area between the ocean and the land. Its sedimentary environment is jointly controlled by multiple dynamic mechanisms such as tides, waves, runoff input, and basin closure. Its geological features combine marine and terrestrial characteristics, making it a complex area and a typical difficult area for paleotopographic inference.

[0102] This embodiment selects multi-source data, including lithological type, paleolatitude and longitude, geochemical indices, and sample age characteristics, as model input. The geochemical indices include Sr / Ca ratio, Mg / Ca ratio, and δ¹⁸O. 18 O-values ​​and δD-values, along with age characteristics including minimum age and early stratigraphic intervals, are used. After normalization and missing value imputation in step S1, feature encoding and feature engineering are performed in step S2. The processed data is then input into the trained LightGBM model for classification and feature contribution analysis.

[0103] The SHAP value and feature interaction analysis method were used to quantitatively evaluate the feature contribution of the ocean-continent transition zone samples. Furthermore, typical microenvironment categories were selected for validation, including coastal zone types, intertidal shallow beach environments, and littoral zones, to test the applicability of the model under different dynamic backgrounds.

[0104] The LightGBM model demonstrated excellent and stable classification performance in the ocean-continent transition zone environment. As shown in Table 1, the model achieved an accuracy of 0.8576 and an F1 score of 0.8529 on the ocean-continent transition zone test set, indicating that the model has strong discriminative ability. Furthermore, the robustness of the model was further tested through 5-fold cross-validation, with accuracies of 0.8437, 0.8697, 0.8414, 0.8570, and 0.8459 for each fold, and a mean cross-validation accuracy of 0.8515. The accuracy remained stable within a narrow range; this extremely low inter-fold fluctuation indicates that the model maintains highly consistent performance across different subsets of the data, demonstrating strong generalization ability, effectively avoiding overfitting, and ensuring the reliability of the prediction results.

[0105] Table 1 shows the overall accuracy, F1 score, and fold-by-fold accuracy and average accuracy of the LightGBM model in the ocean-continent transition zone environment. The table demonstrates that the model exhibits good prediction accuracy and stability when handling complex environmental data from the ocean-continent transition zone.

[0106] Table 1. Overall accuracy, F1 score, and accuracy and average accuracy of the LightGBM model in the ocean-continent transition zone environment.

[0107] environment type overall accuracy Overall F1 Score k-fold cross-validation accuracy (k=5) - cv1 k-fold cross-validation accuracy (k=5) - cv2 k-fold cross-validation accuracy (k=5) - cv3 k-fold cross-validation accuracy (k=5) - cv4 k-fold cross-validation accuracy (k=5) - cv5 cross-validation average accuracy ocean-continent transition zone 0.8576 0.8529 0.8437 0.8697 0.8414 0.8570 0.8459 0.8515

[0108] Feature contribution analysis, such as Figures 1-5 As shown. Figure 1 A scatter plot of the SHAP feature contribution to paleolar latitude and longitude accuracy is used to quantify the marginal impact of paleolar latitude and longitude data accuracy on the model's paleotopographic prediction results. The horizontal axis represents paleolar latitude and longitude accuracy, ranging from 0.0 to 1.0, with values ​​closer to 1 indicating higher accuracy in paleolar latitude and longitude measurement or estimation. The vertical axis represents the SHAP value, with positive values ​​indicating a positive effect on prediction results and negative values ​​indicating an inhibitory effect. The analysis shows a positive correlation between paleolar latitude and longitude accuracy and the model's predictive contribution. High-accuracy paleolar latitude and longitude significantly promotes paleotopographic prediction, while low accuracy tends to have an inhibitory effect or no effective contribution. This indicates that the quality of paleolar latitude and longitude data is a key prerequisite for model prediction; high-accuracy paleolar latitude and longitude can effectively correct climatic gradient biases in geochemical indicators, while low accuracy introduces noise interference.

[0109] Figure 2 A scatter plot analyzing the contribution of paleolatitude to SHAP features was used to illustrate the impact of this spatial feature on model predictions and the differences between the Northern and Southern Hemispheres. The horizontal axis represents paleolatitude, ranging from -75° to 75°, with negative values ​​representing the Southern Hemisphere and positive values ​​representing the Northern Hemisphere; the vertical axis represents SHAP values; the color gradient of the scatter plots corresponds to lithological complexity, with darker colors representing complex lithological combinations and lighter colors representing simpler lithologies. The analysis shows that the promoting effect of paleolatitude on predictions is dominated by the Northern Hemisphere, with SHAP values ​​in the mid-to-high latitude regions of the Northern Hemisphere generally higher than those in the Southern Hemisphere. This phenomenon reflects the asymmetry in continental tectonics between the Northern and Southern Hemispheres during historical periods. The Northern Hemisphere has a richer variety of lithological types and a more complex geological environment, leading to a stronger constraint of paleolatitude on paleotopographic predictions. Paleolatitude and lithology are spatial projections of paleomagnetic and plate tectonics information, and the difference in their contributions from samples in the Northern and Southern Hemispheres reflects the differences in continental tectonics over historical periods.

[0110] Figure 3A horizontal bar chart showing the SHAP contribution values ​​of environmental features in the ocean-continent transition zone is used to rank and display the average contribution intensity of all input features to the model's paleotopographic predictions. The horizontal axis represents the average SHAP value; a higher value indicates a more critical contribution of the feature to the prediction. The vertical axis represents the feature name. Analysis shows a clear hierarchy of feature contributions: lithological features have the highest average SHAP value, ranking first, followed by sample set name and preservation mode. Paleolatitude-longitude related features also rank among the top contributors. Geological structural features and spatial constraint features are the core dependencies of the model's predictions, while temporal features provide temporal constraints, and geochemical features synergistically enhance the effects of geological structural features. This hierarchical distribution demonstrates the rationality of the design of selectively integrating high-value features in this invention, rather than simply overlaying all data.

[0111] Figure 4 This is a SHAP swarm map of the ocean-continental transition zone environment. The distribution of SHAP values ​​for individual samples demonstrates the heterogeneity and dispersion of the predictive impact of each feature. The horizontal axis represents the SHAP value, indicating the marginal impact of that feature on the prediction result in a single sample; the vertical axis represents the feature name, arranged from top to bottom according to the average contribution intensity of the feature; the color gradient of the scatter points corresponds to the magnitude of the feature values. The analysis shows that the scatter point distribution of some features, such as lithology, paleolatitude and longitude, and paleoage, is relatively concentrated, indicating that their predictive contributions are stable and universal, unaffected by individual sample differences; while the scatter point distribution of some features, such as preservation patterns, is relatively dispersed, indicating that their contributions fluctuate greatly and are strongly influenced by the local dynamic environment. This suggests that paleotopographic inference has stable constraint factors and variable influencing factors. Stable factors provide the basic predictive framework, while variable factors explain the predictive differences in local areas. The model can simultaneously capture common patterns and individual differences.

[0112] Figure 5 A scatter plot of the SHAP interaction analysis between lithology and paleolatitude is used to quantify the marginal impact of the synergistic effect of lithology and paleolatitude on model predictions. The horizontal axis represents the lithology coding value, which is the numerical result of lithology type after feature hashing; the vertical axis represents the SHAP interaction value of paleolatitude, representing the marginal contribution of the synergistic effect of paleolatitude and this lithology combination to the prediction; the color gradient of the scatter points corresponds to the intensity of the interaction effect, with red representing a strong interaction effect and blue representing a weak interaction effect. The analysis results show that the interaction effect between lithology and paleolatitude is threshold-dependent, with the strongest interaction at moderate lithological complexity. With a single lithology, the contribution of paleolatitude is relatively independent, and the interaction effect is weak. In areas with drastic changes in dynamic gradient, such as tidal flats and delta fronts, the red scatter points are the most dense, indicating a significant nonlinear interaction relationship between lithology and paleolatitude in these areas. This interaction effect is key to improving the prediction accuracy of the model in complex environments and can explain the topographic differences caused by different lithologies at the same paleolatitude.

[0113] From the perspective of microenvironmental classification in the ocean-continental transition zone, the model demonstrates hierarchical predictive ability for microenvironments with different dynamic backgrounds. Coastal zones and similar environments have a large sample size, maintaining a high level of accuracy. While the sample size for intertidal shallow beach environments is smaller, the accuracy is actually higher because these environments are primarily controlled by tidal forces, resulting in more concentrated and significantly different sedimentary characteristic signals. The accuracy of littoral zone environments decreases due to complex dynamic factors and high noise levels. Analyzing from the perspective of the external forces acting on the environment, environments controlled by a single or dominant external force exhibit more concentrated sedimentary systems in the feature space, leading to higher classification accuracy. Conversely, environments with diverse dynamic mechanisms and complex internal structures exhibit higher uncertainty.

[0114] Experimental results show that the model's recognition performance is significantly improved in small environments with relatively simple sedimentary dynamics or clearly defined controlling factors; while environments with diverse dynamic sources exhibit higher uncertainty. This pattern is highly consistent with geological understanding of sedimentary environments, proving the scientific rationality of the model's predictions.

[0115] In summary, the ocean-continental transition zone is one of the most complex and unstable sedimentary environments, exhibiting high dispersion in lithology, preservation patterns, and geochemical indicators, increasing the difficulty of prediction. However, the multi-source data fusion model based on LightGBM constructed in this invention can still stably identify key control variables of the transitional facies environment, accurately capture the nonlinear relationships between lithology, paleolatitude and longitude, and preservation patterns, maintain hierarchical predictive capabilities for microenvironments under different dynamic backgrounds, and maintain overall reliable predictive performance in highly heterogeneous environments. This embodiment demonstrates that the method of this invention can effectively adapt to the complex sedimentary system of the ocean-continental transition zone, possessing strong generalization ability and geological interpretation value, and providing reliable technical support for paleotopographic reconstruction and sedimentary environment identification.

[0116] This invention's method is applicable to paleotopographic inference in various geological environments, including ocean-continental transition zones, plateaus, orogenic belts, and foreland basins. In geological research, this method can be used to reconstruct plateau uplift processes, simulate geological evolution processes, and explore crustal movements and tectonic activities in different geological environments during different geological periods. This invention integrates multi-source data on stratigraphy, lithology, paleospatial and geochemical data, enabling the revelation of changes in marine sedimentary basins, dynamic processes in subduction zones, and sediment exhumation processes, providing crucial support for cross-environmental research in geology.

[0117] In paleoclimate analysis, the undulations of paleotopography directly influence atmospheric circulation and precipitation distribution. The accurately reconstructed paleotopography using this method can be used to calibrate paleoclimate models and analyze the correlation between topographic changes and climate events, which is of great significance for paleoclimate reconstruction and climate model calibration. In geochemical analysis, this method models the relationship between multi-source geological data and paleotopography, providing new quantitative modeling methods and geological tools for research on different geological environments and geological eras, and promoting the interdisciplinary integration of geochemistry, stratigraphy, and artificial intelligence.

[0118] Compared to traditional empirical formulas and experimental analysis methods, the method of this invention has significant advantages. In terms of efficiency, the LightGBM model can efficiently process multi-source heterogeneous data, significantly shortening data processing and prediction time, and its fully automated modeling process reduces manual intervention. Regarding high accuracy, LightGBM can automatically capture the complex nonlinear coupling relationship between multi-source geological data and paleotopography, avoiding errors and biases inherent in manual analysis, and significantly improving the accuracy of paleoelevation prediction through data-driven nonlinear fitting. In terms of broad applicability, the model possesses strong generalization ability, adapting to changes in different geographical regions and geological eras, and can be extended to various regions and types of geological research worldwide.

[0119] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for inferring paleotopography of an ocean-continent transition zone based on geochemical data and machine learning, characterized in that, The method comprises the following steps: S1, data preparation: collecting multi-source geological data including stratum lithology, paleo latitude and longitude, paleo age and geochemical index, and corresponding elevation data, performing data cleaning, data normalization processing and missing value filling processing on the collected data, and constructing a multi-source geological data set; S2, feature engineering: numerically converting the category type features in the multi-source geological data set, constructing spatial and temporal features, and selecting features related to elevation through correlation analysis to construct a model input feature set; S3, manual labeling: based on the judgment of geology experts and the sample paleo sedimentary environment information, the samples in the data set are labeled to generate a labeled training data set; S4, model construction: the labeled training data set is divided into a training set and a test set, the training set is trained using the LightGBM algorithm, the prediction error is evaluated using the loss function, and a paleo topography inference model is constructed; S5, model evaluation and optimization: the generalization ability of the model is evaluated using the K-fold cross-validation method, the model performance is optimized through hyperparameter adjustment and regularization processing, and the optimized paleo topography inference model is output; S6, application prediction: after pre-processing and feature conversion of the multi-source geological feature data of unknown samples in the same way as the training data, the optimized paleo topography inference model is input, and the paleo topography prediction result is output and visualized.

2. The method of claim 1, wherein, The data preparation in step S1 specifically includes the following sub-steps: S11, data collection: collecting stratigraphic lithology data, paleolatitude data, paleoage data, and geochemical index data of the rock, sediment or fossil sample, the geochemical index data including δ 18 O value, δD value, Sr / Ca ratio value and Mg / Ca ratio value; collecting modern elevation data or geological record elevation data corresponding to the sample at the same time; S12, data cleaning: the distribution of each feature is counted, the mean and standard deviation of each feature value are calculated, data points deviating from the mean by more than 3 times the standard deviation are identified as outliers and removed, and a cleaned data set is obtained; S13, data normalization processing: for each numerical feature in the cleaned data set, Min-Max normalization processing is performed according to the following formula: ; In the formula, X is the normalized feature value, the value range is 0 to 1; X is the original feature value; Xmin is the minimum value of the feature in the data set; Xmax is the maximum value of the feature in the data set; S14, missing value filling processing: for numerical features with missing values in the data set, the mean of all non-missing values of the feature is used for filling, and a complete multi-source geological data set is obtained.

3. The method of claim 1, wherein, The feature engineering in step S2 specifically includes the following sub-steps: S21, numerical conversion of category type features: for stratum lithology types, sedimentary facies types and other category type features, a feature hashing method is used to convert them into numerical features, and the specific operation is as follows: the hash value of each category value is calculated, and the hash value is mapped to a predetermined numerical interval to generate corresponding numerical feature values; S22, spatial and temporal feature construction: for paleo latitude and longitude data, the spherical distance of each sample point to the paleo equator is calculated as a new spatial feature; For paleo age data, segment coding is performed according to geological eras or epochs to generate time segment features; S23, correlation analysis and feature screening: the Pearson correlation coefficient between each feature and the elevation is calculated, which is calculated according to the following formula: ; In the formula, r is a Pearson correlation coefficient, and the value range is -1 to 1, and the greater the absolute value is, the stronger the correlation is; Xi is a characteristic value of the i th sample; is the mean value of the characteristic value of all samples; Yi is an elevation value of the i th sample, and the unit is: meter; is the mean value of the elevation value of all samples, and the unit is: meter; the characteristics with an absolute value of the Pearson correlation coefficient greater than a preset threshold value are screened, and a model input characteristic set is constructed.

4. The method of claim 1, wherein, The manual labeling in step S3 specifically includes the following sub-steps: S31, Labeling standard setting: According to the judgment of geological experts, δ 18 O value, δD value, Sr / Ca ratio and Mg / Ca ratio are determined as the labeling basis reflecting the paleo-depositional environment; S32, sample screening: samples with complete geochemical index, reliable paleo latitude and longitude and paleo age information are screened from the data set to ensure that the screened samples are representative in spatial distribution and time span; S33, label assignment: based on the paleoenvironmental information and geological background of each sample, the geologist assigns the corresponding elevation value of each sample as a label value; S34, annotation quality verification: using the multi-expert independent annotation method, at least two experts independently assign labels to the same batch of samples, compare the consistency of the annotation results of each expert, review and confirm the samples with inconsistent annotations, and generate a labeled training dataset.

5. The method of claim 1, wherein, The model construction in step S4 specifically includes the following sub-steps: S41, data set division: divide the labeled training data set into a training set and a test set according to a 7:3 ratio, wherein the training set is used for model parameter learning, and the test set is used for model performance evaluation; S42, LightGBM model training: using the LightGBM algorithm to iteratively train the training set, which performs the following operations in each iteration: first, discretize the continuous values of each feature into several intervals and construct a histogram, then use the leaf growing strategy to select the leaf node with the maximum split gain from all leaf nodes of the current decision tree to split and generate new child nodes, and iterate until the stop condition is met; S43, loss function calculation: after each iteration, the root mean square error is used as the loss function to evaluate the prediction error of the current model on the training set, calculated according to the following formula: ; In the formula, L is the root mean square error, the unit is: meters; n is the number of samples in the training set; Yi is the actual elevation value of the ith sample, the unit is: meters; is the model predicted elevation value of the ith sample, the unit is: meters; S44, model output: when the loss function value converges or reaches the preset maximum number of iterations, stop training and output the paleotopographic inference model.

6. The method of claim 1, wherein, The model evaluation and optimization in step S5 specifically includes the following sub-steps: S51, K-fold cross-validation: the labeled training dataset is randomly shuffled and evenly divided into K non-intersecting subsets, K is 5; K iterations are performed, and each iteration selects one subset as the validation set and the remaining K-1 subsets as the training set; the model is trained on the training set and the accuracy is calculated on the validation set; after K iterations, the average of the K validation accuracies is calculated as the cross-validation accuracy of the model, calculated according to the following formula: ; In the formula, Acv is the cross-validation accuracy; K is the number of folds, which is 5; Ak is the accuracy on the k-fold validation set; S52, Model performance evaluation: Calculate the F1 score of the model on the test set, calculated as follows: ; In the formula, F1 is the F1 score, the value range is 0 to 1; P is the precision, which represents the proportion of actual positive samples in the samples predicted as positive by the model; R is the recall rate, which represents the proportion of samples correctly predicted as positive by the model among the actual positive samples; S53, hyperparameter adjustment: using the grid search method, traverse all hyperparameter combinations within the preset hyperparameter value range, and select the hyperparameter combination with the highest cross-validation accuracy as the optimal hyperparameter; S54, regularization processing: introduce an L2 regularization term in the model training process to control the model complexity and prevent overfitting, and output the optimized paleotopographic inference model.

7. The method of claim 1, wherein, The application prediction in step S6 specifically includes the following sub-steps: S61, pre-processing of data to be predicted: obtaining the lithology, paleolatitude, paleoage, and geochemical index data of unknown samples, and pre-processing the data according to the same data cleaning and normalization method in step S1; S62, feature conversion: according to the same feature conversion method in step S2, the pre-processed data is numerically converted and spatial-temporal feature is constructed, and a feature vector consistent with the model input feature set format is generated; S63, model prediction: input the predicted feature vector into the optimized paleotopographic inference model, and the model calculates through its internal decision tree set to output the paleotopographic prediction value of each unknown sample; S64, result visualization: plot the predicted paleotopographic value according to the spatial position information of the sample into a contour map, a heat map, or a three-dimensional terrain model to complete the visualization display of the paleotopographic inference result.

8. The method of claim 1, wherein, The step S5 further includes a feature contribution degree analysis sub-step: the SHAP method is used to analyze the contribution degree of each input feature to the model prediction result, and the specific operation is as follows: for each sample to be analyzed, the SHAP value of each feature is calculated, the SHAP value is positive, indicating that the feature has a positive contribution to the prediction result, the SHAP value is negative, indicating that the feature has a negative contribution to the prediction result, and the greater the absolute value of the SHAP value, the higher the contribution degree of the feature; the SHAP values of all samples are summarized, the average value of the absolute values of the SHAP values of each feature is calculated, the average values are sorted from large to small, and the key features with the greatest contribution to the prediction of the ancient landform are identified.

Citation Information

Patent Citations

  • Method for reconstructing paleo-elevation of mountain-building zone based on paleontological data and geochemical data

    CN119942019A

  • HY-2A satellite ocean water vapor inversion method based on machine learning

    CN120632292A