Paleontology diversity analysis method based on dynamic machine learning joint correction model

Through the dynamic interval division and data-driven weight allocation of dynamic machine learning models, the deviation problem of traditional paleodiversity analysis methods in fossil record inhomogeneity is solved, and accurate evaluation of fossil record and efficient identification of key events is achieved.

CN120561582APending Publication Date: 2025-08-29HOHAI UNIV

Patent Information

Application Number
CN202510636552.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-17
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Traditional paleodiversity analysis methods are difficult to adapt to the spatial and temporal inhomogeneity of fossil records, resulting in bias in diversity change assessment and key event identification, and lack of data adaptive mechanisms.

Method used

The dynamic machine learning joint correction model is adopted, and through dynamic interval division and data-driven weight allocation, combined with change point detection, hierarchical clustering and hybrid methods, the natural boundaries of fossil records are adaptively identified and the weight is adjusted to achieve accurate evaluation of diversity analysis.

Benefits of technology

It improves the ability to adapt to imbalances to fossil records, enhances the ability to identify key biological events, and provides a more accurate diversity analysis framework.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561582A_ABST
    Figure CN120561582A_ABST
Patent Text Reader

Abstract

The invention discloses a paleontology diversity analysis method based on a dynamic machine learning joint correction model, and the method comprises the steps: obtaining a paleontology data sample, and carrying out the missing value processing, abnormal value detection, data standardization and time data mapping preprocessing of the paleontology data sample; performing dynamic interval division on the preprocessed paleontology data sample, and realizing adaptive time interval division by adopting multiple algorithms; non-uniform weight distribution is realized through data driving weights; enhanced diversity analysis is carried out on an optimized data set, including interval division processing on the data set, feature evidence extraction and interval weight calculation, integrated interval weighted diversity dynamic calculation, weighted survivor analysis, interval sub-sampling analysis and other methods, so that diversity calculation is more accurate; and outputting a diversity curve, a data table and a performance evaluation result. The method better adapts to the imbalance of fossil records, and the recognition capability of key biological events is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a paleontological biodiversity analysis method based on a dynamic machine learning joint correction model, and belongs to the technical field of paleontological data analysis. Background Art

[0002] Before the mid-20th century, paleontology focused primarily on systematic classification, resulting in the accumulation of a vast amount of literature and fossil data. However, with the rapid development of big data analytics and artificial intelligence, paleontology is shifting from traditional statistical analysis to more complex, dynamic, and multidimensional analytical methods. Furthermore, research on paleobiodiversity (or biological diversity) is attracting increasing attention.

[0003] The study of paleobiodiversity can reveal the history of the biosphere and the relationship between environmental change and fluctuations in diversity. It can also provide insights into how the fossil record informs current biodiversity issues. The functions and services of biodiversity are fundamental to sustainable development and human well-being. These services include primary production, nutrient cycling, water and air purification, mitigation of greenhouse gas emissions and climate change, crop pollination, food and genetic resource provision, disease control, and educational, cultural, and spiritual benefits. Related research can enrich our understanding of the mechanisms of evolution of life on Earth and provide further clues to our understanding of the impact of global biodiversity changes on human evolution.

[0004] In the current era of big data informatization, a wealth of fossil specimen data has been accumulated, providing a solid foundation for studying paleobiodiversity dynamics. Paleobiological data are multi-source and heterogeneous, encompassing a variety of data types, including fossil occurrence records, geological time periods, paleoenvironmental data, and morphological characteristics. These data are often subject to uneven temporal and spatial distribution and preservation biases. While traditional paleobiodiversity analysis methods offer rich computational capabilities and have been widely applied, they suffer from inherent limitations, primarily including fixed time intervals, uniformly weighted statistics, and a lack of data adaptation mechanisms.

[0005] Although traditional paleobiodiversity analysis methods have been widely used in paleontological research and provide a wealth of functions such as diversity indicator calculation, extinction rate and origin rate estimation, and sampling bias correction, they generally have some inherent limitations, mainly manifested as: the preset fixed time interval makes it difficult to capture the key nodes of diversity evolution; the use of uniform weight statistics makes it impossible to reflect the uneven distribution characteristics of fossil records; the lack of data adaptation mechanism causes systematic biases when processing non-steady-state fossil records, etc. Summary of the Invention

[0006] Purpose of the invention: In response to the problems and shortcomings of the existing technology, the present invention provides a paleodiversity analysis method based on a dynamic machine learning joint correction model. Based on the machine learning algorithm, by introducing a dynamic interval division strategy and a data-driven weight allocation mechanism, it solves the statistical bias problem of traditional methods when processing non-steady-state fossil records, and provides a more reliable analysis framework for the study of diversity dynamics.

[0007] Technical Solution: Traditional paleodiversity analysis methods typically use preset uniform time windows and uniformly weighted statistics, making it difficult to adapt to the uneven temporal and spatial distribution of fossil records, leading to deviations in the assessment of diversity changes and the identification of key events. This method identifies the natural boundaries of the fossil record through dynamic interval partitioning, and then uses data-driven weight allocation to enhance the weights of key events. The two modules are closely integrated to form a complete analysis chain. Through enhanced diversity analysis, the system integrates functions such as boosted tree weight calculation and enhanced survivor analysis, realizing the migration of traditional diversity analysis methods to improved versions of diversity analysis methods.

[0008] A paleodiversity analysis method based on a dynamic machine learning joint calibration model includes the following steps:

[0009] Step 1: Obtain paleontological data samples and preprocess the paleontological data samples by missing value processing, outlier detection, data standardization and time data mapping.

[0010] In step 2, the preprocessed paleontological data samples were dynamically partitioned into intervals, and a variety of algorithms were used to achieve adaptive time interval partitioning, including: change point detection, hierarchical clustering (bottom-up clustering and continuity constraints), and hybrid methods (AIC evaluation connection and interval merging classification).

[0011] Step 3: Implement non-uniform weight distribution through data-driven weighting to obtain an optimized dataset. The data-driven weighting methods include: feature importance weighting (XGBoost algorithm, feature importance and adaptive weighting), Bayesian optimization weighting (Gaussian process regression, hierarchical Bayesian analysis and handling uncertainty).

[0012] Step 4 enhances diversity analysis by performing interval partitioning on the dataset, extracting feature evidence, and calculating interval weights to make diversity calculations more accurate. Outputs include "diversity curves" (taxonomic diversity, true vs. optimized comparisons, and extinction and origin rates), "data tables," and "performance evaluation results," providing comprehensive evaluation and visualization. Regarding interval partitioning, various proposed interval partitioning mechanisms, including statistically-based interval partitioning (change point detection), hierarchical clustering, and hybrid methods, can adaptively determine the optimal time window based on the characteristics of the fossil data itself, identifying key nodes of biodiversity change. Regarding weight assignment, data-driven interval weight calculation is achieved through feature importance assessment and Bayesian optimization, overcoming the statistical bias caused by the uniform weight assumption in traditional analysis methods.

[0013] The present invention organically combines dynamic interval division with adaptive weight allocation. Compared with traditional analysis methods, the model better adapts to the imbalance of fossil records, improves the ability to identify key biological events, and has more prominent advantages in processing non-steady-state fossil records.

[0014] The change point detection automatically identifies key change points in data distribution by dynamically analyzing the variance changes of time series, and automatically marks them as interval boundaries when the variance changes exceed an adaptive threshold.

[0015] The hierarchical clustering provides accurate identification of critical moments such as extinction events. Starting from the perspective of data similarity, hierarchical clustering automatically constructs a similarity matrix and forms optimal clusters through Euclidean distance and Ward method. It can automatically adjust the clustering level according to the complexity of the data and group similar time points in the multidimensional feature space into the same interval.

[0016] The hybrid approach first tries multiple interval partitioning methods (change-point detection, hierarchical clustering, and adaptive windowing) in parallel and then quantitatively evaluates the results of each method using the AIC information criterion. The AIC value for each method is calculated using a correlation function based on the residual sum of squares, the number of intervals, and the sample size. The method with the best AIC value is automatically selected as the final solution, achieving data-driven selection.

[0017] The adaptive dynamic interval partitioning can automatically select the partitioning algorithm that best suits the current data characteristics. For each data set, the three algorithms are first executed in parallel to obtain different interval partitioning results. Then, the results of each method are quantitatively evaluated by the AIC information criterion (AIC = 2k + n ln (RSS / n)), where k is the number of intervals (complexity), n is the number of samples, and RSS is the residual sum of squares (fitting degree). The system automatically selects the method with the best AIC value as the final interval partitioning scheme to achieve data-driven selection. This adaptive selection mechanism improves the scientific nature of interval partitioning and provides a more accurate analysis framework for paleobiodiversity research.

[0018] A paleobiodiversity analysis system based on a dynamic machine learning joint calibration model, including:

[0019] The data preprocessing module obtains paleontological data samples and performs missing value processing, outlier detection, data standardization and time data mapping.

[0020] The dynamic interval partitioning module uses multiple algorithms to achieve adaptive time interval partitioning, including: change point detection, hierarchical clustering (bottom-up clustering and continuity constraints), and hybrid methods (AIC evaluation connection and interval merging classification).

[0021] Data-driven weight module to achieve non-uniform weight distribution, including: feature importance weight (XGBoost algorithm, feature importance and adaptive weight), Bayesian optimization weight (Gaussian process regression, hierarchical Bayesian analysis and handling uncertainty).

[0022] After the above processing, the system obtains the "optimized data set", which has realized interval optimization processing, weighted data generation and interval weight mapping.

[0023] The enhanced diversity analysis module makes diversity calculation more accurate by implementing interval partitioning, extracting feature evidence, and calculating interval weights. It also introduces enhanced analysis submodules, including weighted abundance analysis and interval subsampling.

[0024] System outputs include "diversity curves" (taxonomic diversity, original vs. optimized comparisons, and extinction and origin rates), "data tables," and "performance evaluation results," providing comprehensive evaluation and visualization.

[0025] The system implementation process and method are the same and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a system framework diagram of an embodiment of the present invention;

[0027] Figure 2This is a performance diagram of optimized data under subsampling in a specific embodiment;

[0028] Figure 3 This is the optimization effect diagram under different diversity indicators;

[0029] Figure 4 This is a statistical difference analysis diagram before and after optimization;

[0030] Figure 5 This is a comparison chart of extinction rates before and after optimization. DETAILED DESCRIPTION

[0031] The present invention is further illustrated below with reference to specific examples. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0032] Change point detection automatically identifies key change points in data distribution by dynamically analyzing the variance changes of time series. When the variance change exceeds an adaptive threshold, it is automatically marked as an interval boundary.

[0033] Hierarchical clustering provides accurate identification of critical moments such as extinction events. Starting from the perspective of data similarity, hierarchical clustering automatically constructs a similarity matrix and forms optimal clusters through Euclidean distance and Ward method. It can automatically adjust the clustering level according to data complexity and group similar time points in the multidimensional feature space into the same interval.

[0034] The hybrid approach first tries multiple interval partitioning methods (change-point detection, hierarchical clustering, and adaptive windowing) in parallel and then quantitatively evaluates the results of each method using the AIC information criterion. The AIC value for each method is calculated using a correlation function based on the residual sum of squares, the number of intervals, and the sample size. The method with the best AIC value is automatically selected as the final solution, achieving data-driven selection.

[0035] Adaptive dynamic interval partitioning automatically selects the partitioning algorithm that best suits the current data characteristics. For each dataset, the three algorithms are first executed in parallel to obtain different interval partitioning results. The results of each method are then quantitatively evaluated using the AIC information criterion (AIC = 2k + n·ln(RSS / n)), where k is the number of intervals (complexity), n is the number of samples, and RSS is the residual sum of squares (goodness of fit). The system automatically selects the method with the best AIC value as the final interval partitioning scheme, achieving data-driven selection. This adaptive selection mechanism improves the scientific nature of interval partitioning and provides a more precise analytical framework for paleobiodiversity research.

[0036] (1) Interval division based on statistics (change point detection)

[0037] The adaptive interval partitioning algorithm based on the statistical characteristics of data automatically identifies the key change points of data distribution by analyzing the variance changes in the paleontological data time series:

[0038]

[0039] Where V(t) represents the variance in time window t, x i is the sample value, is the sample mean. When the variance ΔV(t) exceeds the preset threshold λ, the system will automatically mark it as the boundary of the interval:

[0040] Boundary={t|ΔV(t)>λ}.

[0041] In this embodiment, a sliding window technique is used to calculate the local variance change. The window size is dynamically adjusted according to the data volume to ensure that the data distribution characteristics can be effectively captured at different scales. At the same time, in order to improve the robustness of the method, a variety of statistical indicators are combined. Not only the variance is considered, but also the trend strength, autocorrelation and interquartile range are calculated. The comprehensive score of the above statistics is used to determine the optimal split point.

[0042] (2) Hierarchical clustering interval division

[0043] The dynamic interval partitioning method based on hierarchical clustering constructs a similarity matrix by calculating a weighted combination of multiple distance metrics (Euclidean distance, Manhattan distance, and correlation distance) between samples (each time point in the time series, each sample is represented by the eigenvector corresponding to that time point. This eigenvector contains a series of biogeological characteristics such as diversity indicators, extinction rate, and origin rate):

[0044] combined_distance=w1·euclidean_dist+w2·manhattan_dist+w3·correlation_dist,

[0045] Among them, combined_distance: the distance between samples after weighted combination; w1, w2, w3: the weight coefficients of each distance metric; euclidean_dist: the Euclidean distance between samples; manhattan_dist: the Manhattan distance between samples; correlation_dist: the correlation distance between samples

[0046] Then we use Ward's method to perform hierarchical clustering to minimize the variance within clusters:

[0047]

[0048] where x iRepresents the feature vector corresponding to the i-th time point in the time series, m A 、m B and m A∪B Represent the centers of cluster A, cluster B, and the merged cluster respectively.

[0049] The hierarchical clustering function of the correlation library is used to maintain the time continuity constraint, that is, the results of all result clustering must be continuous in time.

[0050] Compared with traditional analysis methods that preset fixed time windows and are difficult to adapt to non-uniform fossil records, the dynamic interval partitioning method based on hierarchical clustering can adaptively identify natural boundary intervals in the data without being restricted by uniform time units.

[0051] (3) Hybrid method interval division

[0052] To combine the advantages of both change point detection and hierarchical clustering, a hybrid interval partitioning strategy is used. This method uses the Akaike Information Criterion (AIC) to evaluate the results of multiple partitioning methods and automatically selects the method that best suits the current data characteristics:

[0053] AIC=2k+n*ln(RSS / n)

[0054] Where k is the number of intervals, n is the number of samples, and RSS is the residual sum of squares, which represents the sum of squared deviations of the samples from the mean within the interval.

[0055] The hybrid method not only automatically evaluates and selects the basic method with the smallest AIC value, but also further optimizes the results of the selected method. When the number of intervals generated by the selected method exceeds the target range, interval merging or splitting operations are automatically performed to ensure that the final result is within a reasonable range of interval numbers.

[0056] In practical applications, hybrid methods demonstrate greater adaptability. The system automatically selects the method best suited to the specific biodiversity data characteristics based on the AIC evaluation results. It adaptively selects the most appropriate interval partitioning method, leveraging the respective strengths of change point detection in identifying key change events and hierarchical clustering in discovering natural data groupings.

[0057] The data-driven interval weight assignment method quantifies interval importance through two complementary methods: feature importance-based weight calculation and Bayesian optimization weight assignment. This enables the system to adaptively highlight patterns of diversity change during critical periods. The feature importance-based weight calculation method constructs a feature matrix by combining multiple diversity indicators and utilizes an improved gradient boosting model (based on the traditional gradient boosting tree method, but specifically optimized for the characteristics of paleontological data, such as optimization for sparsity and missing values, and robustness optimization for outliers) to assess the impact of each indicator on the importance of the interval. In particular, this method is specifically optimized for the sparsity and uncertainty of paleontological data, achieving robustness in handling missing values ​​and outliers, making the weight calculation more adaptable. The Bayesian weight optimization method specifically designs a parameter space suitable for the characteristics of paleontological data (the range and distribution of all parameters that the model needs to search during the Bayesian optimization process). It can dynamically adjust weights based on interval length and position, addressing the problems of uneven sampling and preservation bias that traditional methods cannot address. This method implements an innovative weight mapping mechanism and achieves intelligent integration of multi-scale weights. This mechanism not only maps the raw weights generated by the model to specific time periods, but also integrates ecological weights for comprehensive evaluation. Experimental verification has shown that this weighting system can automatically identify and highlight key evolutionary events, such as assigning higher weights to periods before and after extinction events, more accurately reflecting true patterns of diversity change.

[0058] (1) Weight calculation based on feature importance

[0059] The contribution of each feature (referring to the various biogeological indicators corresponding to each time point (sample): diversity index, extinction rate index, origin rate index, etc.) to diversity change is automatically calculated through the random forest / gradient boosting model (Random Forest Model: a nonlinear model that integrates multiple decision trees, Gradient Boosting Model: mainly refers to gradient boosting decision trees, such as XGBoost) and converted into interval weights:

[0060]

[0061] where w i represents the weight of the i-th interval, I i Indicates the importance score of the corresponding feature. First, the feature importance is calculated using the XGBoost model (with an error fallback RandomForest mechanism added), and then the difference in boost weights is further processed by normalization:

[0062]

[0063] in It is the average value of the weight (referring to the importance score of each feature (such as diversity index, extinction rate, etc.) to the interval diversity change or analysis target), and the final weight range is controlled in the range of [0.7, 1.3].

[0064] In the actual implementation, the interval weights are learned from the data through the gradient boosting decision tree model. First, the feature matrix is ​​constructed to represent each interval as a binary feature vector:

[0065]

[0066] Where n is the length of the time series and k is the number of intervals. For time point i in interval j (the entire time series is divided into several consecutive time periods by dynamic interval partitioning method, each time period is an "interval"), if i belongs to interval j, then X interval [i, j] = 1, otherwise 0. This method enables the model (machine learning models for interval feature learning and weight assignment, such as random forests and gradient boosting trees) to learn the importance features of different intervals.

[0067] The model (a machine learning algorithm used for feature importance assessment and interval weight calculation, specifically XGBoost) is trained using the XGBoost algorithm, and its objective function is:

[0068]

[0069] Where l is the loss function and Ω is the regularization term. By minimizing this objective function, the model learns an integration of a series of decision trees, each of which corresponds to a base learner h t From the trained model, we extract feature importance as the basis for interval weights (a numerical coefficient automatically calculated by a machine learning model (such as XGBoost or RandomForest) to measure the impact of each interval (i.e., each time period obtained after dynamic interval division) on overall diversity changes or analysis targets. Interval weights are used in subsequent diversity dynamics analysis, interval weighted statistics, optimization, etc. The higher the weight, the greater the contribution or impact of the interval on overall diversity changes):

[0070]

[0071] Among them, gain k (j) represents the gain of feature j in decision tree k, T represents the total number of decision trees in the XGBoost or RandomForest model, and K represents the Kth tree.

[0072] In the weight application phase, a weight mapping mechanism is implemented to map the original weights generated by the model (XGBoost or RandomForest model used for feature importance assessment and weight generation) to specific time intervals, and to perform a comprehensive evaluation in combination with the ecological weights. First, the original feature weights are obtained through the XGBoost model, and then the feature weights generated by the model are divided according to the time intervals and assigned to each specific time interval. If the number of feature weights is consistent with the number of intervals, they are directly assigned:

[0073] Wi imterval [i]=W feature [i],

[0074] W interval [i]: represents the weight of the i-th time interval, W feature [i]: represents the weight (importance score) of the i-th feature.

[0075] If the number of feature weights is inconsistent with the number of intervals, the feature weights are normalized first, and then aggregated by interval length to obtain the corresponding weights:

[0076]

[0077] W interval [i]: The weight of the i-th time interval, used to map feature weights to interval weights, feat\start i : The feature start index corresponding to the i-th interval, feat\endi: The feature end index corresponding to the i-th interval. norm is the normalized feature weight, feat\start i and feat\endi is the feature index range corresponding to interval i. Finally, by creating an interval weight dictionary, the interval weights are further mapped to geological bins, forming a bin-to-weight mapping dictionary. The fossil records in each interval are assigned corresponding weights according to the bin to which they belong:

[0078] w s =w i ifs∈intervali.

[0079] w s : The weight of the sth fossil record, used to assign interval weights to each fossil record in the interval; w i : The weight of the i-th interval, if s∈intervali: the s-th fossil record belongs to the i-th interval.

[0080] For the calculation of diversity indicators, the weight application is reflected in the weighted processing of the original diversity indicators: sample_weighti =w i (sample_weight i : The weight of the i-th sample (fossil record), w i : Weight of the interval to which the i-th sample (fossil record) belongs. Function: Each sample is weighted by the interval to which it belongs, where sample i belongs to interval i. This weighting strategy allows fossil records from key periods to have greater influence in diversity analyses. To calculate the diversity index, weighting is applied as follows: divRTweighted = divRT·weight, where divRT is the original diversity index and weight is the corresponding weight.

[0081] A weighting method based on feature importance has demonstrated advantages in coral paleobiodiversity analysis. It automatically identifies key periods based on data distribution and enhances the signal expressed during critical transitions. Compared to the uniform weighting method used in traditional analysis, this method adaptively adjusts the weights of geological intervals based on the characteristics of the data, assigning higher weights to key geological features and avoiding the limitations of traditional methods that artificially fix weights. In practice, this method automatically identifies the differences in importance between different geological periods, allowing weights to be dynamically adjusted based on the characteristics of the data to determine the contribution of each interval to the final result.

[0082] (2) Bayesian optimization weight allocation

[0083] Find the optimal weight combination by iteratively optimizing the objective function:

[0084]

[0085] where w * is the optimal weight combination, f(w) is the model performance evaluation function, Ω is the weight parameter space, and w is the candidate weight combination (vector). Gaussian process regression is used to construct a probability model of the performance evaluation function:

[0086]

[0087] f(w): objective function (model performance evaluation function), which is modeled as a Gaussian process in Bayesian optimization. The Gaussian process is used to probabilistically model the objective function. μ(w) is the mean function of the Gaussian process, representing the expected value of the objective function under the weight combination w. k(w, w′) is the kernel function (covariance function) of the Gaussian process, measuring the correlation between different weight combinations w and w′. Modeling the objective function with a Gaussian process facilitates Bayesian optimization in efficiently searching for optimal weights under uncertainty.

[0088] The interval feature matrix is ​​constructed, mask features are created for each interval, features are standardized using StandardScaler, and the Gaussian regression process is configured by multiplying the RBF kernel with the constant kernel:

[0089]

[0090] k(x i ,x j ): RBF kernel function, which measures the similarity between two input feature vectors (such as two weight combinations or interval features), (x i ,x j ): input feature vector (such as feature encoding or weight combination of different intervals), C: amplitude parameter (constant) of kernel function, l: length scale parameter of kernel function, which controls the similarity decay speed, ||x i -x j || 2 : The square of the Euclidean distance, which measures the difference between two eigenvectors. The RBF kernel is used in Gaussian process regression to determine the correlation between different weight combinations and influence the sampling strategy of Bayesian optimization.

[0091] And use standard deviation as the basis for weighting.

[0092] The Bayesian optimization method supports a variety of parameter bounds that are optimized by the system using an objective function that calculates the signal-to-noise ratio of a weighted target value:

[0093]

[0094] SNR: Signal-to-Noise Ratio, used to measure the ratio of the model output signal to the noise, is a choice of objective function; Var(s): Variance of the signal, usually refers to the variance of the target variable such as interval-weighted diversity; The variance of the noise, y^ is the model prediction value, s is the real signal; ∈: a small constant to prevent the denominator from being zero. SNR is used as the objective function to measure the clarity of the model signal after weight distribution. The higher the SNR, the better the weight distribution effect.

[0095] Building on the dynamic interval partitioning mechanism and data-driven weight assignment strategy implemented above, the enhanced diversity analysis module integrates several innovative features. This module automatically detects key change points in time series, uses extinction rate sequences as features to guide interval partitioning, and simultaneously utilizes a multidimensional feature matrix (including multiple metrics such as diversity rate and extinction rate) for weight calculation, creating a weighted version for each diversity metric and implementing a dual mapping from intervals to weights. Furthermore, the survivor analysis function is enhanced to use survivor proportions to guide interval partitioning, and a generalized optimization analysis wrapper is applied to the optimization analysis function.

[0096] In interval partitioning applications, the module intelligently identifies key change points in extinction rates or diversity curves and automatically selects the interval partitioning algorithm that best suits the current data characteristics by finding the optimal interval function, eliminating the need for a pre-set fixed time window. When data quality issues are detected, exception handling mechanisms are automatically applied to ensure result stability.

[0097] In terms of weight distribution, the module integrates multidimensional feature matrix analysis functions, and simultaneously monitors multiple indicators such as true diversity, sampling bias-corrected diversity, single-interval biased diversity, and extinction percentage. By calculating the importance of each interval through the tree weight calculation function, the weight is controlled within the range of [0.7, 1.3] to ensure reasonable differentiation while avoiding the influence of extreme values.

[0098] In practical applications, this set of enhancements improves the ability to identify key events in the fossil record, providing more precise and reliable tools for paleobiodiversity analysis while maintaining compatibility and comparability with traditional methods.

[0099] The specific analysis results are as follows: Figure 2 This is a graph showing the performance of optimized data under subsampling. As can be seen from the figure, the optimized data differs somewhat from the original data. The optimized data exhibits higher diversity values ​​and smoother trends, improving the robustness of the data under sparse sampling conditions and providing a more reliable basis for interpreting the long-term evolution of coral diversity.

[0100] Figure 3 This figure shows the optimization results for different diversity metrics, comparing three key diversity metrics: transtemporal diversity (RT), boundary-crossing diversity (BC), and within-sample bias diversity (SIB). The results show that the weighted RT (solid red line) is slightly higher than the original RT (dashed red line) throughout the entire time period. In particular, the weighted SIB exhibits more pronounced fluctuations than the original SIB between approximately 170 and 150 million years ago, suggesting the presence of potentially significant ecological events during this period.

[0101] Figure 4 It is a statistical difference analysis diagram before and after optimization, showing the difference in diversity indicators before and after optimization under the arithmetic mean and geometric mean calculation methods. Overall, there are obvious systematic differences between the optimized data (solid line) and the original data (dashed line), and both averaging methods show statistically significant improvements. From the time series point of view, the optimization effect varies in different stages of geological history. In the middle Mesozoic, the data showed an upward trend after optimization, while in the middle Cretaceous, there was a certain degree of downward adjustment, which shows that the optimization process can intelligently adjust the diversity estimate according to the quality of the original data. The small graph embedded in the figure focuses on the detailed changes in a specific geological period, which strongly proves the stability of the optimization algorithm when dealing with high-volume periods.

[0102] Figure 5 This graph compares coral extinction rates before and after optimization. Overall, the optimized extinction rate curve follows the original curve, but there are differences at key geological periods (e.g., approximately 200-240 million years ago). This demonstrates that the optimization method improves sensitivity to key evolutionary milestones while maintaining the rationality of overall diversity dynamics.

[0103] According to the results in Table 1, the hierarchical clustering method performed best within the target interval range, generating 30 intervals with an AIC value of -156.25. Although the AIC value of the change point detection method (-145.82) is also low, its number of intervals, 6, is not within the target range.

[0104] Table 1 AIC evaluation comparison table of interval partitioning methods

[0105]

[0106] According to the results in Table 2, the AUC value reached 0.9487, indicating that the model has high discrimination ability; the extremely low diversity PSI (0.0788) and extinction rate PSI (0.0078) indicate that the optimization process fully retains the distribution characteristics of the original data; at the same time, the average Jaccard similarity of 0.9343 and the minimum similarity of 0.8421 demonstrate the high stability and consistency of the interval partitioning method.

[0107] Table 2 Performance evaluation index table

[0108]

Claims

1. A paleodiversity analysis method based on a dynamic machine learning joint calibration model, characterized in that: The steps include: Step 1: Obtain paleontological data samples and perform preprocessing on the paleontological data samples by missing value processing, outlier detection, data standardization, and time data mapping; Step 2: Dynamically partition the preprocessed paleontological data samples using a variety of algorithms to achieve adaptive time interval partitioning, including: change point detection, hierarchical clustering, and hybrid methods; Step 3: Implementing non-uniform weight distribution through data-driven weights to obtain an optimized data set; the data-driven weighting method includes: feature importance weighting and Bayesian optimization weighting method; Step 4: Enhanced diversity analysis is performed on the optimized dataset, including interval partitioning, feature evidence extraction, and interval weight calculation. Diversity calculation is further integrated with interval weighted diversity dynamic calculation, enhanced survivor analysis, and interval subsampling analysis to make the diversity calculation more accurate. The output includes diversity curves, data tables, and performance evaluation results.

2. The paleodiversity analysis method based on the dynamic machine learning joint calibration model according to claim 1, characterized in that: In terms of interval division, the proposed various interval division mechanisms, including statistical-based interval division, hierarchical clustering and hybrid methods, can adaptively determine the optimal time window according to the characteristics of the fossil data itself and identify the key nodes of biodiversity changes; in terms of weight distribution, data-driven interval weight calculation is achieved through feature importance evaluation and Bayesian optimization.

3. The paleodiversity analysis method based on the dynamic machine learning joint calibration model according to claim 1, characterized in that: The change point detection automatically identifies key change points in data distribution by dynamically analyzing the variance changes of time series, and automatically marks them as interval boundaries when the variance changes exceed an adaptive threshold.

4. The paleodiversity analysis method based on the dynamic machine learning joint calibration model according to claim 1, characterized in that: The hierarchical clustering provides accurate identification of critical moments such as extinction events. Starting from the perspective of data similarity, hierarchical clustering automatically constructs a similarity matrix and forms optimal clusters through Euclidean distance and Ward method. It can automatically adjust the clustering level according to the complexity of the data and group similar time points in the multidimensional feature space into the same interval.

5. The paleodiversity analysis method based on dynamic machine learning joint calibration model according to claim 1, characterized in that: The hybrid method first tries multiple interval partitioning methods in parallel, including change point detection, hierarchical clustering, and adaptive windowing, and then quantitatively evaluates the results of each method using the AIC information criterion. The AIC value of each method is calculated based on the residual sum of squares, the number of intervals, and the sample size through a correlation function. The method with the best AIC value is automatically selected as the final solution, achieving data-driven selection.

6. The paleodiversity analysis method based on the dynamic machine learning joint calibration model according to claim 1, characterized in that: The adaptive dynamic interval partitioning method can automatically select the partitioning algorithm that best suits the current data characteristics. For each data set, the three algorithms of change point detection, hierarchical clustering, and adaptive windowing are first executed in parallel to obtain different interval partitioning results. The results of each method are then quantitatively evaluated using the AIC information criterion, where AIC = 2k + n•ln(RSS / n), where k is the number of intervals, n is the number of samples, and RSS is the residual sum of squares. The method with the best AIC value is selected as the final interval partitioning scheme to achieve data-driven selection.

7. The paleodiversity analysis method based on dynamic machine learning joint calibration model according to claim 1, characterized in that: Feature importance weighting methods include XGBoost algorithm, feature importance and adaptive weighting; Bayesian optimization weighting methods include Gaussian process regression, hierarchical Bayesian analysis and handling uncertainty.

8. A paleodiversity analysis system based on a dynamic machine learning joint calibration model, characterized in that: include: Data preprocessing module, which obtains paleontological data samples and performs missing value processing, outlier detection, data standardization and time data mapping; Dynamic interval partitioning module, which uses multiple algorithms to achieve adaptive time interval partitioning, including: change point detection, hierarchical clustering, and hybrid methods; Data-driven weight module to achieve non-uniform weight distribution, including feature importance weight and Bayesian optimization weight; After the above processing, the system obtains an optimized data set, which has realized interval optimization processing, weighted data generation and interval weight mapping; The enhanced diversity analysis module makes diversity calculation more accurate by implementing interval partitioning, extracting feature evidence, and calculating interval weights. It also introduces enhanced analysis function submodules, including weighted abundance analysis and interval subsampling. System outputs include diversity curves, data tables, and performance evaluation results, providing comprehensive evaluation and visualization; diversity curves include taxonomic group diversity curves, original vs. optimized comparison curves, and extinction rate vs. origin rate curves.

Citation Information

Patent Citations

  • Commercial network site selection method, system and device based on open source data mining and medium

    CN112990976A

  • NDVI prediction integrated optimization method and system based on machine learning

    CN117391221A

Cited By

  • Space-time coupling virtual metering intelligent labeling method based on feature space geometry

    CN121542791A