Paleoclimate Change Analysis Methods and Systems Based on Multi-Source Data Fusion
By constructing a Bayesian network structure and fusing multi-source data, the posterior probability of volcanic events and sulfur cycle anomalies is dynamically calculated, which solves the problem of insufficient identification of external driving factors in paleoclimate analysis in existing technologies and realizes accurate analysis of paleoclimate change.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HEBEI GEO UNIVERSITY
- Filing Date
- 2026-01-22
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies have failed to effectively identify external driving factors in paleoclimate analysis, leading to data uncertainty and reduced analytical accuracy.
By collecting multi-source index data from geological sedimentary rocks, a Bayesian network structure was constructed, causal relationships were defined, conditional probability distributions were set, and the posterior probability of volcanic events and sulfur cycle anomalies was dynamically calculated. Time matching and correlation analysis were then performed in conjunction with paleoclimate event labels.
Accurately identifying volcano-driven sulfur cycle disturbances reveals the mechanisms by which volcanic activity influences paleoclimate change, improving the accuracy and reliability of paleoclimate analysis.
Smart Images

Figure CN121561649B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of paleoclimate change analysis technology, specifically to a paleoclimate change analysis method and system based on multi-source data fusion. Background Technology
[0002] The sulfur cycle is closely related to volcanic activity, and the enrichment of mercury (Hg) in sediments and changes in Hg isotopic composition have been found to serve as effective indicators for accurately indicating large-scale volcanic eruptions. Sulfides released from volcanic eruptions enter the atmosphere and form sulfuric acid aerosols, which are highly reflective and can significantly increase the albedo of the Earth's atmosphere, reducing solar radiation reaching the Earth's surface and leading to a series of climate events. By analyzing the time-series evolution characteristics of Hg / Hg isotopes in sedimentary rocks to analyze the sulfur cycle perturbations driven by volcanic events, the impact of volcanic events on paleoclimate events can be indirectly analyzed.
[0003] The existing technology, disclosed in CN117195545A, provides a method for calculating the formation time window and distribution range of source rocks and gypsum-salt caps. This method includes: conducting chronological analysis to determine the fine chronological sequence of stratigraphy; conducting paleomagnetic analysis to determine the paleogeographic evolution history; performing stratigraphic sedimentary facies analysis and paleoclimatic environmental analysis to establish macroscopic and fine paleoclimatic environmental histories; inputting the fine chronological sequence of stratigraphy, paleogeographic evolution history, and macroscopic and fine paleoclimatic environmental histories into a coupled continental-atmospheric circulation-climate-ocean current model for coupled analysis to obtain the coupling relationship of continental-atmospheric circulation-climate-ocean currents under the selected fine chronological constraints, thereby determining the formation time window and distribution range of source rocks and gypsum-salt caps. While this method achieves paleoclimatic analysis through data fusion, it does not consider whether the collected data is from natural evolution or driven by external factors. Therefore, the collected data has a certain degree of uncertainty, resulting in insufficient overall system recognition of externally driven data and reduced accuracy in paleoclimatic analysis.
[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for paleoclimate change analysis based on multi-source data fusion, so as to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] The paleoclimate change analysis method based on multi-source data fusion includes the following specific steps:
[0008] S1: Collect multi-source index data on volcanoes and sulfur cycles from geological sedimentary rocks, standardize the multi-source index data, align the geological time based on geological age, and divide the geological age into several time windows according to the time axis.
[0009] S2: Construct a Bayesian network structure containing volcanic events and observation indicators based on multi-source index data, define the causal relationship between nodes in the Bayesian network structure, and set the conditional probability distribution;
[0010] S3: Generate the probability of volcanic events and sulfur cycle anomalies in different time windows based on the Bayesian network structure, and determine the volcanic events and sulfur cycle anomalies by combining the preset probability thresholds.
[0011] S4: Based on the anomaly detection results, identify volcano-driven sulfur cycle disturbance events and their corresponding time windows. Match the time windows with preset paleoclimate event labels, and perform correlation analysis on the matched volcano-driven sulfur cycle disturbance events and paleoclimate events to analyze the impact of volcanic eruptions on paleoclimate.
[0012] Preferably, the multi-source index data includes, but is not limited to: Hg element content, Hg isotope ratio, volcanic ash layer thickness, S isotope content, organic carbon content, and Fe / S ratio;
[0013] When performing time alignment on multi-source indicator data, the standardized multi-source indicator data are mapped to the same time axis according to geological time. The uniformity of the multi-source indicator data on the time axis is checked. If there are unevenly distributed multi-source indicator data, preliminary data supplementation is performed by interpolation methods, including but not limited to linear interpolation and time window average interpolation.
[0014] Preferably, step S2 includes:
[0015] S201: Define the volcanic event as a binary label and use it as a hidden variable node in the Bayesian network structure. The various multi-source index data are the corresponding observation indicators and are used as observable nodes in the Bayesian network structure respectively.
[0016] S202: Based on geochemical and geological knowledge, a causal relationship chain is established for each node in the Bayesian network structure: volcanic events affect Hg element content, Hg isotope ratio, volcanic ash layer thickness, and S isotope content. S isotope content further affects organic carbon content and Fe / S ratio.
[0017] S203: Check the data integrity of each node. If there is missing data, use the expectation-maximization method to perform secondary data completion.
[0018] S204: Set the probability distribution based on the causal relationship of each node, where the latent variable nodes are the prior probabilities, and each observable node follows a conditional Gaussian distribution.
[0019] Preferably, when the binary label of the volcanic event is 0, it indicates that no volcanic event occurred within the time window; when the binary label is 1, it indicates that a volcanic event occurred within the time window.
[0020] For each observable node, the conditional Gaussian distribution is satisfied by the parent node, and the distribution parameters include the mean and standard deviation.
[0021] The parent node is determined by the causal relationship between the nodes: the volcanic event is the parent node for Hg element content, Hg isotope ratio, volcanic ash layer thickness, and S isotope content; the S isotope content is the parent node for organic carbon content and Fe / S ratio.
[0022] Preferably, the causal relationships among nodes in the Bayesian network structure are represented by a joint probability distribution, the expression of which is:
[0023]
[0024] In the formula This represents the joint probability distribution of the Bayesian network structure. Indicates the first Binary labels for volcanic events within a time window. This represents the prior probability of a volcanic event. Indicates the first A collection of multi-source indicator data within a time window, indexed Indicates the index of the time window. Represents a node The probability distribution, Represents a node The parent node.
[0025] Preferably, in step S3, the probability of a volcanic event is obtained by calculating the posterior probability using a Bayesian network structure combined with Bayes' theorem. The calculation formula is as follows:
[0026]
[0027] In the formula This represents the posterior probability of a volcanic event. Indicates that the parent node is At the same time, multi-source indicator data The joint probability distribution, The marginal probabilities used for normalization are calculated by summing using the method of total probability, and their formula is:
[0028]
[0029] In the formula , These represent multi-source indicator data when a volcanic event has occurred and when no volcanic event has occurred, respectively. The joint probability distribution, , Let represent the prior probabilities of a volcanic event occurring and the probability of no volcanic event occurring, respectively.
[0030] When determining a volcanic event, the posterior probability of the event occurring is compared with a preset probability threshold. If the following conditions are met:
[0031]
[0032] Then it is believed that in the first A volcanic event occurred within a time window, where... This represents the posterior probability of a volcanic event occurring. This represents the preset probability threshold.
[0033] Preferably, in step S3, the probability of sulfur cycle anomalies is obtained based on the conditional Gaussian distribution of the S isotope content at the corresponding nodes, specifically as follows:
[0034] Based on the conditional Gaussian distribution of the S isotope content at the corresponding nodes, the corresponding mean and standard deviation are obtained.
[0035] The normal range of the sulfur cycle is defined based on the mean and standard deviation of S isotope content. When the S isotope content is within the normal range, the sulfur cycle is considered normal, and when the S isotope content is outside the normal range, the sulfur cycle is considered abnormal.
[0036] Define the upper and lower limits of the normal range as the mean and... Sum of standard deviations, mean and The difference of two standard deviations This represents a multiplier factor ranging from 2 to 4.
[0037] The probability of a normal sulfur cycle is calculated based on the cumulative distribution function of the standard Gaussian distribution, and then the probability of an abnormal sulfur cycle is calculated based on the probability of a normal sulfur cycle.
[0038] The posterior probability of volcanic events is used to weight the probability of sulfur cycle anomalies to generate the posterior probability of sulfur cycle anomalies.
[0039] When determining sulfur cycle anomalies, the posterior probability of the sulfur cycle anomaly is compared with a preset probability threshold. If the posterior probability of the sulfur cycle anomaly is not lower than the preset probability threshold, then it is considered to be an anomaly in the [missing information - likely a specific timeframe]. An anomaly in the sulfur cycle occurred within a specific time window.
[0040] Preferably, if both a volcanic event and a sulfur cycle anomaly occur within a certain time window, then a volcano-driven sulfur cycle disturbance event is considered to exist within that time window.
[0041] A paleoclimate change analysis system based on multi-source data fusion, wherein the analysis system is used to perform the above-mentioned analysis methods, specifically including:
[0042] The data acquisition module is used to collect multi-source index data on volcanoes and sulfur cycles in geological sedimentary rocks, standardize the multi-source index data, align the data based on geological time, and divide the geological time into several time windows according to the time axis.
[0043] The algorithm construction module is used to construct a Bayesian network structure containing volcanic events and observation indicators, define the causal relationships between nodes in the Bayesian network structure, and set the conditional probability distribution.
[0044] The comprehensive judgment module is used to generate the probability of volcanic events and sulfur cycle anomalies in different time windows, and to judge volcanic events and sulfur cycle anomalies in combination with preset probability thresholds.
[0045] The data analysis module is used to identify volcano-driven sulfur cycle disturbance events and their corresponding time windows, and to match the time windows with preset paleoclimate event labels in order to analyze the impact of volcanic eruptions on paleoclimate.
[0046] Compared with the prior art, the beneficial effects of the present invention are:
[0047] This invention collects and standardizes multi-source geochemical indicators related to volcanoes and the sulfur cycle, and combines geological methods to achieve time alignment, ensuring data consistency and continuity. Based on a Bayesian network, a multivariate causal structure is constructed, with volcanic events as latent variables and conditional Gaussian distributions used to model observed indicators, uncovering causal relationships and statistical characteristics among multiple indicators and resolving data uncertainty and missing data issues. By dynamically calculating the posterior probability of volcanic events and sulfur cycle anomalies, probability-based objective judgment is achieved, accurately identifying volcano-driven sulfur cycle disturbances. Furthermore, time matching and correlation analysis are performed using paleoclimate event labels, thereby revealing the mechanism by which volcanic activity influences paleoclimate change. Attached Figure Description
[0048] Figure 1 This is a schematic diagram of the overall method flow of the present invention;
[0049] Figure 2 This is a schematic diagram of step S2 of the present invention;
[0050] Figure 3This is a schematic diagram of the overall module structure of the analysis system of the present invention. Detailed Implementation
[0051] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0052] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0053] Example:
[0054] Please see Figures 1-2 The present invention provides a technical solution:
[0055] The paleoclimate change analysis method based on multi-source data fusion includes the following specific steps:
[0056] S1: Collect multi-source index data on volcanoes and sulfur cycles from geological sedimentary rocks, standardize the multi-source index data, align the data based on geological time, and divide the geological time into several time windows according to the time axis.
[0057] Multi-source indicator data include, but are not limited to: Hg element content, Hg isotope ratio, volcanic ash layer thickness, S isotope content, organic carbon content, and Fe / S ratio.
[0058] When collecting geological sedimentary rocks, clean sampling tools and containers must be used to avoid sample contamination. For the various parameters in the multi-source index data, Hg elemental content, Hg isotope ratio, and S isotope content can be detected by mass spectrometry. The thickness of the volcanic ash layer can be directly obtained from the geological sedimentary rock profile, and its thickness and grain size distribution can be measured under a microscope. Organic carbon content can be detected by elemental analyzer. The Fe content in the Fe / S ratio can be detected by X-ray fluorescence spectroscopy, and the S content can be detected by elemental analyzer. The ratio can then be calculated based on the detection results.
[0059] These indicators are collected because the sulfur cycle is closely related to volcanic activity, and the enrichment of mercury (Hg) and changes in Hg isotopes in sediments (such as geological sedimentary rocks) have been found to be effective indicators for accurately indicating large-scale volcanic eruptions. Analyzing the time-series evolution of Hg / Hg isotopes in sedimentary rocks allows for the analysis of volcanic events, thus indirectly analyzing the sulfur cycle. Sulfides from volcanic eruptions entering the atmosphere form sulfuric acid aerosols, which are highly reflective and can significantly increase the Earth's atmospheric albedo, reducing solar radiation reaching the Earth's surface and leading to a series of climate events (such as short-term / long-term global / local cooling). Without identifying these indicators and determining whether they have external drivers (such as volcanic eruptions), the collected data will be uncertain, reducing the overall analytical accuracy of the system. Z-score standardization can be used to standardize these data to eliminate the influence of dimensions.
[0060] When performing time alignment on multi-source index data, the standardized multi-source index data are mapped to the same time axis according to geological time (such as U-Pb zircon dating, magnetostratigraphy, and C isotope stratigraphy). The uniformity of the multi-source index data on the time axis is checked. If there are multi-source index data with uneven distribution (i.e., data missing at some nodes of the time axis), preliminary data supplementation is performed using interpolation methods, including but not limited to linear interpolation and time window average interpolation.
[0061] By standardizing the time scale of multi-source indicator data, causal analysis can be performed on the same data at the same point in time, avoiding inference biases caused by time mismatch. This not only enhances the scientific ability to interpret the causal relationship between volcanic events and sulfur cycle perturbations and their impact on paleoclimate, but also enables subsequent identification of volcano-driven sulfur cycle perturbation events to more accurately correspond to paleoclimate event labels, providing a solid basis for research on paleoclimate change mechanisms.
[0062] S2: Construct a Bayesian network structure containing volcanic events and observation indicators based on multi-source indicator data, define the causal relationships between nodes in the Bayesian network structure, and set the conditional probability distribution.
[0063] Volcanoes and sulfur cycle-related indicators suffer from measurement errors, missing data, and uncertainties. Bayesian networks, which express causal relationships between variables in the form of directed acyclic graphs, can make robust inferences and complete the data. Therefore, they are suitable for describing the multi-level and multi-path influence mechanisms of volcanic events on geochemical indicators. They can transform geological and geochemical knowledge into prior structures and probability distributions, and combine actual observation data for parameter learning and posterior updates, systematically revealing the dependency structure between events and observed indicators. This provides a quantitative basis for the determination of volcanic events and sulfur cycle anomalies.
[0064] Step S2 includes:
[0065] S201: Define the volcanic event as a binary label and use it as a hidden variable node in the Bayesian network structure. The various multi-source index data are the corresponding observation indicators and are used as observable nodes in the Bayesian network structure respectively.
[0066] S202: Based on geochemical and geological knowledge, a causal relationship chain is established for each node in the Bayesian network structure: volcanic events affect Hg element content, Hg isotope ratio, volcanic ash layer thickness, and S isotope content. S isotope content further affects organic carbon content and Fe / S ratio.
[0067] S203: Check the data integrity of each node. If there is missing data, use the expectation-maximization method to perform secondary data completion.
[0068] S204: Set the probability distribution based on the causal relationship of each node, where the latent variable nodes are the prior probabilities, and each observable node follows a conditional Gaussian distribution.
[0069] Since historical volcanic events are geological events that cannot be directly observed, yet they are the main drivers of the entire system, they are set as latent variable nodes to reflect their "unknown state but significant impact" characteristics, facilitating the inference of their occurrence probability through observational indicators. Multi-source indicator data, on the other hand, are all directly measurable or inferable variables; they are used as observation nodes to represent the actual input of the observational data, making it easier for the model to extract information from the data. The causal chain is a reasonable construction based on geochemical geological knowledge: volcanic events directly affect Hg and its isotopes, volcanic ash layers, and sulfur isotopes, reflecting the material output and chemical signals of volcanic activity. Sulfur isotopes further affect organic carbon content and the Fe / S ratio, reflecting the indirect impact of sulfur cycle disturbances on organic geochemical processes and environmental conditions.
[0070] Regarding the probability distribution of each node, the prior probability of the latent variable nodes can be obtained through geological history statistics, such as statistically analyzing the frequency of volcanic eruptions in similar geological regions or study areas at different geological periods, and converting it into the probability of occurrence per unit time or time window. Alternatively, a uniform distribution (unbiased prior) can be used, or values can be assigned based on expert knowledge. Furthermore, since the elemental content under a single environment or process in geology and geochemistry approximates a standard Gaussian distribution and conforms to the central limit theorem, a conditional Gaussian distribution is adopted for observable nodes to express the phenomenon that "the distribution parameters of child nodes differ under different parent node states."
[0071] A binary label of 0 for a volcanic event indicates that no volcanic event occurred within that time window, while a binary label of 1 indicates that a volcanic event occurred within that time window. Binarizing volcanic events is done because they are essentially discrete events, and binary labels accurately reflect their characteristics. Furthermore, it simplifies model construction and probability inference, making it easier to directly determine whether an event has occurred.
[0072] For each observable node, the conditional Gaussian distribution is satisfied by the parent node, and the distribution parameters include the mean and standard deviation.
[0073] The parent node is determined by the causal relationship between the nodes: the volcanic event is the parent node of Hg element content, Hg isotope ratio, volcanic ash layer thickness, and S isotope content; the S isotope content is the parent node of organic carbon content and Fe / S ratio.
[0074] By automatically adjusting the conditional Gaussian distribution parameters with the parent node, the model can sensitively capture the changes in indicators caused by volcanic events, effectively distinguish between event and non-event states. Moreover, the clear parent-child relationship and conditional distribution allow multiple observation indicators to work together, ensuring that the Bayesian network node settings and probability distribution conform to the actual process, thus enhancing the comprehensive identification ability of complex volcanic-driven sulfur cycle disturbance events.
[0075] In a Bayesian network structure, the causal relationships between nodes are represented by a joint probability distribution, which is expressed as:
[0076]
[0077] In the formula This represents the joint probability distribution of the Bayesian network structure. Indicates the first Binary labels for volcanic events within a time window. This represents the prior probability of a volcanic event. Since a volcanic event may or may not occur, the prior probability also includes... , Two scenarios, Indicates the first A collection of multi-source indicator data within a time window, indexed This indicates the index of the time window. The time window can be set based on the uncertainty of geochronological determination (such as the U-Pb dating error, which is usually in the range of thousands to tens of thousands of years). The selected time window length should not be less than the maximum uncertainty to avoid temporal confusion of geological events within the same window. The specific length is determined based on expert experience. Represents a node The probability distribution (i.e., the conditional Gaussian distribution that depends on the parent node). Represents a node The parent node.
[0078] The first The multi-source indicator data within each time window are respectively labeled as: Hg element content. Hg isotope ratio Thickness of volcanic ash layer S isotope content Organic carbon content Fe / S ratio ,So This can be expressed as:
[0079]
[0080] Expressing the causal relationships among nodes in a Bayesian network through a joint probability distribution not only aligns with scientific understanding in geology and geochemistry but also significantly improves the accuracy and interpretability of multi-source data fusion and event determination.
[0081] S3: Generate the probability of volcanic events and sulfur cycle anomalies within different time windows based on a Bayesian network structure, and determine the volcanic events and sulfur cycle anomalies by combining them with a preset probability threshold.
[0082] In step S3, the probability of a volcanic event is calculated using a Bayesian network structure combined with Bayes' theorem, and the formula is as follows:
[0083]
[0084] In the formula This represents the posterior probability of a volcanic event, i.e., given multi-source indicator data. Under what circumstances, the probability of a volcanic event occurring / not occurring. Indicates that the parent node is At the same time, multi-source indicator data The joint probability distribution, The marginal probabilities used for normalization are calculated by summing using the method of total probability, and their formula is:
[0085]
[0086] In the formula , These represent multi-source indicator data when a volcanic event has occurred and when no volcanic event has occurred, respectively. The joint probability distribution, , Let represent the prior probabilities of a volcanic event occurring and the probability of no volcanic event occurring, respectively.
[0087] When determining a volcanic event, the posterior probability of the event occurring is compared with a preset probability threshold. If the following conditions are met:
[0088]
[0089] Then it is believed that in the first A volcanic event occurred within a time window, where... This represents the posterior probability of a volcanic event occurring. This represents the preset probability threshold, which is generally between 0.75 and 0.9.
[0090] The probability of sulfur cycle anomalies is obtained based on the conditional Gaussian distribution of S isotope content at corresponding nodes, with the specific logic as follows:
[0091] Based on the conditional Gaussian distribution of the S isotope content at the corresponding nodes, the corresponding mean and standard deviation are obtained.
[0092] The normal range of the sulfur cycle is defined based on the mean and standard deviation of S isotope content. When the S isotope content is within the normal range, the sulfur cycle is considered normal, and when the S isotope content is outside the normal range, the sulfur cycle is considered abnormal.
[0093] Define the upper and lower limits of the normal range as the mean and... Sum of standard deviations, mean and The difference of two standard deviations This represents a multiplier factor ranging from 2 to 4.
[0094] The probability of a normal sulfur cycle is calculated based on the cumulative distribution function of the standard Gaussian distribution, and then the probability of an abnormal sulfur cycle is calculated based on the probability of a normal sulfur cycle.
[0095] The posterior probability of volcanic events is used to weight the probability of sulfur cycle anomalies to generate the posterior probability of sulfur cycle anomalies.
[0096] When determining sulfur cycle anomalies, the posterior probability of the sulfur cycle anomaly is compared with a preset probability threshold. If the posterior probability of the sulfur cycle anomaly is not lower than the preset probability threshold, then it is considered to be an anomaly in the [missing information - likely a specific timeframe]. An anomaly in the sulfur cycle occurred within a specific time window.
[0097] Since the S isotope content is an observable node whose parent node is a volcanic event, the conditional Gaussian distribution of this observable node can be expressed as:
[0098]
[0099] in This represents the conditional Gaussian distribution of S isotope content under volcanic activity. The binary label value representing the volcanic event is either 0 or 1. The mean is Standard deviation is The conditional Gaussian distribution function, the specific function expansion form of which will not be elaborated here.
[0100] After obtaining the mean and standard deviation, the upper and lower limits of the normal range for the sulfur cycle can be defined. The normal range can be expressed as: The interval is marked as In other words, as long as the sulfur isotope content is within this range, the sulfur cycle is considered normal; otherwise, it is considered abnormal. Therefore, the cumulative distribution function of the standard normal distribution can be used to calculate the probability of a normal sulfur cycle, as shown in the following formula:
[0101]
[0102] In the formula This indicates the probability that the sulfur cycle is functioning normally. This represents the cumulative distribution function.
[0103] Once the probability of a normal sulfur cycle is obtained, the probability of an abnormal sulfur cycle can be obtained, that is:
[0104]
[0105] In the formula This indicates the probability of an anomaly in the sulfur cycle.
[0106] Since the probability of sulfur cycle anomaly obtained at this point is based on the volcanic event as the parent node, but the volcanic event is an unobservable latent variable node, it is necessary to combine it with the posterior probability of the volcanic event state for weighted calculation to obtain the posterior probability of sulfur cycle anomaly. This can also be calculated using the total probability method, i.e.:
[0107]
[0108] In the formula This represents the posterior probability of a sulfur cycle anomaly, indicating the confidence level of a sulfur cycle anomaly given all multi-source data. Compared to... In this regard, the posterior probability is more accurate and can be dynamically updated with the Bayesian network structure, thus more realistically reflecting the probability of sulfur cycle anomalies occurring in actual situations.
[0109] S4: Based on the anomaly detection results, identify volcano-driven sulfur cycle disturbance events and their corresponding time windows. Match the time windows with preset paleoclimate event labels, and perform correlation analysis on the matched volcano-driven sulfur cycle disturbance events and paleoclimate events to analyze the impact of volcanic eruptions on paleoclimate.
[0110] If a volcanic event and a sulfur cycle anomaly both occur within a certain time window, it is considered that a volcano-driven sulfur cycle disturbance event exists within that time window.
[0111] Specifically, paleoclimate event tagging can collect information on paleoclimate events (such as glacial periods, warm periods, dry periods, cold events, greenhouse events, etc.) regionally or globally through paleoclimate databases, literature reviews, or research results. This information is accompanied by time intervals and event characteristics, mapping the time periods of each paleoclimate event to a unified time axis, and standardizing event names, types, and durations to form a computable event tag library. The frequency with which the overlap between the time windows of paleoclimate event tags and the time windows corresponding to volcano-driven sulfur cycle disturbance events exceeds a threshold (e.g., 50%) is statistically analyzed. Then, statistical methods (such as cross-correlation and Granger causality tests) can be used to verify the correlation between different paleoclimate events and volcanic events. Finally, the influence relationship between volcanic events and this type of paleoclimate event can be quantified based on the correlation (e.g., three levels of influence: strong / medium / weak).
[0112] This step effectively connects key aspects of geochemical data processing and paleoclimate research, from Bayesian network anomaly detection results to specific disturbance event identification, and then to correlation analysis with paleoclimate events.
[0113] Please see Figure 3 This application also provides a paleoclimate change analysis system based on multi-source data fusion. The analysis system is used to perform the above-mentioned analysis methods, specifically including:
[0114] The data acquisition module is used to collect multi-source index data on volcanoes and sulfur cycles in geological sedimentary rocks, standardize the multi-source index data, align the data based on geological time, and divide the geological time into several time windows according to the time axis.
[0115] The algorithm construction module is used to construct a Bayesian network structure that includes volcanic events and observation indicators, define the causal relationships between nodes in the Bayesian network structure, and set the conditional probability distribution.
[0116] The comprehensive judgment module is used to generate the probability of volcanic events and sulfur cycle anomalies in different time windows, and to judge volcanic events and sulfur cycle anomalies in combination with preset probability thresholds.
[0117] The data analysis module is used to identify volcano-driven sulfur cycle disturbances and their corresponding time windows, and to match the time windows with preset paleoclimate event labels to analyze the impact of volcanic eruptions on paleoclimate.
[0118] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0119] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.
[0120] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0121] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A paleoclimate change analysis method based on multi-source data fusion, characterized in that, The specific steps include: S1: Collect multi-source index data on volcanoes and sulfur cycles from geological sedimentary rocks, standardize the multi-source index data, align the geological time based on geological age, and divide the geological age into several time windows according to the time axis. S2: Construct a Bayesian network structure containing volcanic events and observation indicators based on multi-source index data, define the causal relationship between nodes in the Bayesian network structure, and set the conditional probability distribution; S3: Generate the probability of volcanic events and sulfur cycle anomalies in different time windows based on the Bayesian network structure, and determine the volcanic events and sulfur cycle anomalies by combining the preset probability thresholds. S4: Based on the judgment results, identify volcano-driven sulfur cycle disturbance events and their corresponding time windows, match the time windows with preset paleoclimate event labels, and perform correlation analysis on the matched volcano-driven sulfur cycle disturbance events and paleoclimate events to analyze the impact of volcanic eruptions on paleoclimate. The multi-source index data includes: Hg element content, Hg isotope ratio, volcanic ash layer thickness, S isotope content, organic carbon content, and Fe / S ratio. Step S2 includes: S201: Define the volcanic event as a binary label and use it as a hidden variable node in the Bayesian network structure. The various multi-source index data are the corresponding observation indicators and are used as observable nodes in the Bayesian network structure respectively. S202: Based on geochemical and geological knowledge, a causal relationship chain is established for each node in the Bayesian network structure: volcanic events affect Hg element content, Hg isotope ratio, volcanic ash layer thickness, and S isotope content. S isotope content further affects organic carbon content and Fe / S ratio. S203: Check the data integrity of each node. If there is missing data, use the expectation-maximization method to perform secondary data completion. S204: The probability distribution is set based on the causal relationship of each node, where the latent variable nodes are the prior probabilities, and each observable node follows a conditional Gaussian distribution. When the binary label of the volcanic event is 0, it means that no volcanic event occurred within the time window; when the binary label is 1, it means that a volcanic event occurred within the time window. For each observable node, the conditional Gaussian distribution is satisfied by the parent node, and the distribution parameters include the mean and standard deviation. The parent node is determined by the causal relationship between the nodes: the volcanic event is the parent node for Hg element content, Hg isotope ratio, volcanic ash layer thickness, and S isotope content; the S isotope content is the parent node for organic carbon content and Fe / S ratio.
2. The paleoclimate change analysis method based on multi-source data fusion according to claim 1, characterized in that: When performing time alignment on multi-source indicator data, the standardized multi-source indicator data are mapped to the same time axis according to geological time. The uniformity of the multi-source indicator data on the time axis is checked. If there are unevenly distributed multi-source indicator data, preliminary data supplementation is performed by interpolation methods, including but not limited to linear interpolation and time window average interpolation.
3. The paleoclimate change analysis method based on multi-source data fusion according to claim 2, characterized in that: The causal relationships among the nodes in the Bayesian network structure are represented by a joint probability distribution, the expression of which is: In the formula This represents the joint probability distribution of the Bayesian network structure. Indicates the first Binary labels for volcanic events within a time window. This represents the prior probability of a volcanic event. Indicates the first A collection of multi-source indicator data within a time window, indexed Indicates the index of the time window. Represents a node The probability distribution, Represents a node The parent node.
4. The paleoclimate change analysis method based on multi-source data fusion according to claim 2, characterized in that: In step S3, the probability of a volcanic event is calculated using a Bayesian network structure combined with Bayes' theorem, and the formula is as follows: In the formula This represents the posterior probability of a volcanic event. Indicates that the parent node is At the same time, multi-source indicator data The joint probability distribution, Indicates the use of normalization The marginal probabilities are obtained by summing using the method of total probability, and the formula is: In the formula , These represent multi-source indicator data under conditions of volcanic events and no volcanic events, respectively. The joint probability distribution, , Let represent the prior probabilities of a volcanic event occurring and the probability of no volcanic event occurring, respectively. When determining a volcanic event, the posterior probability of the event occurring is compared with a preset probability threshold. If the following conditions are met: Then it is believed that in the first A volcanic event occurred within a time window, where... This represents the posterior probability of a volcanic event occurring. This represents the preset probability threshold.
5. The paleoclimate change analysis method based on multi-source data fusion according to claim 4, characterized in that: In step S3, the probability of sulfur cycle anomalies is obtained based on the conditional Gaussian distribution of the S isotope content at the corresponding nodes. The specific logic is as follows: Based on the conditional Gaussian distribution of the S isotope content at the corresponding nodes, the corresponding mean and standard deviation are obtained. The normal range of the sulfur cycle is defined based on the mean and standard deviation of S isotope content. When the S isotope content is within the normal range, the sulfur cycle is considered normal, and when the S isotope content is outside the normal range, the sulfur cycle is considered abnormal. Define the upper and lower limits of the normal range as the mean and... Sum of standard deviations, mean and The difference of two standard deviations This represents a multiplier factor ranging from 2 to 4. The probability of a normal sulfur cycle is calculated based on the cumulative distribution function of the standard Gaussian distribution, and then the probability of an abnormal sulfur cycle is calculated based on the probability of a normal sulfur cycle. The posterior probability of volcanic events is used to weight the probability of sulfur cycle anomalies to generate the posterior probability of sulfur cycle anomalies. When determining sulfur cycle anomalies, the posterior probability of the sulfur cycle anomaly is compared with a preset probability threshold. If the posterior probability of the sulfur cycle anomaly is not lower than the preset probability threshold, then it is considered to be an anomaly at the [missing information - likely a specific timeframe]. An anomaly in the sulfur cycle occurred within a specific time window.
6. The paleoclimate change analysis method based on multi-source data fusion according to claim 5, characterized in that: If a volcanic event and a sulfur cycle anomaly occur within a certain time window, it is considered that a volcano-driven sulfur cycle disturbance event exists within that time window.
7. A paleoclimate change analysis system based on multi-source data fusion, characterized in that: The analysis system is used to execute the analysis method as described in any one of claims 1-6, specifically including: The data acquisition module is used to collect multi-source index data on volcanoes and sulfur cycles in geological sedimentary rocks, standardize the multi-source index data, align the data based on geological time, and divide the geological time into several time windows according to the time axis. The algorithm construction module is used to construct a Bayesian network structure containing volcanic events and observation indicators, define the causal relationships between nodes in the Bayesian network structure, and set the conditional probability distribution. The comprehensive judgment module is used to generate the probability of volcanic events and sulfur cycle anomalies in different time windows, and to judge volcanic events and sulfur cycle anomalies in combination with preset probability thresholds. The data analysis module is used to identify volcano-driven sulfur cycle disturbance events and their corresponding time windows, and to match the time windows with preset paleoclimate event labels in order to analyze the impact of volcanic eruptions on paleoclimate.
Citation Information
Patent Citations
Method and device for calculating formation time window and distribution range of hydrocarbon source rock and gypsum salt cover layer
CN117195545A
Markov chain Monte Carlo-based paleoclimate inversion method, system and equipment
CN119476054A
Retrodicting Source-Rock Quality And Paleoenvironmental Conditions
US20100175886A1