Method and system for correlating geological datasets

GB2643657APending Publication Date: 2026-02-25UNIVERSITY OF DURHAM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
GB2025017076
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2026-02-25

AI Technical Summary

Technical Problem

Existing methods for correlating multiple geological datasets are time-consuming, subjective, and lack quantitative uncertainty estimates, particularly when using visual alignment and traditional dynamic time warping (DTW), which do not accommodate prior knowledge or provide sedimentation rate information.

Method used

A method and system for aligning multiple geological datasets using Bayesian inference and Markov chain Monte-Carlo (MCMC) techniques to determine z scale factors and offsets, allowing simultaneous correlation of datasets with probabilistic uncertainties, incorporating lithology information to reduce computation and uncertainty.

Benefits of technology

Provides accurate, quantitative information on sedimentation rates and ages with reduced uncertainty by aligning datasets on a common scale, enabling efficient correlation of multiple datasets with probabilistic uncertainties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Disclosed is a computer-implemented method for correlating multiple quantitative geological datasets obtained at one or more geographic locations. The method comprises: receiving multiple datasets including a first dataset and one or more second datasets, each dataset comprising a plurality of data points corresponding to measurements of a geological property of material obtained at different z coordinates at a respective geographic location; and aligning the multiple datasets on a common z scale by: determining a measure of correlation of the multiple datasets by fitting a single curve to the data points of the multiple datasets and calculating a deviation of the data points of the multiple datasets from the fitted curve; determining one or more separate z scale factors and z offsets to apply to each of the one or more second datasets based at least in part on the measure of correlation of the multiple datasets and corresponding predefined probability distributions for the z scale factors and z offsets for each second dataset; and modifying the one or more second datasets by applying each z scale factor and z offset to the corresponding second dataset. Also disclosed is a system for implementing the method.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]METHOD AND SYSTEM FOR CORRELATING GEOLOGICAL DATASETS Technical Field The disclosure related generally to a computer-implemented method and system for correlating multiple quantitative geological datasets obtained at one or more geographic locations. Background Understanding the stratigraphy of sedimentary rocks underpins the study and reconstruction of the evolution of life from the fossil record, as well as the exploration of subsurface resources such as water, minerals, hydrocarbons such as oil and gas, and geothermal energy. In many instances, detailed geological or stratigraphic data is only available from isolated outcrops or wells, and correlating these data is necessary to understand the three-dimensional structures of the subsurface. This allows for the reconstruction of relative and sometimes absolute ages of strata containing fossils or resources, and for evaluating e.g., the continuity and extent of resource-bearing deposits. State of the art techniques for stratigraphic correlation using quantitative geological data, such as geochemical or geophysical rock properties, include visual correlation done manually, and computational (automated) correlation using dynamic time warping or other methods. These approaches present challenges for accurate quantitative analysis. Visually aligning data sets entails matching peaks, troughs and trends in one dataset to those of another. This may be done by plotting the data next to each other, and drawing correlation lines, and / or by shifting one data set relative to the next until it roughly matches (e.g., Bowyer et al. "Implications of an integrated late Ediacaran to early Cambrian stratigraphy of the Siberian Platform, Russia"^GSA Bulletin^135, no.9-10 (2023): 2428- 2450). However, visual alignment methods are time consuming and incur large uncertainty as they are dependent on the individual performing the method. Furthermore, the uncertainty of a particular alignment solution is difficult to quantify, and it is more difficult still to manually explore and evaluate which of several potential solutions is more likely. Computational evaluation of the alignment between two or more datasets can overcome some of the problems associated with visual correlation. In particular, dynamic time warping (DTW) is a technique commonly used to find an optimal alignment between two signals. The DTW approach works on a pair-wise basis, whereby a pair of datasets are stretched and shifted in a non-linear fashion to achieve alignment, resulting in each individual data point in one dataset being mapped to one or more data points in another, or vice versa (see e.g., US 11,391,856 B2). However, traditional implementations of DTW only finds one optimal solution for a pair of datasets without uncertainty estimates, does not accommodate prior knowledge, and does not yield quantitative information on the sedimentation rates or ages in the datasets. There is thus a need for an improved robust technique for simultaneously correlating multiple (i.e. more than two) geological datasets and providing quantitative information on sedimentation rates and ages along with their respective uncertainties, whilst also allowing for use of more than one type of proxy data in alignment. Aspects and embodiments of the invention are set out below and in the appended claims, and aim to address at least some of the above-described problems, and related problems. According to a first aspect of the disclosure, there is provided a method for correlating multiple quantitative geological datasets obtained at one or more geographic locations. The method comprises: receiving multiple datasets including a first dataset and one or more second datasets, each dataset comprising a plurality of (measurement) data points corresponding to measurements of a geological property of material obtained at different z coordinates at a respective geographic location; and aligning the multiple datasets on a common z scale. The geological property may be selected from a group containing: stratigraphic (proxy) data, geochemical data, geophysical data, palaeontology data, palaeobiological data, biomolecular data, and spectroscopic data. The z-coordinates of the datasets may correspond to a measurement height or depth relative to a reference point at the respective geographic locations. Aligning the multiple datasets on a common z scale comprises: determining a measure of correlation of the multiple datasets; determining a separate z scale factor and z offset to apply to each of the one or more second datasets based at least in part on the measure of correlation of the multiple datasets and preferably also corresponding predefined probability distributions for the z scale factor and z offset for each second dataset; and modifying the one or more second datasets by applying each z scale factor and z offset to the corresponding second dataset. Preferably, the measure of correlation is determined by fitting a single curve, e.g. a cubic spline curve, to the data points of the multiple datasets (which can be referred to as composite data) and calculating a deviation of the data points of the multiple datasets from the fitted curve. In some embodiments, the one or more second datasets include a plurality of second datasets. Each dataset preferably comprises data of the same measurement type which is to be correlated, i.e. the same measured geological property or signal. In some examples, each dataset can comprise multiple different data types measured at the same geographic location (e.g. geophysical data and geochemical data). In this case, the multiple datasets can be correlated based on the multiple types of measurement data. The z offset(s), z scale factor(s) and optionally any further parameters used in the correlation may be referred to herein collectively as alignment parameters or an alignment parameter set. The method may advantageously allow simultaneous correlation of any number of geological datasets based on one or multiple measurement data types, and provides useful quantitative information preferably with probabilistic uncertainties, including the z scale factor and z offset as an output of the model. In some embodiments, each dataset can be modelled as having a single z scale factor. In this case, aligning the multiple datasets on a common z scale comprises determining a z scale factor and z offset to apply to all the z-coordinates of the corresponding second dataset. This may provide a simple correlation option with reduced processing requirements. Optionally, the method may include determining one or more further z offsets for applying to a subsection or at least a portion of data points of at least one of the second datasets. This may allow gaps in the data (i.e. disconformities) to be accounted for in the correlation / alignment process. For example, a further offset may be applied to data points on one side (i.e. above or below, on the height or depth scale) of an identified / determined gap in the data (e.g. applied at a predefined gap position). Preferably, each dataset can be modelled as having multiple z scale factors. In this case, aligning the multiple datasets on a common z scale preferably further comprises: partitioning each of the one or more second datasets into a plurality of contiguous subsections, determining at least one separate z offset to apply to each second dataset and a separate z scale factor to apply to each subsection of each second dataset based at least in part on the measure of correlation of the multiple datasets and preferably also corresponding predefined probability distributions for the z offset and z scale factors for each second dataset. The probability distributions for z offset and z scale factor may be the same or different for each subsection of a given second dataset. The at least one z offset preferably includes a z offset to apply to all the data points of the respective second dataset, and optionally one or more separate (further) z offsets to apply to one or more respective subsections of at least one of the second datasets. The one or more further z offsets are preferably applied at respective locations / positions (i.e. z-coordinates) of one or more identified gaps in the respective second dataset. Aligning the multiple datasets may then further comprise modifying the one or more second datasets by applying each determined z scale factor to the corresponding subsection of the corresponding second dataset, and applying each determined z offset to the corresponding dataset or subsection thereof. This allows quantitative information to be obtained subsection-wise for each dataset of the plurality of second datasets resulting in a high granularity of useful quantitate information. Provision of subsection-wise offsets allows gaps in the data (disconformities) to be accounted for in the correlation. In some embodiments, wherein only the one or more second datasets are modified to align to the first dataset, correlation is performed on a common height scale, whereby each z offset corresponds to a height or depth offset, and each z scale factor corresponds to a relative sedimentation rate with respect to the first dataset (i.e. a reference dataset). This allows quantitative correlation to be performed without absolute age estimates for any of the data points. In some embodiments, at least one of the multiple datasets preferably further includes at least one reference age data point corresponding to a (estimated) geological age of material at a given z coordinate at the respective geographic location. In this case, correlation can be performed on a common age scale, wherein the z offsets correspond to an age offset, and the z scale factors correspond to an absolute sedimentation rate. Aligning the multiple datasets may then comprise: converting the multiple datasets from a height or depth scale to an age scale based on initial sedimentation rates and initial age offsets for the respective first and second datasets; and determining a separate final sedimentation rate and final age offset to apply to each of the first and second datasets based at least in part on the measure of correlation of the converted multiple datasets, the at least one reference age data point, and preferably also corresponding predefined probability distributions for the sedimentation rates and age offsets for each of the first and second datasets; modifying the first dataset by applying the corresponding final sedimentation rates and age offsets; and modifying the one or more second datasets by applying the corresponding final sedimentation rates and age offsets. Preferably, where each dataset can be modelled as having multiple sedimentation rates, aligning the multiple datasets on a common age scale further comprises: partitioning the first dataset and the one or more second datasets into a plurality of contiguous subsections; converting the multiple datasets from a height or depth scale to an age scale based on initial sedimentation rates and initial age offsets for each dataset or subsection thereof; determining at least one separate final age offset to apply to each dataset and a separate final sedimentation rate to apply to each subsection of each of the first and second datasets based at least in part on a measure of correlation of the converted multiple datasets, the at least one reference age data point, and preferably also corresponding predefined probability distributions for the sedimentation rates and age offsets for the first and second datasets. The probability distributions for the sedimentation rates and age offsets may be the same or different for each subsection of a given dataset. The at least one final age offset preferably includes an age offset to apply to all the data points of the respective dataset, and optionally one or more separate (further) age offsets to apply to one or more respective subsections of at least one of the datasets. The one or more further age offsets are preferably applied at respective locations of one or more identified gaps in the respective dataset. Aligning the multiple datasets may then further comprise: modifying the first dataset by applying the final sedimentation rates the corresponding subsections of the first dataset and applying the at least one final age offset to first dataset or the corresponding subsections thereof; and modifying the one or more second datasets by applying the final sedimentation rates corresponding to the subsections of the one or more second datasets and applying each final age offset to the corresponding second dataset of subsection thereof. Alternatively, where each dataset is modelled as having a single sedimentation rate, aligning the multiple datasets on a common age scale may comprise determining a separate final sedimentation rate and a separate final age offsets to apply to all the z-coordinates of each of the first and second datasets. Optionally, one or more further final age offsets may be determined for applying to a subsection of data points of at least one dataset, e.g. to account for gaps as described above. Where the first and / or second datasets are partitioned, the partitioning is preferably based at least in part on lithology information associated with the respective dataset. Lithology information preferably partitions datasets in stratigraphic intervals or subsections (on the height / depth scale) that may have different relative or absolute sedimentation rates (i.e. different lithology types), and may further indicate locations of gaps or disconformities. Preferably, at least some partitions between subsections are associated with a change in lithology type as indicated by the lithology information. Advantageously, incorporating known lithologies into the correlation may reduce the amount of computation required in order to output a correlation with a given threshold uncertainty. Furthermore, it may reduce the uncertainty associated with the optimal (most likely) correlation / alignment solution. In such examples, the lithology information may be used as a constraint on the alignment of the multiple datasets such that subsections associated with the same lithology type within a given dataset or across all datasets share the same z scale factor. Where subsections associated with the same lithology type share the same z scale factor across all datasets, aligning the multiple datasets may further comprise: determining, for at least one of the multiple datasets, a subsection-agnostic multiplier to be applied to the z scale factors of all subsections of the respective dataset, based at least in part on the measure of correlation; and modifying the at least one of the multiple dataset by applying the determined z scale factors, z offsets and multiplier for the corresponding dataset. This may, for example, be useful in scenarios where lithologies are shared across the different datasets and the relative sedimentation rates for the different lithological types / units are assumed to be constant across the datasets, but the actual sedimentation rates are expected to vary systematically between the datasets, for example due to varying distances from a sediment source. Gap positions are preferably predefined / predetermined, e.g. provided in user input data or included in lithology data. As such, the method may comprise determining a further z offset (or age offset) to apply to at least one of the datasets at a predefined / predetermined gap position in the respective dataset. The method may further comprise determining, in at least one of the datasets, a position (i.e. a z-coordinate) of a gap in the data points, and determining a further z offset (or age offset) to apply to the respective dataset at the position of the gap. The gap position(s) may be determined based on lithology information or user input data associated with the datasets. The z offsets and z scale factors may further be defined based on a user-defined sedimentation rate model for the datasets. The method may comprise receiving user input data for the datasets specifying one or more of: a sedimentation rate model for the multiple datasets; one or more gap positions of one or more datasets; whether correlation / alignment is to be performed on a height scale or an age scale; and one or more parameters for defining the probability distributions of the alignment parameters. Whether the datasets are modelled as including one or multiple sedimentation rates can be specified in the sedimentation rate model. The sedimentation rate model may be selected from a group comprising: (i) one sedimentation rate per dataset; (ii) multiple sedimentation rates per dataset (involving partitioning the datasets into subsections); (iii) multiple sedimentation rates per dataset based on lithology information, wherein occurrences of the same lithology in the datasets share the same sedimentation rate (involving partitioning the datasets into subsections based on lithology information); (iv) multiple sedimentation rates per dataset based on lithology information, wherein occurrences of the same lithology in the datasets can have separate sedimentation rates; (v) as in (iii) but additionally including sedimentation rate multipliers (e.g. applied to all but one of the datasets). The inclusion and positions of any gaps may also be specified in any of the sedimentation rate models. Preferably, determining the z scale factors and z offsets comprises performing Bayesian inference. In such cases, the predefined probability distributions for the z scale factors and z offsets are respective predefined prior distributions for the z scale factors and z offsets. Performing a Bayesian inference may comprise: determining the measure of correlation of the multiple datasets; defining a likelihood for the correlation of the multiple datasets based at least in part on the measure of correlation; performing Bayesian inference to determine posterior probability distributions for the z offsets and the z scale factors for each second dataset or each subsection thereof, based on the likelihood and the corresponding predefined prior distributions; and determining the z-offsets and z scale factors to be applied to each second dataset or each subsection thereof based on the corresponding posterior probability distributions. Preferably, where the correlation is performed on an age scale, the predefined probability distributions for the sedimentation rates and age offsets are respective predefined prior distributions for the sedimentation rates and age offsets, and wherein performing the inference comprises: determining a measure of correlation of the converted multiple datasets; defining a likelihood for the correlation of the converted multiple datasets based on the measure of correlation of the converted multiple datasets and the at least one reference age data point; performing Bayesian inference to determine posterior probability distributions for the final sedimentation rates and final age offsets for each dataset or subsection thereof, based on the likelihood and the corresponding predefined prior distributions; and obtaining the z-offsets and z scale factors to be applied to each dataset or each subsection thereof from the corresponding posterior probability distributions. In such examples, the at least one reference age data point is preferably used as a constraint placed on the likelihood to induce the Bayesian inference to produce final sedimentation rates and age offsets that tie specific data points in the aligned multiple datasets to the geological age associated with the at least one reference age data point. Advantageously, performing Bayesian inference allows final z offsets and z scale factors to be determined that include some prior knowledge or belief and considers the likelihood of the received data based on this belief. It also may allow uncertainties for the parameters to be determined based on the posterior probability distributions. As such, the method may further comprise determining an uncertainty for each final z offset and z scale factor based on the corresponding posterior probability distributions. The priors may be defined based at least in part on user input data. Preferably, the Bayesian inference is performed using a Markov chain Monte-Carlo (MCMC) method. In such examples, the MCMC method may comprises performing a MCMC chain including: performing an initial iteration of: generating initial proposals for the z offsets and the z scale factors for each dataset or subsection thereof determining a measure of correlation of the multiple datasets resulting from applying the initial proposals; determining a posterior probability density for the initial proposals based on a prior probability of the initial proposals and the likelihood of the correlation; and performing one or more further iterations of: generating further proposals for the z offsets and z scale factors for each dataset or subsection thereof; determining a new measure of correlation of the multiple datasets resulting from applying the further proposals (e.g. by fitting a new single curve to the multiple datasets as aligned based on the further proposals); determining a posterior probability density for the further proposals based on a prior probability of the further proposals and a revised likelihood of the correlation, wherein the revised likelihood is based on the new measure of correlation. The MCMC method may further comprise generating final posterior probability distributions for the z offsets and z scale factors for each dataset or subsection thereof based on the proposals; and obtaining final z offsets and z scale factors for each dataset or subsection from the corresponding final posterior probability distribution. The final posterior probability distributions may comprise multiple peaks corresponding to different possible alignments (i.e. different possible final z offsets and the z scale factors that can describe the data). In this case, the method may further comprise, determining optimal or most likely final z offsets and the z scale factors (i.e. a most likely alignment) based on a cluster analysis of the z offsets and the z scale factors proposed in a plurality of iterations of the MCMC chain (i.e. a cluster analysis on the history of proposals). This may involve determining a probability for each possible alignment solution based on the cluster analysis and determining the optimal or most likely final z offsets and the z scale factors (i.e. a most likely alignment) based on the corresponding probabilities. The further iterations may further comprise: determining an acceptance probability based on the posterior probability density and the posterior probability density of the previous iteration; and accepting or rejecting the further proposals based on the acceptance probability. In this case, the final posterior probability distributions may be generated based on the accepted proposals. The use of MCMC is particularly beneficial for numerically calculating the posterior distribution and performing an intelligent search in a vast high dimensional space. The initial proposals for the z offsets and the z scale factors may be generated by sampling from a corresponding initial distribution, and preferably wherein the corresponding initial distribution is the corresponding prior probability distribution. In some examples, the further proposals for the z offsets and the z scale factors are generated by one or more of: (i) sampling from the corresponding prior probability distribution; (ii) sampling from a corresponding normal distribution defined with respect to the respective z offset and the z scale factor proposed in the previous iteration with a standard deviation based on the respective z offsets and the z scale factors proposed in a plurality of the previous iterations; (iii) sampling from a corresponding multivariate normal distribution defined with respect to the respective z offset and the z scale factor proposed in the previous iteration with a covariance matrix based on the respective z offsets and the z scale factors proposed in a plurality of the previous iterations; (iv) sampling from one of a plurality of a corresponding multivariate normal distributions which is centred closest to the z offset and the z scale factor proposed in the previous iteration, wherein the plurality of multivariate normal distributions are defined based on a cluster analysis of the z offsets and the z scale factors proposed in a plurality of previous iteration; and (v) sampling a further proposal for one of the z offset and the z scale factor according to any of (i) to (iv), while keeping the other of the z offset and the z scale factor constant. In this case, the method may comprise determining a proposal type to use based on a probability for each proposal type. The probability for each proposal type may be based on the acceptance probability of the respective proposal type, i.e. proposal types that are rejected often are chosen less frequently. Optionally, in some examples the MCMC chain is a target chain, and the method further comprises performing one or more further MCMC chains in parallel to the target chain, wherein the posterior probability densities for each further chain are modified by applying a predefined temperature factor, and wherein the method comprises, after the step of accepting or rejecting the further proposals: determining a swap probability for each further MCMC chain based on the posterior probability density of the target MCMC chain and the posterior probability density of the respective further MCMC chain; and swapping the proposals for the z offsets and the z scale factors generated in the target MCMC chain with those of a further MCMC chain based on the swap probabilities. Beneficially, this avoids the chain becoming trapped at isolated peaks of a posterior probability. In some examples the MCMC method is performed according to a Metropolis-Hasting algorithm or approach. In some examples, the MCMC method is performed according to a Metropolis-within-Gibbs algorithm or approach. This combination of MCMC approaches may be particularly effective for cases (such as this one) where the posterior distribution is difficult to sample from directly. The method may comprise outputting the modified datasets, optionally with one or more of the fitted curve, final z offsets and z scale factors, and corresponding uncertainties. Outputting may comprise outputting a plot or a graphical representation of the modified datasets (and optionally with the fitted curve and / or uncertainties) on the common z scale. Preferably, the method comprises generating a probabilistic height-height model or age-height / age- depth model for material at the respective geolocations based at least in part on the aligned multiple datasets and the final z offsets and z scale factors. The height-height model may define a height or depth in the first (reference) dataset for any height or depth in the one or more second datasets preferably together with a probabilistic uncertainly. The age-height model may define an age for any point within the z-range of the aligned multiple datasets. Preferably, the age-height or age-depth model further defines a measure of uncertainty for each age. The uncertainties in the height-height model or age-height / age-depth model can be obtained from the final posterior distributions of the corresponding z offsets and z scale factors. The method comprises outputting the probabilistic height-height model or age-height / age-depth model. Optionally, the method may comprise generating or outputting a graphical representation of the height-height model or age-height / age-depth model. According to a second aspect of the disclosure, there is provided a system for correlating quantitative geological datasets obtained at one or more geographic locations. The system comprises one or more processors and memory storing instructions that, when executed by the one or more processors, cause the one or more processors (or system) to perform the method of the first aspect. According to a third aspect of the disclosure, there is provided a non-transient computer readable medium, comprising instructions that, when executed by one or more processors or a computer, cause the one or more processors or a computer to perform the method of the first aspect. Any method feature as described herein may also be provided as a system feature, and vice versa. As used herein, means plus function features may be expressed alternatively in terms of their corresponding structure. Any, some and / or all features in one aspect of the disclosure may be applied to other aspects of the disclosure, in any appropriate combination or sub-combination. In particular, system aspects may be applied to method aspects, and vice versa. It should also be appreciated that particular combinations of the various features described and defined in any aspect of the disclosure can be implemented and / or supplied and / or used independently. The disclosure extends to methods, systems, computer readable mediums and devices substantially as herein described and / or as illustrated with reference to the accompanying figures. The disclosure also extends to any novel aspects or features described and / or illustrated herein. In this specification the word 'or' can be interpreted in the exclusive or inclusive sense unless stated otherwise. Brief Description of Drawings In order that the disclosure can be well understood, aspects and embodiments will now be discussed by way of example only with reference to the accompanying drawings, in which: Figure 1 shows example raw datasets of geological data; Figure 2 shows a schematic diagram of a dynamic time warping technique; Figures 3(a) to 3(d) show further example geological datasets; Figure 4 shows a method for correlating multiple geological datasets according to embodiments of the disclosure; Figure 5 shows a method for aligning multiple geological datasets on a common z scale according to embodiments of the disclosure; Figure 6 shows an example partitioning of a second dataset into a plurality of contiguous subsections; Figure 7(a) shows example lithology data corresponding to the geological datasets shown in figures 3(a) to 3(d); Figure 7(b) shows example partitioning of the second datasets in figures 3(a) to 3(d) based on the corresponding lithology data in figure 7(a); Figure 8 shows a method for aligning multiple geological datasets on a common age scale according to embodiments of the disclosure; Figures 9(a) to 9(d) show example datasets partitioned into subsections; Figure 9(e) shows the multiple datasets of figures 9(a) to 9(d) aligned on a common age scale according to the method of figure 4; Figures 10(a) and 10(b) show illustrative examples of prior probability distributions; Figure 11(a) shows an example of a single curve fitted to data from all datasets; Figure 11(b) show an example of an age constraint used to determine a likelihood; Figure 12 shows a method for numerically obtaining z scale factors and z offsets for correlating multiple geological datasets according to embodiments of the disclosure; Figure 13 shows an example of a final posterior probability density distribution for a sedimentation rate of a dataset; Figure 14 illustrates a probabilistic age-height model for the datasets in figure 9 produced from the method of the present disclosure; Figures 15(a) and (b) show further example geological datasets corresponding to different locations; Figure 16(a) to 16(d) show different alignment solutions for the datasets of figures 15(a) and 15(b) produced by the correlation method of the present disclosure; Figure 17 shows the posterior samples of the parameters ^ and ^ for the dataset in figure 15(b) in a histogram plot showing clusters corresponding to the possible alignment solutions; Figure 18 shows the resulting probabilistic depth model for the datasets in figure 15(a) and 15(b) produced from the correlation; Figure 19 shows an alignment solution for the datasets shown in figure 7(b) on a common age scale; Figures 20(a) – 20(c) illustrate probabilistic age-height models for the datasets of figure 7(b) produced from the method of the present disclosure; Figure 21 shows a schematic block diagram of an example system for implementing the methods of the present disclosure; and Figure 22 shows a schematic block diagram of an example computing device of the system. It should be noted that many of the figures are diagrammatic and may not be drawn to scale. Relative dimensions and proportions of parts of these figures may have been shown exaggerated or reduced in size, for the sake of clarity and convenience in the drawings. The same reference signs are generally used to refer to corresponding or similar features in modified and / or different embodiments. Detailed Description In the present disclosure, the term “geological” data may be used to refer to any data or measurement of a geological property of a material which varies across depth or height and can thereby provide information on strata. In some instances, geological data is referred to herein as stratigraphic data. Geological or stratigraphic data may include but is not limited to: geochemical data such as an elemental concentration or an isotope ratio such as carbon isotopes (δ13C) and sulphur isotopes (δ34S), Fe2O3 concentrations, Si / Al elemental ratios, or87Sr / 86Sr etc.; geophysical data such as a gamma ray signal, neutron porosity measurements, resistivity measurements, a material density or seismic velocity; biomolecular data such as sterane abundance, spectroscopic data such as Raman data, and palaeontological and palaeobiological data such as the relative or absolute abundance of particular species, total taxonomic diversity / abundance, the presence or absence of species or groups thereof, or morphological information (such as size of fossil, number of spines, etc.) associated with particular groups. Correlations of these geological data are attempted as it is assumed that there are common trends in the data that reflect geological events such as changes in the sediment supply or seawater composition. Geological data may be collected invasively, for example by drilling a borehole or well into the ground through the strata and taking measurements of a geological property of material at varying depths. In other examples, geological data may be collected non-invasively, for example, by taking measurements of a geological property of material at varying heights on a mountain face, surface, or rock outcrop. Indeed, in some cases geological or stratigraphic data can be simulated, for example for testing purposes. The nature of obtaining the data may be specific to the type of data obtained. Herein sedimentation rate refers to the rate at which sedimentary material is deposited over time expressed in terms of meters per unit of time (typically in units of meters per million years, m / Myr, or centimeters per thousand years, cm / kyr), and age is typically expressed in units of million years ago, Ma. By way of example, 1 Myr would typically denote a duration (i.e. the length of an interval of time), whereas 1 Ma would refer to a specific point in time (analogous to 2024 denoting the present year). Figure 1 shows example first and second datasets 100A, 100B of geological data each comprising a plurality of data points 11A, 11B corresponding to measurements of a geological property of material obtained at different z coordinates (height in this example) at a respective geographic location. Traditionally, correlating stratigraphic datasets such as the first and second datasets 100A, 100B would be done manually by eye, whereby the datasets 100A, 100B are plotted side-by-side and equivalent features such as peaks and troughs are identified from visual inspection and correlation lines are drawn between the corresponding features of the datasets 100A, 100B. This may also involve shifting one dataset relative to the next until it roughly aligns. By way of example, features F1a, F2a, F3a in the first dataset 100A can be visually correlated with corresponding features F1b, F2b, F3b of the second dataset 100B, as indicated by the dashed correlation lines C1, C2, C3. The end product of this approach is thus a visual diagram of the datasets 100A, 100B including numerous correlation lines linking height or depths in the different datasets 100A, 100B. Visual alignment techniques, although commonly used in the field, have several problems. They are time consuming, have high levels of uncertainty which are difficult to quantify, are subjective and require skilled interpretation to derive any quantitative information, such as age and sedimentation rates. This uncertainty becomes particularly prevalent when visually aligned datasets are further processed to convert from a height or depth scale to an age scale and / or when data is obtained from geological formations without any corresponding age information to tie a particular height or depth to an age (e.g. radiometric dates). In addition, there can often be multiple possible alignment solutions which are challenging to evaluate manually. Computational correlation techniques aim to automatically manipulate and align datasets which addresses some of the problems with visual alignment. Common computational approaches are based on dynamic time warping (DTW). DTW is a technique to find the optimal alignment between two signals. The DTW approach works on a pair-wise basis, whereby a pair of datasets such as datasets 100A, 100B are stretched and shifted in a non-linear fashion to achieve alignment of individual data points. In general, DTW works by calculating a cost matrix that represents distances between points in one dataset to points in another. For each pair of points from the pair of datasets the cost of aligning that pair of points with one another is computed, and the path through the cost matrix that minimises the total cumulative cost is taken as the best or optimal alignment. As such, DTW maps each individual data point in one dataset to one or more data points in another, or vice versa. This is illustrated schematically in figure 2 whereby the dashed lines indicate mappings between data points 11A, 11B of schematic datasets 10A 10B. Despite accurately mapping and aligning data points in stratigraphic data, in its basic form, DTW is limited to aligning pairs of datasets and only finds one optimal solution without uncertainty estimates. It also does not accommodate prior knowledge or known radiometric dates, and does not provide quantitative information for the datasets, such as sedimentation rates as an output to the method. In practice, there are often multiple datasets, e.g. 3, 4 or more datasets, which require correlating. Figures 3(a) to 3(d) show first to fourth example geological datasets 100A, 100B, 100C, 100D each comprising a plurality of data points corresponding to measurements of a geological property (δ13C) of material obtained at different z coordinates (height in these examples) at respective geographic locations. As seen, the height scale for each dataset is different. There is thus a requirement for techniques that can simultaneously correlate multiple datasets, preferably more than two datasets, and provide quantitative information on sedimentation rates and ages and their respective uncertainties. Figure 4 shows a computer-implemented method 400 for automatically correlating multiple quantitative geological datasets obtained at one or more geographic locations according to embodiment of the disclosure. The method 400 can be implemented on a standalone computing device or on a networked computing device such as a client-server system as described below with reference to figures 21 and 22. The method 400 can be applied, for example, to the multiple datasets 100A-100D described above with reference to figure 3. The method 400 can be defined broadly as encompassing two main steps. At step 410, multiple datasets 100A-100D are received, including a first dataset (for example, dataset 100A of figure 3(a)) and one or more second datasets (for example, datasets 100B-100D of figures 3(b)-3(d)), each containing a plurality of data points corresponding to a geological property at different z-coordinate associated with a respective geographic location. The datasets 100A-100D can comprise measured or simulated geological data, and the respective geographic locations may be the same or different. Each dataset preferably comprises the same data or signal type which is to be correlated, often referred to as a proxy type (e.g. a carbon isotope ratio δ13C signal). In some examples, each dataset can comprise multiple different data types measured at the same geographic location (e.g. some geophysical data, and some geochemical such as gamma ray signals). In this case, like data types are correlated with like data types, but their correlation process is combined and interrelated (i.e. not independent) based on the assumption that each data type records the same sedimentary changes, as described in more detail below. In this way, the datasets can be correlated based on the multiple data types. At step 420a, 420b, the multiple datasets 100A-100D are aligned to a common z scale by determining separate z scale factors and z offsets to apply to at least one of the multiple datasets 100A-100D based at least in part on a measure of correlation of the multiple datasets 100A-100D and, preferably, corresponding predefined probability distributions for the z scale factors and z offsets, as discussed in more detail below. Each of the at least one of the datasets 100A-100D is then modified by applying the determined z scale factors and z offsets to the corresponding dataset(s) 100A-100D. The common z-scale can include a vertical distance scale such as height or depth, or an absolute age scale if prior information relating to age is known. In the context of the present disclosure, the modified datasets 100A-100D which have been aligned according to the method 400 are substantially correlated, such that their signals vary in a common way along the z scale (i.e. age or height / depth). Correlation of the multiple datasets 100A-100D can be achieved generally two ways depending on whether any absolute age estimates are available for any data points of the datasets 100A-100D. In one embodiment, in step 420a, the multiple datasets 100A-100D are aligned on a common height or depth scale. In this case, one of the multiple datasets 100A-100D, i.e. the first dataset 100A, is selected as a reference dataset and the rest of the multiple datasets 100A-100D, i.e. the one or more second datasets 100B-100D, are shifted and rescaled relative to the first (reference) dataset by applying z offsets and z scale factors to align to the first (reference) dataset 100A. The common z- scale is thus the z-scale (i.e. height or depth) of the first (reference) dataset 100A, and the multiple datasets 100A-100D can be aligned without absolute age estimates associated with any of the datasets 100A-100D. This approach provides relative sedimentation rates (related to the z scale factors) for each of the one or more second datasets 100 defined with respect to the first (reference) dataset 100. A relative sedimentation rate indicates whether the corresponding second dataset 100B- 100D has a higher or lower sedimentation rate than the first (reference) dataset 100A. In another embodiment, in step 420b, the multiple datasets 100A-100D are aligned on a common age scale. In this case, all of the multiple datasets 100A-100D are converted to age scale and shifted and scaled relative to each to achieve alignment, whereby the age scale is determined using absolute age estimates associated with at least one of the multiple datasets 100A-100D. In this case, the z offsets and z scale factors become age offsets and absolute sedimentation rates. This approach provides absolute sedimentation rates for each dataset 100A-100D. In preferred embodiments, the measure of correlation of the multiple datasets 100A-100D is determined by fitting a single curve to the composite measurement data from the multiple datasets 100A-100D and calculating a deviation of the data points from the fitted curve. The z offsets and z scale factors for the datasets are then determined probabilistically using Bayesian inference, by exploring a range of different z offsets and z scale factors (i.e. shift and sedimentation rate parameters) for the datasets 100A-100D and evaluating the fit of different alignments of the multiple datasets 100A- 100D (through the measure of correlation), corresponding to the different z offsets and z scale factors, in a Bayesian framework. In this case, prior probability distributions are placed on the z offsets and z scale factors for each dataset 100A-100D and a likelihood is defined based on the measure of correlation of the multiple datasets 100A-100D to quantify how well each alignment explains the observed data, as will be described in more detail below. The method 400 can thus advantageously simultaneously align or correlate any number of datasets from various geographic locations or sites on a common z scale (age or height) through evaluating the fit of the single curve to the composite data to yield absolute or relative sedimentation rates with probabilistic uncertainties. The resulting final fitted curve, sedimentation rates, and uncertainties can be used to produce a probabilistic age-depth or age-height model for strata associated with the multiple datasets 100A-100D, as will be described in more detail below. In contrast to DTW-based techniques, the method 400 employs an interval-based approach rather than a point-based approach and can readily incorporate different kinds of geological information in the Bayesian priors and likelihood, allowing for a practical way of dealing with complex geological data and generating realistic sedimentation rate information. Alignment on a common height scale This section discusses embodiments where the method 400 is used to obtain a correlation of multiple geological datasets on a common height scale. This corresponds to step 420a, whereby the one or more second datasets 100B-100D are shifted and rescaled relative to the first (reference) dataset 100A. In such examples, the z-coordinates of the datasets 100A-100D correspond to a measurement height h relative to a reference point at the respective location, each z offset corresponds to a height offset ^, and each z scale factor corresponds to a relative sedimentation rate ^ with respect to the first dataset 100A. As such, each data point 11A-11D in the datasets 100A-100D comprises a measured property, for example, a ratio of carbon isotopes (δ13C), and a corresponding height h at which the property was measured, as shown in figure 3. Of course, it is understood that because the z- coordinates are in practice measured relative to a reference point, height may take positive or negative values and thus may also include ‘depth’. At step 420a the datasets 100A-100D are aligned to a common height scale by determining separate z scale factors and z offsets to apply to each of the second datasets 100B-100D based at least in part on the measure of correlation of the multiple datasets 100A-100D, and preferably also corresponding predefined probability distributions for the z scale factors and z offsets. Preferably, the z scale factors and z offsets are determined by performing a Bayesian inference (see below). In one example implementation, each dataset 100A-100D is modelled as having a single sedimentation rate. Thus, aligning the datasets 100A-100D requires determining, for each of the second datasets 100B-100D, a single height offset ^^^to shift the entire respective second dataset 100B-100D up or down and anchor the first (or last) data point to a specific height h in the first dataset 100A, and a single relative sedimentation rate ^^^to scale (stretch or squeeze) the given second dataset relative to the reference section 100A. In this case, for any height ℎ^in a given one of the second datasets 100B-100D, the corresponding height hr in the first (reference) dataset 100A can be calculated according to: where ℎ^,^is the height of the first (or last) data point of the given second dataset 100B-100D. A relative sedimentation rate ^^^< 1 implies that the corresponding second dataset 100B-100D has a lower (absolute) sedimentation rate than the first (reference) dataset 100A, and consequently the corresponding second dataset 100B-100D has to be stretched to match the first (reference) dataset 100A. Conversely, a relative sedimentation rate ^^^> 1 implies a higher (absolute) sedimentation rate than the first (reference) dataset 100A, and will lead to the corresponding second dataset 100B-100D being squeezed to match the first (reference) dataset 100A. In this example, each second dataset 100B-100D can be converted from its height scale to the reference height scale of the first (reference) dataset 100A using equation 1 for the correlation / alignment process. Further, each second dataset 100B-100D can be modified in step 420 by applying the determined z scale factor ^^^and z offset ^^^to the corresponding dataset using equation 1.   In another example implementation, instead of having one relative sedimentation rate ^^^per second dataset 100B-100D, each second dataset 100B-100D is modelled as having a plurality of sedimentation rates ^^^. In this case, step 420a comprises steps 4201a-4203a as shown in figure 5. As step 4201a, each of the one or more second datasets 100B-100D is partitioned into a plurality of contiguous subsections 100B-l, 100C-l, 100D-l, and at step 4202a a relative sedimentation rate ^^^,^is determined for each subsection of each second dataset 100B-100D based on the measure correlation between the multiple datasets 100A-100D, and preferably also the corresponding probability distributions for the relative sedimentation rates ^^^,^for each second dataset 100B-100D. Once relative sedimentation rates ^^^,^are determined, in step 4203a each second dataset 100B-100D is modified by applying the relative sedimentation rates ^^^,^to the corresponding subsections to thereby align to the first (reference) dataset 100A. In this case, for any height ℎ^in a given one of the second datasets 100B-100D, the corresponding height hr in the first (reference) dataset 100A can be calculated according to: where, ^^^,^is the number of subsections 100B-l, 100C-l, 100D-l encountered from heights ℎ^to ℎ^,and ℎ^,^ െ ℎ^,^ି^ is the height of the subsection 100B-l, 100C-l, 100D-l (for ^^ ^ ^^^,^) , or, if ^^ ൌ ^^^,^ (i.e.the last subsection 100B-l in which height ℎ^is located), the height from the last data point (top) of the previous subsection 100B-l until height ℎ^. In this example, each second dataset 100B-100D is modified in step 4203a by applying the determined z scale factors ^^^,^and z offset ^^^to the corresponding dataset using equation 2. Figure 6 shows an example partitioning of a second dataset 100B into a plurality of contiguous subsections 100B-1 to 100B-l. In figure 6, the subsections 100B-1 to 100B-l are similar in height, however, this is not essential. In general, each dataset 100B-100D can be partitioned into equal or unequal height subsections 100B-l, 100C-l, 100D-l which can be defined arbitrarily / randomly, based on lithology information 200B-200D associated with the respective dataset 100B-100D, or based on locations of certain features in the datasets 100B-100D which may indicate a change in lithology. Preferably, the partitioning in step 4201a is based at least in part on known lithology information or data 200B-200D associated with the respective dataset 100B-100D. A lithology is a known term in the art, referring to a physical characteristic of a material, such as a colour, texture, grain size, and composition which can be used to indicate or classify a lithology type, such as a rock or a sediment type. Lithology data 200B-200D for each dataset 100B-100D comprises a plurality of lithology or stratigraphic intervals or segments 200B-l, 200C-l, 200D-l representing lithology information or type corresponding to a given range of heights (or depths) in the respective dataset 100B-100D. Lithology data 200B-200D may further indicate the presence and locations of gaps G or disconformities (see below). Where lithology information 200B-200D is used, partitions between adjacent subsections 100B-l, 100C-l, 100D-l (i.e. the start or end of a subsection 100B-l) are associated with a change in lithology. In this way, the datasets 100B-100D can be divided up into lithological units. Incorporating known lithologies into the correlation may reduce the amount of computation required in order to output a correlation with a given threshold uncertainty. Furthermore, it may reduce the uncertainty associated with the optimal (most likely) correlation / alignment solution. Figure 7(a) shows example lithology data 200B-200D corresponding to the geological datasets 100B- 100D shown in figure 7(b), illustrating the lithological segments 200B-1, 200C-l, 200D-l (shaded blocks) which can be modelled as having different sedimentation rates. I this example, each different shade corresponds to a different lithology type, and a gap G is indicated in lithology data 200D. Lithology data 200B-200D can be received, for example, at the same point of receiving the multiple datasets 100B-100D in step 410. Figure 7(b) shows example partitioning of the datasets 100B-100D based on the corresponding lithology data 200B-200D in figure 7(a). The sedimentation rates ^^^,^across the subsections of the datasets 100B-100D can be treated or modelled in various ways. In practice, a sedimentation rate model will be chosen or defined for the multiple datasets 100A-100D being correlated (e.g. one sedimentation rate per dataset or multiple, lithological units or not, etc.). In one example, relative sedimentation rates ^^^,^are treated / modelled independently in all subsections of each second dataset 100B-100D. In another example, some subsections in different positions within a given dataset or across the datasets 100B-100D can share relative sedimentation rates ^^^,^. In a further example, where lithology data 200B-200D is used to partition to the datasets 100B-100D the lithology data 200B-200D can be used as a constraint on the alignment of the multiple datasets 100A-100D whereby subsections associated with the same lithology type within a given dataset or across all datasets can share the same relative sedimentation rate ^^^,^. In some examples, the sedimentation rate model is further expanded with the addition of a dataset- specific or subsection-agnostic sedimentation rate multiplier ^^^, in which case equation 2 is modified to: This may, for example, be useful in scenarios where lithologies are shared across the different datasets 100A-100D and the relative sedimentation rates for the different lithological types / units are assumed to be constant across the datasets 100A-100D, but the actual sedimentation rates are expected to vary systematically between the datasets 100A-100D, for example due to varying distances from a sediment source. In this case, where subsections associated with the same lithology type across all datasets share the same relative sedimentation rate ^^^,^, step 4202a can further comprise determining, for at least one of the multiple datasets 100A-100D, a subsection-agnostic sedimentation rate multiplier ^^^to be applied to the relative sedimentation rates ^^^,^of all subsections of the respective dataset based at least in part of the measure of correlation, and step 4203a comprises modifying the respective dataset by applying the determined height offsets ^s, relative sedimentation rates ^^^,^, and multiplier ^^^for the corresponding dataset. Geological datasets 100A-100D can include “disconformities” or gaps G where there is a jump in measured geological properties representing missing data, e.g. due to erosion of sediment at the respective location or other hiatuses such as periods of non-deposition without erosion, as is known in the art. An example gap G is shown in figures 7(a) and 7(b). Gaps G and gap locations may be identified from lithology data 200B-200D or directly from the data (e.g. by detecting discontinuities). In preferred examples, gaps G of height ^g can be accommodated for in the method 400 by introducing additional height offsets ^g at the position of the gap G (which is preferably predefined / user-defined and / or included in the lithology data) or determining a height offset ^^^,^and a relative sedimentation rate ^^^,^, for each subsection. For example, where the second datasets 100B-100D are not partitioned into subsections, equation 1 can be modified to include gaps G according to: where ^^^,^is the number of gaps G encountered up to height ℎ^. Equation 2 can be adapted in a similar manner to accommodate gaps G and subsections. The parameters ^^^,, ^^^, and optionally ^^^and ^gare referred to herein collectively as alignment parameters ^ or an alignment parameter set ^. Alignment on a common age scale This section discusses embodiments where the method 400 is used to obtain a correlation of multiple geological datasets 100A-100D on an age scale. This corresponds to step 420b, whereby all of the multiple datasets 100A-100D (including the first dataset 100A) are shifted and scaled to align on a common absolute age scale (i.e. no dataset is designated as a reference dataset). In such examples, the z offsets correspond to an age offset ^, the z scale factors correspond to an absolute sedimentation rate ^, and at least one of the multiple datasets 100A-100D further includes at least one reference age data point 300A, 300B, 300C, 300D corresponding to a geological age of material at a given z coordinate at the respective geographic location. In this embodiment, the multiple datasets 100A-100D are aligned by first converting them from a height scale to an age scale based on initial alignment parameters ^ including at least initial sedimentation rates ^^^and initial age offsets ^^^(and optionally also ^^^and ^g) for the respective dataset 100A-100D, and then determining final alignment parameters ^ including one or more final sedimentation rates ^^^and final age offsets ^^^(and optionally also final ^^^and ^g) to apply to each dataset 100A-100D based at least in part on the measure of correlation of the converted multiple datasets 100A-100D, the at least one reference age data point 300A-300D, and preferably also the corresponding predefined probability distributions for the alignment parameters ^ for each dataset 100A-100D. Preferably, at least the initial sedimentation rates ^^^and initial age offsets ^^^of the initial alignment parameters ^ are obtained by sampling from the corresponding predefined probability distributions, and the final alignment parameters ^ are determined via a Bayesian inference, as described in more detail below. Where gaps G and multipliers are included in the sedimentation rate model, initial values for these are also sampled from the corresponding predefined probability distributions. The datasets 100A-100D are then modified by applying the corresponding final alignment parameters ^ including at least the final sedimentation rates ^^^and final age offsets ^^^. In the simplest case, where the datasets 100A-100D are modelled with a single sedimentation rate ^^^(i.e. not partitioned), analogous to equation 1 the heights hsin each dataset 100A-100D can be converted to age (^^) according to: where in this case, ^^^is the bottom age, rather than height, of the respective dataset 100A-100D. As described above, sedimentation rates ^^^are absolute sedimentation rates defined on the age scale, rather than relative to a reference section. As described in the above section with reference to aligning on a common height scale, in preferred examples the datasets 100A-100D (now including the first dataset 100A) are each partitioned into subsections 100A-l, 100B-l, 100C-l, 100D-l, and more preferably into lithological units using additional lithological data 200A-200D, whereby the sedimentation rates ^^^,^for each dataset 100A-100D can be modelled in the same various ways described. Additionally, gaps G and / or sedimentation rate multipliers can optionally be included in the same way. In this case, equations 2 to 4 can also be modified accordingly for an analysis on the age scale. In a preferred example, step 420b comprises steps 4201b-4204b as shown in figure 8. Step 4201b comprises partitioning 4201b each of the datasets 100A-100D into a plurality of subsections 100A-l, 100B-l, 100C-l, 100D-l. In step 4202b, each dataset 100A-100D is converted from a height (or depth) scale to an age scale based on initial alignment parameters ^ including at least initial sedimentation rates ^^^,^and initial age offsets ^^^(and optionally also and ^g) for the respective dataset 100A- 100D. Step 4203b comprises determining separate final alignment parameters ^ including at least final sedimentation rates ^^^,^and a final age offsets ^^^to apply to each subsection of each dataset 100A- 100D based at least in part on a measure of correlation of the converted multiple datasets 100A-100D, the at least one reference age data point, and preferably also the corresponding predefined probability distributions for the alignment parameters ^ for each datasets 100A-100D. Where gaps G and / or sedimentation rate multipliers are included in the correlation analysis, steps 4202b-4204b are adapted accordingly to include all the alignment parameters ^ considered in the correlation. Figures 9(a)-9(d) show example multiple datasets 100A-100D partitioned into subsections as described above. Dataset 100A includes corresponding reference age data points 300A as indicated by the asterisks. Figure 9(e) shows the multiple datasets 100A-100D aligned on a common age scale according to the method 400. Figure 9(e) also shows the fitted single curve 10 used to determine the measure of correlation of the datasets 100A-100D and align the datasets, as will be discussed in more detail below. The associated uncertainty 20 of the fit is indicated by the dashed lines. Parameter determination In preferred embodiments, Bayesian inference is used to determine the alignment parameters ^ including at least the final z scale factors ^^ and z offsets ^^ (the shift and sedimentation rate parameters) to be applied to each dataset 100A-100D or subsection thereof. Specifically, a range of different alignment parameters ^ for the datasets 100A-100D are explored and the fit of different alignments of the datasets 100A-100D, corresponding to the different applied alignment parameters ^ (quantified by the fit of a single curve through all the data points of the datasets), is evaluated in a Bayesian framework. Herein, an “alignment” refers to the alignment of the composite data (of a given measurement signal) from all datasets 100A-100D when plotted on a common z-scale having been updated or aligned (i.e. shifted and / or scaled) according to the alignment parameters ^. Prior probability distributions ^^^^^^are placed on the alignment parameters ^ for each dataset 100A-100D, and a likelihood is defined based on the measure of correlation of the multiple datasets 100A-100D, as determined from the calculated deviation from the fitted single curve, (and optionally the age reference points where available) to quantify how well each alignment explains the observed data. Posterior probability distributions for the alignment parameters ^ for each dataset 100A-100D or each subsection thereof are determined based on the likelihood and the corresponding predefined prior distributions. The final alignment parameters ^ to be applied to each dataset 100A-100D or each subsection thereof can then be determined from the corresponding posterior probability distributions. The purpose of a prior probability ^^^^^^in Bayesian inference is to provide some representation and weighting of prior knowledge to the algorithm. In practice, the skilled person will understand that when inferring the final probabilities by performing Bayesian inference over a large number of iterations, the initial prior distribution may be any sensible distribution and is primarily designed to reduce the size of the probability space. Illustrative examples of prior probability distributions for the z offset ^^ and z scale factor ^^ are shown in figures 10(a) and 10(b) for the case of age offset ^^ and absolute sedimentation rate ^^. The z offset distribution is a uniform distribution between specified bounds whilst the z scale factor ^^ distribution is a truncated log-normal distribution centred on a given sediment rate. Using a log scale for the z scale factor ^^ may ensure that ratios are treated symmetrically and ^^ is positive. Prior distributions for gap lengths ^ may take an exponential form. The likelihood of the composite data from all datasets given ^, ^^^^^^^^^^^|^^^, quantifies how well an alignment resulting from applying the alignment parameters ^ explains the observed data. As described above, in preferred embodiments, after the datasets 100A-100D are shifted and scaled by applying a given set of alignment parameters ^, the degree of correlation of the datasets 100A-100D is measured by fitting a single curve 10 to all the datasets, for example a cubic spline. To calculate the likelihood, it is assumed that the errors, i.e. the difference between the observed data, y (the composite data points from all datasets 100A-100D being correlated), and the predicted data, ypred, from the curve 10 are independently and identically distributed according to a normal distribution with a mean of 0 and standard deviation ^. An example of a cubic spline curve 100 fitted to the composite data points from multiple datasets 100A-100D on the age scale is shown in figure 11(a). In one example, the likelihood of an observed (measured) data point ‘yi’ given the applied alignment parameters ^^, is defined according to: where ^ is the standard deviation obtained from the fitted curve 10, and ypred,i is the corresponding predicted data point. The log-likelihood for the entire dataset is then represented by equation 7: log^^^^^|^^^ ൌ ∑^ log^^^^^^|^^^, ^ 7)Where correlation is performed on an age scale the reference age data 300A-300D, which includes discrete age estimates ^^^(e.g. from radiometric dates), is incorporated as age constraints with, for example, mean ages ^^^^^^and uncertainties given by standard deviations ^^^ௗ, as illustrated in figure 11(b). In this case, the probability density of a reference age date point ^^^is then calculated according to: where ^^^^^ௗ,^is the age predicted by the age-height transform at height ℎ^,^, the height in the given dataset at which date ^^^was obtained. The log-likelihood for all age constraints ^^^is then calculated in the same manner as the log likelihood defined in equation 7: log^^^^^|^^^ ൌ ∑^ log^^^^^^|^^^ , ^9^and the overall likelihood, if absolute age constraints are included, is: log^^^^^, ^^|^^^ ൌ log^^^^^|^^^ ^ log^^^^^|^^^. ^10^As such, the likelihood of the data y given ^^ is a product of the probability densities of each data point ^^^of all the datasets and of all absolute age information. It will be appreciated that reference age date points ^^^of a given dataset are shifted by the z offsets and scale factors for that dataset. Posterior probabilities, ^^^^^|^^^^^^^^^, for each alignment are then inferred by combining the prior ^^^^^^and the likelihood ^^^^^^^^^^^|^^^ using Bayes theorem according to: ^^^^^|^^^^^^^^^ ∝ ^^^^^^^^^^^|^^^^^ ^^^^^^,   ^11^where the posterior probability can be thought of as the updated prior given the data observed. The above approach allows simultaneous correlation of multiple datasets 100A-100D. The data in the datasets being correlated is of the same type, e.g. ^13C signals from one dataset 100A are correlated with ^13C signals from another dataset 100B. However, the above approach can be extended to correlation using multiple data types, e.g. where each dataset comprises multiple different types of measurement data (e.g. ^13C signals and gamma signals), or where some of the datasets contain one type of measurement data (e.g. ^13C signals) and others contain another type of measurement data (e.g. gamma ray signals). In this case, a first single curve 10 is fitted to the composite data of the first type (e.g. ^13C signals) and a second single curve 10 is fitted to the composite data of the second type (e.g. gamma signals), and the deviations associated with each curve 10 can be included in the likelihood in a similar way to that described above. As such, like data types are still correlated with like data types, but their correlation process is combined and interrelated (i.e. not independent) through the likelihood based on the assumption that each data type varies systematically with time (height). In this way, multiple datasets (which may not all contain the same signal type) can be correlated based on multiple data types. In preferred implementations, the above Bayesian inference is performed numerically using Markov chain Monte-Carlo (MCMC) methods. MCMC can be used to perform an intelligent search in a vast high dimensional space so is well suited to inferring the posterior distributions in the present case. Whilst it is understood that any appropriate method for inferring the posterior distribution could be used, for example, MCMC techniques such as Hamiltonian Monte Carlo, slice sampling and hit-and- run, the inventors have found that the Metropolis-Hastings based approaches are particularly effective. Figure 12 shows a MCMC method 900 for obtaining z scale factors and z offsets for correlating multiple geological datasets 100A-100D according to a preferred implementation based on the Metropolis- Hastings approach. The method 900 sits within the step 420a, 420b of aligning the multiple datasets on common z scale. At step 905, initial proposals for the z offsets and the z scale factors for each dataset or subsection are generated in accordance with the methods described above with reference to figure 8 and the above discussion of predefined distributions. In some examples, initial proposals ^ are also generated for sedimentation rate multipliers and gap heights ^ in the same manner. For each unknown parameter, initial proposals ^ are sampled independently from their respective priors. At step 910 a measure of correlation of the multiple datasets 100A-100D resulting from applying the initial proposals ^ is determined by fitting a single curve 10 (e.g. a cubic spline curve) to the composite data of all the datasets 100A-100D and determining a deviation ^ of the data from the curve 10. At step 915 a posterior probability density for the initial proposals ^ is determined based on a prior probability P(^) of the initial proposals and the likelihood of the alignment . At step 920 further proposals ^’ for the z offsets and z scale factors for each dataset 100A-100D or subsection thereof are generated. At step 925 a new measure of correlation of the multiple datasets resulting from applying the further proposals is determined by again fitting a single curve 10 to the composite data of all the datasets 100A-100D and determining a deviation ^ of the data from the curve 10. At step 930 a posterior probability density for the further proposals is determined based on a prior probability of the further proposals and a revised likelihood of the alignment. The revised likelihood is determined, for example, the same manner as described above with reference to step 915. At step 935 an acceptance probability is determined based on the posterior probability density and the posterior probability density of the previous iteration. At step 940 the further proposals are accepted or rejected based on the acceptance probability. In one example, the acceptance probability is calculated as: where ^^^^^^is the posterior probability density of the previously proposed parameter values ^, and ^^^^^′^ is the posterior probability density of the currently proposed values ^’. The posterior probabilitydensities ^^^^^^and ^^^^^′^can be calculated according to:^^^^^^ ൌ ^^^^^^ ൈ ^^^^^^^^^^^|^^^, (20)where ^^^^^^ is the prior probability of ^^, and ^^^^^^^^^^^|^^^ the likelihood of the data given ^^ (the same equation can be used for the currently proposed values ^’). Proposals are accepted if A > 1. Preferably, if A < 1, a random number R can be generated uniformly between 0 and 1, and the proposals are accepted if A > R, else rejected if A < R. This allows more of the parameter space to the explored. Steps 920 to 940 described above are iteratively performed until a stopping condition is met. The various iterations can be referred to as a MCMC chain. Any appropriate stopping condition may be defined. In one example, the stopping condition is a predetermined number of iterations or a predetermined number of accepted proposals. In another example, a stopping condition may be defined that assesses whether the MCMC chain has converged on a set of parameters ^’, preferably with sufficiently high effective sample sizes. This may involve evaluating the variation of the posterior probability densities for the parameters ^’ over the course of the MCMC chain, or the variation in the sampled parameters themselves. Where multiple MCMC chains are run in parallel (see discussion of parallel tempering below), this may involve evaluating where all parallel MCMC chains have converged on similar posterior distributions for the parameters ^’. In step 945, once the stopping condition is met, final posterior probability distributions for the z offsets and z scale factors for each dataset 100A-100B or subsection thereof is generated based on the accepted proposals. At step 950 final z offsets and z scale factors for each dataset 100A-100B or subsection are obtained from the corresponding final posterior probability density distribution. Figure 13 shows an example of a final posterior density distribution for a sedimentation rate determined from the above approach from which a final sedimentation rate can be obtained. For example, the final sedimentation rate can be obtained from the mean or median of the distribution. In preferred examples, at step 905 the initial proposals for the parameters ^ are generated by sampling them from the corresponding prior distributions or custom distributions. Preferably, in step 920 the further proposals ^’ can be generated in one of several different ways in an adaptive proposal approach to enforce a broad search of the parameter space. In one implementation, further proposals ^’ are generated by sampling from the prior or custom distributions for a predefined number of iterations. After the predefined number of iterations the method can enter an adaptive phase whereby the further proposals ^’ are generated by any one of the following proposal types: 1) Proposing from the prior or a custom distribution. 2) Adaptive independent (univariate) proposals: Proposals ^’ for each parameter are selected independently from other parameter values. Proposals are dependent on the current state of the parameter ^^^, and sampled from a normal distribution ^^^^^^,^^^^, where is a standard deviation that is estimated based on the history of the MCMC chain, i.e. based on the sampled ^^^from previous iterations. 3) Adaptive dependent (multivariate) proposals: Proposals for the parameters ^^’ are selected jointly and are dependent on the current state of the parameters ^^’. Proposals are sampled from a multivariate normal distribution ^^^^^^^^^′,^^^, where ^^ is a covariance matrix that is estimated based on the history of the MCMC chain, i.e. based on the sampled proposals ^^′^from previous iterations. 4) Shifting some parameters (e.g. the z offsets) while keeping the other parameters constant. This can accelerate the convergence of the MCMC chain in cases where some datasets 100A- 100D are aligned with each other, but are offset relative to other datasets 100A-100D. Any appropriate combination of proposal types may be used throughout the adaptive phase of the iterations of the MCMC chain. In one example, proposal types are selected dynamically during the adaptive phase of the iterations based on a probability that broadly corresponds to the relative acceptance probability of the respective proposal type, i.e. proposal types that are rejected often are chosen less frequently. Adaptation for types 2) and 3), and the adjustment of proposal type probabilities ends after the adaptation phase i.e. the standard deviation or covariance matrix for the proposals remains fixed after the adaptation has ended. Posterior samples may only be used from after the adaptation has ended, to ensure the correct convergence of the MCMC chain. For type 2) or 3) proposals, the adaptation can be expanded to include the estimation of clusters based on the distribution of sampled from previous iterations. Proposal standard deviations or covariances are then chosen based on which cluster centre is closest to the current state of ^^. Additional measures can be taken to avoid the MCMC chain becoming trapped at isolated peaks of the posterior probability distribution. In some examples, a parallel tempering framework is implemented. This involves running several MCMC chains in parallel. The target MCMC chain, the MCMC chain from which the posterior samples will be taken, is left unaltered. The other MCMC chains are tempered, i.e. their unnormalised log posterior probabilities are raised to the power of 1 / ^^, with ^^ being a predefined temperature. The higher ^^, the more “flattened” the posterior probability landscape becomes, and the easier it is for the MCMC chain to explore the landscape. Frequently, chain swaps are proposed, during which the ^^ of different chains are exchanged with a probability that is higher if the hotter chain has the higher (untempered) posterior probability. As such, in this example, after the step 940 of accepting or rejecting the further proposals, the method 900 further comprises determining a swap probability for each further MCMC chain based on the unnormalized posterior probability of the target MCMC chain and the unnormalized posterior probability of the respective further MCMC chain; and swapping the proposals for the z offsets and the z scale factors in the target MCMC chain with those of a further MCMC chain based on the swap probabilities. The initial temperatures T are selected with a geometric spacing, meaning each temperature T is set at a fixed ratio higher than the previous one, ensuring an exponential increase across the sequence. Specifically, the base ratio for the temperature T increase is calculated to adjust with the total number of MCMC chains. As the number of MCMC chains increases, the ratio decreases, leading to a finer temperature gradient. Temperatures T can be updated in the adaptive phase of the MCMC to increase the swap rates of chains. This can be done by tracking swap rates between pairs of neighbouring MCMC chains when ordered by temperatures. For example, if the swap rate of a first MCMC chain 1 (cold) with a second MCMC chain 2 (medium temperature) is higher than the swap rate of the second MCMC chain 2 with a third MCMC chain 3 (high temperature), the swap rates can be evened by bringing the temperature of the second MCMC chain 2 closer to that of the third MCMC chain 3 The final posterior probability distributions for the final parameters ^ may be determined differently depending on the specific approach for performing Bayesian inference. In a specific example, the Bayesian inference is performed based on a Metropolis-Hastings within Gibbs sampling approach which combines two known independent MCMC approaches. Gibbs sampling was found particularly useful in sampling from an entire probability space in cases (such as this one) where the posterior distribution is difficult to sample from directly. In a specific implementation, the overall measure of correlation ^ of all the composite data in all the datasets 100A-100D is determined (at steps 910 and 925) based on analysis of deviations from individual spline curves fitted to each dataset 100A-100D. In this example, spline coefficients are sampled from a multivariate normal distribution of the form: ^^ ∼ ^^^^^^^^^^^,^^^ (12)where ^^ is given by. ^^^ℎ^^ are cubic B-splines at a set of knots at heights ℎ^, ^^ is the composite stratigraphic signal of all datasets, and ^^ the overall standard deviation of the spline curve 10. The other element needed for sampling from the posterior of ^^ is ^^, given by ^^ ൌ ^^^ ^ ^^^^^ି^, (14)where ^^ is a smoothing parameter, ^^ is a penalty matrix to prevent the spline 10 from overfitting the data, and The standard deviation ^^ can be fixed as where ^^ is the number of datasets 100A-100D, and ^^^is the standard deviation of individual splines fitted to the individual datasets on their unaltered height scales of a given dataset 100A-100D. These individual splines are fitted solely for the purpose of determining an average standard deviation that can be used for calculating the likelihood of the overall spline curve 10 (a single curve 10 is fitted for all other purposes). Alternatively, ^^ can be estimated within the Gibbs sampling scheme from the data, by placing aconjugate Gamma prior on its inverse (precision, ^^ ൌ 1 / ^^ଶ): The smoothing parameter ^^ is estimated by placing a Gamma prior on ^^: Figure 14 illustrates a probabilistic age-height or age-depth model for material at the respective geolocations generated from the final curve 10 fitted to the datasets 100A-100D aligned on the age scale using the final age offsets and final sedimentation rates, following the procedure described above. The age-height model defines an estimated absolute age for any point within the z-range of the aligned datasets 100A-100D, including interpolated points between the measured data points. The resulting relationship for dataset 100C is shown. This also means that any number of correlation lines can be drawn on the age scale between any pairs of datasets. The age-height or age-depth model further includes a measure of uncertainty for each age, obtained from the final posterior distributions of the sedimentation rates and age offsets. The final sedimentation rates ^ and their posterior probability distributions for each dataset or subsection thereof can also be output, with figure 13 showing an example posterior probability distribution for the sedimentation rate of a given subsection. If gaps G are present, a posterior probability distribution of their duration in Myr (or their height in m, if on the height scale) is also obtained. In the following specific example, two datasets associated with Permian-Triassic strata in the UK are correlated on a common height (depth) scale: a first dataset 100A including measurements of chemostratigraphic data (sulphur isotope ratio ^34S) as a function of depth from a drill core from Staithes (Cleveland Basin, Yorkshire, UK), and a second dataset 100B including measurements of ^34S a function of depth from a well in the North Sea Basin. The ^34S data is assumed to reflect the coeval isotopic composition of the sea water. The received (unmodified) first and second datasets 100A, 100B are shown figure 15. The first dataset 100A from Staithes is selected as the reference dataset (as it includes more data points). In this example, a simple sedimentation model is specified (by the user) whereby each dataset 100A, 100B includes a single z offset ^ and z scale factor ^. In this case, depths in the second dataset 100B (at the North Sea Basin site) can be translated to corresponding depths in the reference dataset 100A (at the Staithes site) using equation 1. Appropriate prior distributions for ^, P(^), and ^, P(^), are also specified for the data (by the user). In this example, P(^) is a uniform distribution and P(^) is a log- normal prior distribution (log^). To allow for partial overlap with the reference datasets 100A, ^ is allowed to range from 2000 to 500. It is initially assumed that sedimentation rates in the two datasets 100A, 100B similar, such that the initial relative sedimentation rate is around 1. As such, in this example, P(^) has a mean of log(1) = 0 and a standard deviation of 0.5. The method 900 is performed to estimate the posterior distributions for ^ and ^, in this example with a stopping condition of 1000 iterations. Figure 16(a) shows the (unaltered) reference dataset 100A along with the modified second data 100B resulting from applying the final z offset ^ and z scale factor ^ and the fitted curve 10. The final z offset ^ and z scale factor ^ were obtained from the median of their respective final posterior distributions of the most likely alignment. In this example, the results suggest more than one (4) distinct final alignment solutions. Figure 16(a) is the most likely alignment with a probability of 77%, while figures 16(b) to 16(d) show three other possible alignment solutions with respective probabilities of 17 %, 6 % and 0.3 %. The different discrete possible final alignments are determined by performing clustering analysis on the posterior samples of the parameters ^ and ^ for the second dataset 100B. This is illustrated in figure 17 which shows the posterior samples of the parameters ^ and ^ for the second dataset 100B in a histogram plot from which different discrete clusters A1, A2, A3 can be identified corresponding to the possible alignment solutions. Cluster A1 is the most likely alignment solution, then A2, and then A3 (the fourth possible, least likely, alignment is not visible on this scale). Figure 18 shows the resulting probabilistic depth model for the two datasets 100A, 100B produced from the correlation. The model includes the relationship between the depths in the first (reference) dataset 100A (at the Staithes suite) and the corresponding depths in the second dataset 100B (at the North Sea Basin site) through equation 1 (solid line), and the corresponding uncertainty at each depth position (shaded region). Specific Example 2 In this specific example, three datasets associated with different geographic sites in Morocco and Siberia are correlated on a common age scale. The datasets considered are the datasets 100B-100D shown in figure 7(b) which correspond to measurements of chemostratigraphic data (carbon isotope ratio ^13C) as a function of height at the respective sites. Corresponding lithology data 200B-200D as shown in figure 7(a). In this example, each dataset 100B-100D includes at least one reference date point 300B-300D corresponding to a radiometric date obtained at a given height (not shown). The received geological datasets 100B-100D are segmented into subsections 100B-l to 100D-l based on the corresponding lithology data 200B-200D as shown in figure 7(b). Datasets 100B to 100D correspond to ‘Morocco section 7’, ‘Morocco section 3’ and ‘Siberia’ respectively. The objective of this specific example is to obtain a correlation of the stratigraphic data on an age scale and to obtain probabilistic age-height models. In this example, a sedimentation model is specified (by the user) whereby each subsection can have separate independent sedimentation rate ^s,l. A single offset ^sis used for each dataset, and an additional offset ^g for the gap G in dataset 100D. Appropriate prior distributions for ^, ^, ^g are again specified (by the user) for the data. As with the previous example, P(^) is a uniform distribution and P(^) is a log-normal prior distribution (log^). P(^g) is specified as an exponential distribution. Based on the sampled Initial proposals for ^s, ^s,l, and ^g were sampled from their respective prior, and the datasets 100B-100D were converted to an age scale in accordance with equation 5. The method 900 is performed to estimate the posterior distributions for the alignment parameters ^ (^s, ^s,l, and ^g). In this example, the posterior probability distributions for the final parameters ^ was determined using the Metropolis-Hastings within Gibbs approach run for 5000 iterations with 12 parallel MCMC chains (one cold and 11 tempered). As in the previous example, cluster analysis of the results suggests more than one (2) distinct final alignment solutions, one with probability of 96 % and another with a probability of 3 %. Figure 19 shows the most likely alignment, in which it can be seen that all three modified datasets 100B-100D are plotted and aligned on a common age scale. Based on the final correlation of the most likely solution, probabilistic age-height models were obtained for each dataset 100B-100D as illustrated in figures 20(a) to 20(c). These models include the relationship between height and age (solid line) along with a corresponding age uncertainty for each subsection (shaded region). Figure 21 shows a schematic block diagram of an example system 1000 for correlating quantitative geological datasets 100A-100D obtained at one or more geographic locations. The system 1000 is configured, via appropriate hardware components, software and / or program instructions, to perform the method steps described herein, including providing stratigraphic correlation and further resulting products including a probabilistic age-height or age depth model that may, e.g. be provided to an end user. The system 1000 comprises one or more computing devices 1500 that include one or more processing devices or modules 1502, and one or more memory components 1504, 1506, 1518 containing logic, instructions and / or software that when executed by the one or more processing devices 1502 cause the system 1000 to perform at least some of the operations and steps discussed herein. In one implementation, the system 1000 can be implemented on a single computing device 1500, e.g. a personal computer (PC), a tablet computer, a set-top box (STB), a Personal Digital Assistant (PDA), a cellular telephone, a web appliance, a server, a network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine. In another implementation, the system 1000 can be implemented in a computer network comprising a collection of computing devices 1500 that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods described herein. The computing devices 1500 are preferably in wired or wireless communication without each to facilitate the exchange of instructions, data and information. For example, the computing devices 1500 can be connected, e.g., networked, to each other in a Local Area Network (LAN), a Wide Area Network (WAN), an intranet, an extranet, or the Internet. In one example, the system 1000 can be implemented in a client- server network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Figure 22 shows a schematic block diagram of an example computing device 1500 in more detail. The computing device 1500 comprises one or more processing devices 1502, and at least one of a main memory 1504a (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SD RAM) or Rambus DRAM (RD RAM), etc.), a static memory 1506 (e.g., flash memory, static random access memory (SRAM), etc.), and a secondary memory e.g., a data storage device 1518, which communicate with each other via a bus 1530. The processing device 1502 represents one or more general-purpose processors such as a microprocessor, central processing unit, or the like. More particularly, the processing device 1502 may be a complex instruction set computing (CISC) microprocessor, reduced instruction set computing (RISC) microprocessor, very long instruction word (VLIW) microprocessor, processor implementing other instruction sets, or processors implementing a combination of instruction sets. Processing device 1502 may also be one or more special-purpose processing devices such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), network processor, or the like. Processing device 1502 is configured to execute the processing logic (instructions 1522) for performing at least some of the operations and steps discussed herein. The computing device 1500 may further include a network interface device 1508. The computing device 1500 also may include a display unit 1510 (e.g., a liquid crystal display (LCD)), an alphanumeric input device 1512 (e.g., a keyboard or touchscreen), a cursor control device 1514 (e.g., a mouse or touchscreen), and an audio input and / or output device 1516 (e.g., a speaker or microphone). The data storage device 1518 may include one or more machine-readable storage media (or more specifically one or more non-transitory computer-readable storage media) 1528 on which is stored one or more sets of instructions 1522 embodying any one or more of the methodologies, steps or functions described herein. The instructions 1522 may also reside, completely or at least partially, within the main memory 1504 and / or within the processing device 1502 during execution thereof by the computer system 1500. The main memory 1504 and / or the processing device 1502 also constituting computer-readable storage media. The various methods described above may be implemented by a computer program. The computer program may include computer code arranged to instruct one or more computing devices 1500 to perform the functions of one or more of the various methods described above. The computer program and / or the code for performing such methods may be provided to a computing device, on one or more computer readable media or, more generally, a computer program product. The computer readable media may be transitory or non-transitory. The one or more computer readable media could be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, or a propagation medium for data transmission, for example for downloading the code over the Internet. Alternatively, the one or more computer readable media could take the form of one or more physical computer readable media such as semiconductor or solid-state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a rigid magnetic disc, and an optical disk, such as a CD-ROM, CD-R / W or DVD. In an implementation, the modules, components and other features described herein can be implemented as discrete components or integrated in the functionality of hardware components such as ASICS, FPGAs, DSPs or similar devices. A "hardware component" is a tangible (e.g., non-transitory) physical component (e.g., a set of one or more processors, memory etc.) capable of performing certain operations and may be configured or arranged in a certain physical manner. A hardware component may include dedicated circuitry or logic that is permanently configured to perform certain operations. A hardware component may be or include a special-purpose processor, such as a field programmable gate array (FPGA) or an ASIC. A hardware component may also include programmable logic or circuitry that is temporarily configured by software to perform certain operations. Accordingly, the phrase "hardware component" should be understood to encompass a tangible entity that may be physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein. In addition, the modules and components can be implemented as firmware or functional circuitry within hardware devices. Further, the modules and components can be implemented in any combination of hardware devices and software components, or only in software (e.g., code stored or otherwise embodied in a machine-readable medium or in a transmission medium). Although physical hardware components have been described, it would be understood that the above methods could be carried out in a cloud-based environment and / or on virtualized hardware. As such, this phrase "hardware implementation" used in this section encompasses both physical hardware or virtualized hardware in a singular function or shared computing environment, which may be located in dedicated equipment space or hosted equipment in either private or public cloud. Unless specifically stated otherwise, as apparent from the above discussion, it is appreciated that throughout the description, discussions utilizing terms such as "receiving", "determining", "comparing", "correlating", "maintaining," "identifying", "providing", "applying", "generating" or the like, refer to the actions and processes of a computing system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computing system's registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission or display devices. It will be understood that the present disclosure has been described above purely by way of example, and modifications of detail can be made within the scope of the disclosure. In particular, although the method is described in the context of correlating measurements of geological property on both a common height / depth and an absolute age scale, it will be appreciated that the methods disclosed herein could be applied to any appropriate data that requires correlating and is not limited to geological data. Each feature disclosed in the description, and (where appropriate) the claims and drawings may be provided independently or in any appropriate combination. Although the appended claims are directed to particular combinations of features, it should be understood that the scope of the disclosure also includes any novel feature or any novel combination of features disclosed herein either explicitly or implicitly or any generalisation thereof, whether or not it relates to the same invention as presently claimed in any claim and whether or not it mitigates any or all of the same technical problems as does the present disclosure. Reference numerals appearing in the claims are by way of illustration only and shall have no limiting effect on the scope of the claims.

Claims

CLAIMS 1. A computer-implemented method for correlating multiple quantitative geological datasets obtained at one or more geographic locations, the method comprising: receiving multiple datasets including a first dataset and one or more second datasets, each dataset comprising a plurality of data points corresponding to measurements of a geological property of material obtained at different z coordinates at a respective geographic location; and aligning the multiple datasets on a common z scale by: determining a measure of correlation of the multiple datasets by fitting a single curve to the data points of the multiple datasets and calculating a deviation of the data points of the multiple datasets from the fitted curve; determining one or more separate z scale factors and z offsets to apply to each of the one or more second datasets based at least in part on the measure of correlation of the multiple datasets and corresponding predefined probability distributions for the z scale factors and z offsets for each second dataset; and modifying the one or more second datasets by applying each z scale factor and z offset to the corresponding second dataset.

2. The method of claim 1, wherein aligning the multiple datasets on a common z scale further comprises: partitioning each of the one or more second datasets into a plurality of contiguous subsections, determining a separate z scale factor to apply to each subsection of each second dataset based at least in part on a measure of correlation of the multiple datasets and corresponding predefined probability distributions for the z scale factors for each subsection of the second dataset; and modifying the one or more second datasets by applying each determined z scale factor to the corresponding subsection of the corresponding second dataset.

3. The method of claim 1, wherein aligning the multiple datasets on a common z scale comprises determining a z scale factor and z offset to apply to all the z-coordinates of the corresponding second dataset.

4. The method of claim 1, 2 or 3, wherein the z-coordinates of the datasets correspond to a measurement height or depth relative to a reference point at the respective geographic locations, each z offset corresponds to a height or depth offset, and each z scale factor corresponds to a relativesedimentation rate with respect to the first dataset, and wherein the modified one or more second datasets are substantially aligned to the first dataset on a common height or depth scale.

5. The method of claim 1, wherein at least one of the multiple datasets further include at least one reference age data point corresponding to a geological age of material at a given z coordinate at the respective geographic location, wherein the z-coordinates of the datasets correspond to a measurement height or depth relative to a reference point at the respective geological location, wherein the z offsets correspond to an age offset, and the z scale factors correspond to an absolute sedimentation rate, and wherein aligning the multiple datasets on a common z scale comprises aligning the multiple datasets on a common age scale by: converting the multiple datasets from a height or depth scale to an age scale based on initial sedimentation rates and initial age offsets for the respective first and second datasets; and determining one or more separate final sedimentation rates and final age offsets to apply to each of the first and second datasets based at least in part on a measure of correlation of the converted multiple datasets, corresponding predefined probability distributions for the sedimentation rates and age offsets for each of the first and second datasets, and the at least one reference age data point; modifying the first dataset by applying the corresponding final sedimentation rates and age offsets; and modifying the one or more second datasets by applying the corresponding final sedimentation rates and age offsets.

6. The method of claim 5, wherein aligning the multiple datasets on a common age scale further comprises: partitioning the first dataset and the one or more second datasets into a plurality of contiguous subsections; converting the multiple datasets from a height or depth scale to an age scale based on initial sedimentation rates and initial age offsets for each subsection of each dataset; determining a separate final sedimentation rate to apply to each subsection of each of the first and second datasets based at least in part on a measure of correlation of the converted multiple datasets, corresponding predefined probability distributions for the sedimentation rates for the subsections of each of the first and second datasets, and the at least one reference age data point; modifying the first dataset by applying the corresponding final sedimentation rates to the subsections of the first dataset; andmodifying the one or more second datasets by applying the corresponding final sedimentation rates to the subsections of the one or more second datasets.

7. The method of claim 5, wherein determining a separate final sedimentation rate and final age offset to apply to each of the first and second datasets comprises determining a final sedimentation rate and final age offset to apply to all the z-coordinates of the corresponding dataset.

8. The method of claim 2, 4 or 6, wherein the partitioning is based at least in part on lithology information associated with the respective dataset, and preferably wherein at least some partitions between subsections are associated with a change in lithology type indicated by the corresponding lithology information.

9. The method of claim 8, wherein the lithology information is used as a constraint on the alignment of the multiple datasets such that subsections associated with the same lithology type within a given dataset or across all datasets share the same z scale factor.

10. The method of claim 9, where subsections associated with the same lithology type across all datasets share the same z scale factor, aligning the multiple datasets further comprises: determining, for at least one of the multiple datasets, a subsection-agnostic multiplier to be applied to the z scale factors of all subsections of the respective dataset, based at least in part on the measure of correlation; and modifying the at least one of the multiple dataset by applying the determined z scale factors, z offsets and multiplier for the corresponding dataset.

11. The method of any preceding claim, wherein the one or more second datasets include a plurality of second datasets.

12. The method of any preceding claim, wherein determining the z scale factors and z offsets comprises performing Bayesian inference.

13. The method of claim 12 when dependent on any of claims 1 to 4 or 9 to 11, wherein the predefined probability distributions for the z scale factors and z offsets are respective predefined prior distributions for the z scale factors and z offsets, and wherein performing a Bayesian inference comprises: determining the measure of correlation of the multiple datasets; defining a likelihood for the correlation of the multiple datasets based at least in part on the measure of correlation; performing Bayesian inference to determine posterior probability distributions for the zoffsets and the z scale factors for each second dataset or each subsection thereof, based on the likelihood and the corresponding predefined prior distributions; and determining the z-offsets and z scale factors to be applied to each second dataset or each subsection thereof based on the corresponding posterior probability distributions.

14. The method of claim 12 when dependent from any of claims 5 to 11, wherein the predefined probability distributions for the sedimentation rates and age offsets are respective predefined prior distributions for the sedimentation rates and age offsets, and wherein performing the inference comprises: determining a measure of correlation of the converted multiple datasets; defining a likelihood for the correlation of the converted multiple datasets based on the measure of correlation of the converted multiple datasets and the at least one reference age data point; performing Bayesian inference to determine posterior probability distributions for the final sedimentation rates and final age offsets for each dataset or subsection thereof, based on the likelihood and the corresponding predefined prior distributions; and obtaining the final z-offsets and z scale factors to be applied to each dataset or each subsection thereof from the corresponding posterior probability distributions.

15. The method of claim 14, wherein the at least one reference age data point is used as a constraint placed on the likelihood to induce the Bayesian inference to produce final sedimentation rates and age offsets that tie specific data points in the aligned multiple datasets to the geological age associated with the at least one reference age data point.

16. The method of any of claims 12 to 15, wherein the Bayesian inference is performed using a Markov-chain Monte-Carlo (MCMC) method.

17. The method of claim 16, wherein the MCMC method comprises performing a MCMC chain including: performing an initial iteration of: generating initial proposals for the z offsets and the z scale factors for each dataset or subsection thereof determining a measure of correlation of the multiple datasets resulting from applying the initial proposals; determining a posterior probability density for the initial proposals based on a prior probability of the initial proposals and the likelihood of the correlation; and performing one or more further iterations of: generating further proposals for the z offsets and z scale factors for eachdataset or subsection thereof; determining a new measure of correlation of the multiple datasets resulting from applying the further proposals; determining a posterior probability density for the further proposals based on a prior probability of the further proposals and a revised likelihood of the correlation, wherein the revised likelihood is based on the new measure of correlation; generating posterior probability distributions for the final z offsets and z scale factors for each dataset or subsection thereof based on the proposals; and obtaining the final z offsets and z scale factors for each dataset or subsection from the corresponding posterior probability distributions.

18. The method of claim 17, wherein the MCMC method further comprises, in each further iteration: determining an acceptance probability based on the posterior probability density and the posterior probability density of the previous iteration; and accepting or rejecting the further proposals based on the acceptance probability, and wherein the posterior probability distributions for the final z offsets and z scale factors for each dataset or subsection thereof is generated based on the accepted proposals.

19. The method of claim 17 or 18, wherein the initial proposals for the z offsets and the z scale factors are generated by sampling from a corresponding initial distribution, and preferably wherein the corresponding initial distribution is the corresponding prior probability distribution.

20. The method of claim 17, 18 or 19, wherein the further proposals for the z offsets and the z scale factors are generated by one or more of: (i) sampling from the corresponding prior probability distribution; (ii) sampling from a corresponding normal distribution defined with respect to the respective z offset and the z scale factor proposed in the previous iteration with a standard deviation based on the respective z offsets and the z scale factors proposed in a plurality of the previous iterations; (iii) sampling from a corresponding multivariate normal distribution defined with respect to the respective z offset and the z scale factor proposed in the previous iteration with a covariance matrix based on the respective z offsets and the z scale factors proposed in a plurality of the previous iterations; (iv) sampling from one of a plurality of a corresponding multivariate normal distributions which is centred closest to the z offset and the z scale factor proposed in the previous iteration, wherein the plurality of multivariate normal distributions are defined based on a cluster analysis of the z offsets and the z scale factors proposed in a plurality of previous iterations; and(v) sampling a further proposal for one of the z offset and the z scale factor according to any of (i) to (iv), while keeping the other of the z offset and the z scale factor constant.

21. The method of any of claims 17 to 20, wherein the MCMC chain is a target chain, and the method further comprises performing one or more further MCMC chains in parallel to the target chain, wherein the posterior probability densities for each further chain are modified by applying a predefined temperature factor, and wherein the method comprises, after the step of accepting or rejecting the further proposals: determining a swap probability for each further MCMC chain based on the posterior probability density of the target MCMC chain and the posterior probability density of the respective further MCMC chain; and swapping the proposals for the z offsets and the z scale factors generated in the target MCMC chain with those of a further MCMC chain based on the swap probabilities.

22. The method of any of claims 16 to 21, wherein the MCMC method is performed according to a Metropolis-Hastings or Metropolis-within-Gibbs approach.

23. The method any of claims 16 to 22, wherein the posterior distributions for the final z offsets and the z scale factors comprise multiple peaks associated with different possible alignment solutions, and the method comprises: determining a most likely peak in the final posterior distribution for the final z offsets and the z scale factors based on a cluster analysis of the proposals generated over a plurality of iterations of the MCMC; and obtaining the final z offsets and z scale factors for each dataset or subsection from the most likely peaks in the corresponding posterior probability distributions.

24. The method of any preceding claim, wherein the single curve used to determine the measurement of alignment is a cubic spline curve.

25. The method of any preceding claim, wherein the geological property is selected from a group containing: a geochemical property, a geophysical property, a palaeontology property, and a palaeobiological property, a biomolecular data, and a spectroscopic data.

26. The method of any of claims 5 to 7, or any claim directly or indirectly dependent on claims 5 to 7, further comprising: generating a probabilistic age-height or age-depth model for material at the respective geolocations based at least in part on the aligned multiple datasets and the final age offsets and sedimentation rates, wherein the age-height model defines an age for any point within the z-range ofthe aligned datasets, and preferably wherein the age-height or age-depth model further defines a measure of uncertainty for each age, the uncertainty obtained from the final posterior distributions of the sedimentation rates and age offsets.

27. A system for correlating quantitative geological datasets obtained at one or more geographic locations, the system comprising one or more processors and memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform the method of any of claims 1 to 26.

28. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any of claims 1

Citation Information

Patent Citations

  • Stochastic Dynamic Time Warping for Automated Stratigraphic Correlation

    US20220011457A1