Systems and methods for supervised learning of permeability of a formation
Patent Information
- Application Number
- CN202080033019.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-11
- Filing Date
- 2020-03-09
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2040-03-09
AI Technical Summary
虽然已经开发了多种利用不同测井和岩心测量来确定渗透率的方法,但围绕数据异方差的问题仍然是一个挑战
Smart Images

Figure CN114207550B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of U.S. Provisional Application No. 62 / 815,714, filed March 8, 2019, and U.S. Provisional Application No. 62 / 816,566, filed March 11, 2019, the contents of which are incorporated herein by reference. Technical Field
[0003] This disclosure relates to supervised learning of formation permeability, and more specifically, to systems and methods for supervised learning of formation permeability. Background Technology
[0004] In reservoir evaluation, permeability is a crucial and essential petrophysical parameter for geological formations (also known as formations). Specifically, accurate permeability determination is vital for understanding the value of a reservoir. Furthermore, it is an important input in reservoir models for predicting oil and gas production. Although various methods have been developed to determine permeability using different logging and core measurements, the issue of data heteroscedasticity remains a challenge. Summary of the Invention
[0005] This summary is provided to introduce some concepts that will be further described in the detailed description below. This summary is not intended to identify the essential features of the claimed subject matter, nor is it intended to help limit the scope of the claimed subject matter.
[0006] In embodiments of this disclosure, a method for characterizing rock strata samples is provided. The method for characterizing rock strata samples may include obtaining multiple datasets characterizing the rock strata samples. The method for characterizing rock strata samples may also include training a neural network to generate a computational model. Furthermore, the method for characterizing rock strata samples may additionally include using multiple datasets as input to the computational model, wherein the computational model may be implemented by a processor that derives an estimate of the permeability of the rock strata samples.
[0007] The computational model may include one or more of the following features. The computational model may be based on a trained artificial neural network. The computational model may further derive values representing the uncertainty associated with the estimate of the permeability of the rock strata sample. The computational model may be based on a trained Bayesian neural network and / or employ an artificial neural network using discarded Bayesian inference. Multiple datasets may include data derived from nuclear magnetic resonance (NMR) measurements of the rock strata sample. Multiple datasets may include T2 feature data. T2 feature data can be derived by encoding the T2 distribution of the rock sample using a kernel based on singular value decomposition (“SVD”), and then mapping the T2 distribution data to T2 features in a reduced-dimensional space. Multiple datasets may include mineralogical data corresponding to the rock strata sample. The rock strata sample may be selected from rock cuttings, cores, drill cuttings, rock outcrops, or rock strata and coal surrounding boreholes.
[0008] In another embodiment of this disclosure, a system for characterizing rock strata samples is provided. The system for characterizing rock strata may include a memory storing multiple datasets characterizing the rock strata samples. The system for characterizing rock strata may further include a processor configured to train a neural network to generate a computational model, wherein the multiple datasets are input into the computational model, and wherein the computational model is implemented by the processor, which derives an estimate of the permeability of the rock strata samples.
[0009] It may include one or more of the following features. The computational model may be based on at least one of an artificial neural network or a Bayesian neural network trained on it. The computational model may further derive values representing the uncertainty associated with the estimate of the permeability of the rock formation sample. The computational model may be based on an artificial neural network trained using Bayesian inference with discarding.
[0010] In another embodiment of this disclosure, a method for supervised learning of rock physical parameters of a formation is provided. The method may include obtaining multiple datasets representing samples. The method may further include providing a neural network with one or more discarded parameters. Low-fidelity datasets associated with the multiple datasets can be used to train the computational model.
[0011] It may include one or more of the following features. The computational model can be fine-tuned using a high-fidelity dataset. The neural network may be a Bayesian neural network. The first autoencoder can be trained using a low-fidelity dataset. The second autoencoder can be trained using a high-fidelity dataset. At least one parameter associated with the first or second autoencoder may be frozen. Attached Figure Description
[0012] The invention is illustrated in the accompanying drawings by way of example rather than limitation, wherein the same reference numerals denote similar elements, wherein:
[0013] Figure 1 These are diagrams depicting embodiments of the system according to this disclosure;
[0014] Figure 2 It is a graph representing the permeability of rock samples according to this disclosure;
[0015] Figure 3 This is a schematic diagram depicting an embodiment of the method according to the present disclosure;
[0016] Figure 4 This is a block diagram depicting an embodiment of the system according to the present disclosure;
[0017] Figure 5 These are diagrams depicting embodiments of the system according to this disclosure;
[0018] Figure 6 These are diagrams depicting embodiments of the system according to this disclosure;
[0019] Figure 7 These are diagrams depicting embodiments of the system according to this disclosure;
[0020] Figures 8A-8D It is a graph representing the results of the supervised learning process according to this disclosure;
[0021] Figures 9A-9B It is a graph representing the results of the supervised learning process according to this disclosure;
[0022] Figure 10 These are diagrams depicting embodiments of the system according to this disclosure. Detailed Implementation
[0023] The following discussion pertains to certain implementations and / or embodiments. It should be understood that the following discussion can be used to enable those skilled in the art to make and use any subject matter now or hereafter defined by the patent "claims" in any of the patents disclosed herein.
[0024] Specifically, the claimed combination of features is not limited to the embodiments and illustrations contained herein, but includes modifications of those embodiments that are included within the scope of the appended claims and combinations of elements of different embodiments. It should be understood that in the development of any such actual embodiment, as in any engineering or design project, many implementation-specific decisions can be made to achieve the developer's specific goals, such as complying with system-related and business-related constraints that may vary from embodiment to embodiment. Furthermore, it should be understood that such development efforts may be complex and time-consuming, but remain routine tasks of design, manufacture, and production for those skilled in the art who benefit from this disclosure. Unless explicitly stated as "critical" or "essential," nothing in this application is considered critical or essential to the claimed invention.
[0025] It should also be understood that although the terms first, second, etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms can be used to distinguish individual elements. For example, a first object or step may be referred to as a second object or step, and similarly, a second object or step may be referred to as a first object or step, without departing from the scope of this disclosure. The first object or step and the second object or step are two separate objects or steps, but they should not be regarded as the same object or step.
[0026] refer to Figure 1 This paper illustrates a method for supervised learning of rock physical parameters and a method for characterizing rock strata samples, both referred to herein as a supervised learning process 10. For the purposes of the following discussion, it should be understood that the supervised learning process 10 can be implemented in various ways. For example, the supervised learning process 10 can be implemented as a server-side process, a client-side process, or a server-side / client-side process.
[0027] For example, supervised learning process 10 can be implemented as a purely server-side process via supervised learning process 10s. Alternatively, supervised learning process 10 can be implemented as a purely client-side process via one or more of client applications 10c1, 10c2, 10c3, and 10c4. Still alternatively, supervised learning process 10 can be implemented as a server-side / client-side process via server-side supervised learning process 10s combined with one or more of client applications 10c1, 10c2, 10c3, 10c4, and 10c5. In such an example, at least a portion of the functionality of supervised learning process 10 can be performed by supervised learning process 10s, and at least a portion of the functionality of supervised learning process 10 can be performed by one or more of client applications 10c1, 10c2, 10c3, 10c4, and 10c5.
[0028] Therefore, the supervised learning process 10 used in this disclosure may include any combination of supervised learning process 10s, client application 10c1, client application 10c2, client application 10c3, client application 10c4 and client application 10c5.
[0029] The supervised learning process 10s can be a server application and can reside on and be executed by computing device 12, which can be connected to network 14 (e.g., the Internet or a local area network). Examples of computing device 12 may include, but are not limited to: personal computer, server computer, a series of server computers, minicomputer, mainframe computer, or dedicated network device.
[0030] The instruction set and subroutines of the supervised learning process 10s, which can be stored on a storage device 16 coupled to the computing device 12, can be executed by one or more processors (not shown) and one or more memory architectures (not shown) included within the computing device 12. Examples of storage device 16 may include, but are not limited to: hard disk drives; tape drives; optical drives; RAID devices; NAS devices; storage area networks; random access memory (RAM); read-only memory (ROM); and all forms of flash storage devices.
[0031] Network 14 can be connected to one or more secondary networks (e.g., network 18), examples of which may include, but are not limited to, local area networks (LANs), wide area networks (WANs), or intranets.
[0032] The instruction sets and subroutines of client applications 10c1, 10c2, 10c3, 10c4, and 10c5 can be stored (respectively) on storage devices 20, 22, 24, 26, and 28 (respectively) coupled to client electronic devices 30, 32, 34, 36, and 38, and can be executed by one or more processors (not shown) and one or more memory architectures (not shown) coupled to client electronic devices 30, 32, 34, 36, and 38. Examples of storage devices 20, 22, 24, 26, and 28 may include, but are not limited to: hard disk drives; tape drives; optical drives; RAID devices; random access memory (RAM); read-only memory (ROM); and all forms of flash memory storage devices.
[0033] Examples of client electronic devices 30, 32, 34, 36, and 38 may include, but are not limited to, personal computers 30 and 36, laptop computers 32, mobile computing devices 34, notebook computers 36, netbook computers (not shown), server computers (not shown), game consoles (not shown), data-enabled television consoles (not shown), and dedicated network devices (not shown). Each client electronic device 30, 32, 34, 36, and 38 may run an operating system.
[0034] Users 40, 42, 44, 46, and 48 can access the supervised learning process 10 directly through network 14 or through secondary network 18. Additionally, the supervised learning process 10 can be accessed via link 50 through secondary network 18.
[0035] Various client electronic devices (e.g., client electronic devices 30, 32, 34) can be directly or indirectly coupled to network 14 (or network 18). For example, personal computer 28 is shown as being directly coupled to network 14. Furthermore, laptop computer 30 is shown as being wirelessly coupled to network 14 via a wireless communication channel 52 established between laptop computer 30 and wireless access point (WAP) 54. Similarly, mobile computing device 32 is shown as being wirelessly coupled to network 14 via a wireless communication channel 56 established between mobile computing device 32 and cellular network / bridge 58, which is shown as being directly coupled to network 14. For example, WAP 48 can be an IEEE 802.11a, 802.11b, 802.11g, 802.11n, Wi-Fi, and / or Bluetooth device capable of establishing a wireless communication channel 52 between laptop computer 30 and WAP 54. Additionally, personal computer 34 is shown as being directly coupled to network 18 via a hardwired network connection.
[0036] Generally as described above, some and / or all of the functionality of the supervised learning process 10 may be provided by one or more client applications 10c1-10c5. For example, in some embodiments, the supervised learning process 10 (and / or the client functionality of the supervised learning process 10) may be included within and / or interact with client applications 10c1-10c5, which may include client e-applications, web browsers, or other applications. Various additional / alternative configurations may be utilized equivalently.
[0037] Various models have been developed over the years for determining the permeability of geological rocks. Specifically, known models involve measurements using well logging and core data to determine permeability. For example, models such as the KSDR and Timur-Coates models may be included. Estimates of permeability from rock samples based on earlier studies of sandstone are called... This can be derived empirically, as shown in Equation 1 below:
[0038]
[0039] Equation 1
[0040] In Equation 1, φ is the porosity of the rock sample. It is a rock sample The logarithmic mean of the distribution, where A is a formation-related scalar factor. Parameters b and c can be determined empirically through core measurement calibration. Furthermore, the estimated permeability of the rock sample is called... This can be derived from the surface relaxation rate of the rock sample, as shown in Equation 1 below:
[0041]
[0042] Equation 2
[0043] In Equation 2, the relaxation rate It can be estimated by comparing diffusion relaxation (D-T2) plots or NMR and micromolecular pulse-poured (MICP) data. Unfortunately, estimating the relaxation rate using any of the methods mentioned above is time-consuming and challenging. Estimating the surface relaxation rate from a mineralogical perspective is also challenging because even if there is a potential relationship between the relaxation rate and the concentration of paramagnetic elements (Fe and Mn), the exact functional form of this relationship remains unknown.
[0044] The Timur-Coates model for the permeability of rock samples is called... As shown in equation 3 below:
[0045]
[0046] Equation 3
[0047] In Equation 3, B is a scalar factor, and FFV and BFV represent the free fluid volume and bound fluid volume of the rock sample, respectively. FFV and BFV can be derived from the rock sample... The data were obtained from a distribution with appropriate mineralogically relevant cutoff values. Furthermore, the parameters in the aforementioned model (such as A and B) can be calculated through laboratory tests, where the actual permeability of the ground rock was measured using helium or nitrogen flow rate measurements on core samples, and the results are compared with measurements from the core samples. Distribution correlation. Different studies have also shown a strong correlation between porosity and permeability in different rock types. Figure 2 This concept is explained. Figure 2 This demonstrates a strong permeability-porosity correlation across different rock types or mineralogical studies. Specifically, Figure 2The graph shows permeability versus porosity from four datasets of sand and sandstone, illustrating the decrease in permeability and porosity as the pore size decreases with changes in minerals.
[0048] Furthermore, deep learning (DL)-based regression methods can be used to capture with high accuracy the highly complex underlying nonlinearities present in data collected on the permeability of geological rocks. However, DL-based regression methods present challenges, particularly for large-scale applications in industrial settings. For example, one challenge might be that some (if not most) algorithms cannot handle the highly heteroscedastic nature of data common in industrial applications, particularly regarding noise fidelity. Existing solutions focus on processing low-fidelity and high-fidelity datasets based on the noise in the output labels and combining them into a more accurate model. To achieve better accuracy, the model can be trained on low-fidelity data and then fine-tuned on high-fidelity data. Incorporating noisy output labels into the machine learning workflow is another option. However, introducing input noise appears unresolved when jittering the training data with known uncertainties.
[0049] Another example of the challenges posed by using deep learning-based regression methods includes the potential lack of Value of Information (VOI) aspects when making predictions using conventional algorithms. They typically provide point estimates but offer no indication of the associated predictive uncertainty. Two distinct deep learning approaches have been proposed, for instance, to incorporate uncertainty. One approach involves incorporating one or more uncertainties via Bayesian inference. The basic idea is to assign a mean and variance to each weight parameter in the neural network and then perform backpropagation based on variational inference, rather than the usual sampling-based approach. This allows computation to be done on GPUs, significantly increasing its speed. The second variational approach uses the concept of "dropout." Traditionally, dropout has been used by models to prevent overfitting. They work by randomly shutting down neurons during the training phase of the neural network based on a certain Bernoulli parameter. This forces neurons to be independent of each other and forces them to learn unique, distinguishable features. Dropout segments remain active during training but are inactive during testing. During testing, their outputs are essentially multiplied by a factor dependent on the differential Bernoulli parameter to maintain the expected activation as likely as it might have been during training. By maintaining uncertain activations during testing, dropout can be further used to estimate uncertainty. Therefore, such a network can be approximated as a deep Gaussian process. However, a significant drawback is that the choice of the Bernoulli parameter for each dropout segment is manually set, i.e., it is a hyperparameter. Optimizing the hyperparameter requires grid search, which can be computationally prohibitively expensive for complex neural networks. The Concrete Dropout segment essentially replaces the discrete Bernoulli distribution of the dropout with its continuous relaxation (i.e., concrete distribution relaxation). This relaxation allows for the reparameterization of the distribution and, in essence, makes the parameters learnable. When combined with heteroscedasticity loss (instead of the usual MSE), architectures with Concrete Dropout segments can produce well-calibrated predictive uncertainty. Furthermore, all machine learning techniques only work if the test dataset is within the distribution of the training dataset. When this assumption is violated, and the input test data points do not conform to the distribution (OOD), conventional methods fail silently without issuing an alarm that the model may be failing.
[0050] Now for reference Figure 3A flowchart illustrating an example of a supervised learning process 10 is provided. In some embodiments, the supervised learning process 10 may include obtaining 302 multiple datasets characterizing rock formation samples. The supervised learning process 10 may further include training 304 a neural network to generate a computational model. Furthermore, the supervised learning process 10 may include using 306 multiple datasets as input to the computational model, wherein the computational model may be implemented by a processor that derives an estimate of the permeability of the rock formation samples. The supervised learning process 10 may consider the heteroscedasticity of the data (i.e., the level of variation in noise fidelity across one or more samples). In addition to point estimation, the supervised learning process 10 may also provide one or more confidence intervals for the predicted measurement. The supervised learning process 10 may include a calibrated metric for checking one or more uncertainties.
[0051] In some embodiments, the supervised learning process 10 may use one or more calibrated uncertainties to classify test data points as either within or ODD relative to training data points, and may also provide a metric for measuring the performance of the prediction uncertainty as an indicator of the ODD data. Generally, the supervised learning process 10 may give equal importance to all input data points. For example, in oilfield applications, different data points may come from different legacy tools and have different degrees of uncertainty. Furthermore, the supervised learning process 10 typically provides a point estimate of the desired output by minimizing the l2 norm between one or more predictions and ground reality. The supervised learning process 10 can provide one or more point estimates of the predicted measurement along with one or more confidence intervals, while balancing the precision and accuracy of the output. Additionally, the supervised learning process 10 may include one or more metrics that can measure the calibration of one or more confidence intervals. Furthermore, the supervised learning process 10 may use one or more calibrated uncertainties to classify test data points as either within or ODD relative to training data points, and may also provide a metric for measuring the performance of the prediction uncertainty as an indicator of the ODD data. The supervised learning process 10 can ensure the following three points: (1) respecting the noise information present in the input and output of the model in the training dataset; (2) providing one or more confidence intervals for each predicted output on the test dataset and performing quality checks on one or more uncertainties obtained; and (3) effectively handling imbalanced multifidelity datasets.
[0052] In some embodiments, the supervised learning process 10 may present a novel probabilistic programming pipeline, also known as a computational workflow and computational model, for determining the petrophysical parameters (i.e., permeability) of geological strata. The supervised learning process 10 may use NMR relaxation (T2) data measured from geological strata samples, along with corresponding mineralogical data. The pipeline can be flexible to modify the feature space based on other measurements. In some embodiments, the supervised learning process 10 may use machine learning within the programming pipeline to learn a set of input features x (i.e., permeability). The computational model (i.e., the mapping of f) between the real-valued output y (i.e., permeability) and the actual output y (i.e., permeability) minimizes the mean squared error (MSE), as shown in Equation 4 below:
[0053]
[0054] Equation 4
[0055] In some embodiments, the supervised learning process 10 may include a preprocessing stage for preparing a set of input features. For example, the preprocessing stage may encode the T2 distribution of kernel-paired samples using Singular Value Decomposition (SVD), and then map the T2 distribution data encoded in a higher-dimensional space (i.e., 64-dimensional) to T2 features in a reduced-dimensional space (i.e., 6-dimensional). The general workflow of the programming pipeline is as follows: Figure 4 As shown. First, the T2 distribution 402 is encoded using an SVD-based encoder 404. Then, the NMR features of sample 406 (i.e., porosity, bound-free volume ratio, and the T2 log-mean derived from the NMR measurements of the sample), the T2 features 408 in the reduced-dimensional space (i.e., as the output of the preprocessing stage), and the mineralogical data 410 corresponding to the rock sample (i.e., the concentration of a set of mineral components commonly found in geological rock samples) are combined and input into a machine learning model 412 to predict the permeability 414 of the rock sample. The quality of the computational model can be determined through cross-validation.
[0056] In some embodiments, NMR measurements can be performed on rock cuttings, cores, drill cuttings, or other samples from geological formations using an NMR spectrometer in a surface laboratory or at a surface well site. Alternatively, NMR measurements can be performed using NMR logging tools on portions or samples of geological formations surrounding the borehole as part of wellbore logging measurements.
[0057] In some embodiments, mineralogical data can be determined from X-ray diffraction or infrared spectroscopy measurements of rock samples performed in a surface laboratory or at a surface well site. Such spectroscopic measurements can be performed on rock cuttings, cores, drill cuttings, or other samples from geological formations using an NMR spectrometer in a surface laboratory or at a surface well site. These methods can be considered “direct measurements” of mineralogical data because each produces a spectrum containing information about mineral properties (i.e., the location of the spectral signal, typically plotted on the horizontal axis) and about the concentration of mineral components (the intensity of the spectral signal, typically plotted on the vertical axis). However, these methods are generally not available downhole within the wellbore.
[0058] In some embodiments, methods can be employed to determine the mineralogical composition of a portion or sample of one or more geological formations surrounding the borehole based on well logging measurements. However, these methods are more challenging because the determination may rely on “indirect measurements.” Well logging measurements commonly used to infer mineralogical composition are induced neutron gamma-ray spectroscopy / spectroscopy. In short, fast neutrons and thermal neutrons, generated by naturally occurring radioactive material or a pulsed neutron generator contained within the well logging probe housing, interact with elemental nuclei in the formation, producing transient (i.e., inelastic) and trapped gamma rays, respectively, in a local volume surrounding the well logging probe. These generated gamma rays penetrate the formation, and a portion of them are detected by a scintillation detector contained within the well logging probe housing. The spectrum produced by the detector signal contains contributions from gamma rays representing elemental nuclei in the formation. This spectrum is typically plotted as a count rate (vertical axis) versus energy (horizontal axis) and includes information about elemental identity (based on the characteristic energies of the gamma rays) and information about the elemental abundance (based on the number of counts) of certain elements commonly found in reservoir rocks (e.g., Si, Al, Ca, Mg, K, Fe, S, etc.). This measurement of elemental abundance does not provide a direct mineralogical measurement because the same finite number of common rock-forming elements are contained in far greater numbers of common rock-forming minerals. An example is the element silicon (Si), which is found in common sedimentary minerals such as quartz (SiO2), opal (SiO2·nH2O), potassium feldspar (KAlSi3O8), plagioclase ([Na,Ca]Al[Al,Si]Si2O8), and many other silicate minerals, including clay and mica group minerals. Nevertheless, there must be a relationship between the concentration of macroelements in a rock sample and its mineral composition concentration, because minerals have a finite range of chemical compositions due to their defined crystal structures. Therefore, the concentration of macroelements in a rock sample is determined by the mineralogical characteristics of the rock sample.
[0059] In some embodiments, methods can be employed to derive mineralogical estimates from measurements of a large number of elemental concentrations. These methods may rely on the derivation of one or more mapping functions to positively model the prediction of mineral composition concentrations based on elemental concentrations. One set of methods is a linear regression model based on an empirical linear relationship between the concentrations of one or more elements and one or more minerals of interest. Another set of methods is radial basis functions, also known as nearest neighbor mapping functions. This approach has been applied to determine formation mineralogy based on elemental concentrations obtained from gamma-ray spectroscopic logging measurements performed from the wellbore. Other methods can also be used to determine the mineralogical data of rock samples.
[0060] In some embodiments, the supervised learning process 10 may use one or both of two methods that take into account the heteroscedasticity of the input data (e.g., varying levels of uncertainty in NMR measurements and different elemental characteristics of rock samples), the need to combine multi-fidelity datasets (e.g., core and well logging data), and the need to obtain output data and its uncertainty estimates. These two methods are Bayesian regression methods, namely Bayesian Neural Networks (BNNs), and discarding methods.
[0061] Now for reference Figure 5 A diagram illustrating the general structure of an artificial neural network (ANN) is provided, which can be used to implement a supervised learning process using either a BNN or a discarding method. The ANN 502 can be configured to determine the permeability of a rock sample based on a highly nonlinear combination of input features. Input features can include a predetermined number (e.g., 6) of T2 features reduced from a T2 distribution determined using an SVD-based kernel. Specifically, T2 feature data can be derived by encoding the T2 distribution of the rock sample using an SVD-based kernel and then mapping the T2 distribution data to T2 features in a reduced-dimensional space. Other input features include NMR-based input data (e.g., porosity, bound-free volume ratio, T2 log-mean) and mineralogical input data. Nonlinearity can be introduced into the ANN through a nonlinearity or activation function applied to each neuron of the ANN. The ANN can be trained by fixing the kernel weights 504 and updating the neural network weights 506 by minimizing the MSE loss function. To improve performance, a batch normalization step can be performed at each layer of the ANN.
[0062] In some embodiments according to this disclosure, the supervised learning process 10 may utilize a computational model based on a trained ANN. The computational model may further derive values representing the uncertainty associated with the estimation of the permeability of the rock formation sample.
[0063] In some embodiments, according to this disclosure, the supervised learning process 10 may utilize a computational model based on training an artificial neural network employing discarded Bayesian inference. To determine the uncertainty in the permeability prediction of a rock sample, instead of simply determining the point estimate of permeability, as... Figure 5 As shown, ANNs can be implemented as Bayesian neural networks or using discarded Bayesian inference.
[0064] In some embodiments according to this disclosure, the supervised learning process 10 may utilize a computational model based on a trained BNN. For the BNN, it may be assumed that each weight of the BNN follows a normal distribution, characterized by two parameters μ and σ, which introduce uncertainty in penetration rate prediction, such as... Figure 6 As shown. In general, Figure 6 A simplified BNN architecture is shown. By sampling from the BNN weights, the penetration value k and the posterior distribution p(k) can be predicted. The training process for BNNs can vary because different loss functions, called the lower bound on evidence (ELBO), can be minimized based on Bayesian theory. This function estimates how close the true posterior distribution is to its approximation.
[0065] In some embodiments, an alternative way to obtain penetration uncertainty can be achieved using an ANN that applies dropout during both the training and testing phases, which is an approximate Bayesian inference process. Methods such as Concrete Dropout can improve this process by automatically determining the probability parameter p used to detach each neuron of the ANN.
[0066] In some embodiments, in addition to providing uncertainty in the output prediction of permeability from rock samples, permeability determination can be obtained from heteroscedastic, multifidelity datasets, such as datasets from different well logging and core measurements. To include different uncertainties in the input features, the initial dataset can be augmented by sampling each data point around its mean using the standard deviation of the noise in each feature. By doing so, the network can become more robust to noisy input data. Furthermore, if one or more data streams come from different sources (i.e., well logging and core measurements), and it is assumed that one source provides more accurate data than the other, the network can... Figure 7 The training process is illustrated in two consecutive steps. Typically, low-fidelity data is data with a low signal-to-noise ratio (SNR), such as downhole logging data, while high-fidelity data is data with a high SNR, usually laboratory data or field data with station measurements. Here, SNR is one of the main drivers of fidelity.
[0067] Figure 7The diagram illustrates sampling of input and output from low-fidelity data 702, with a neural network application 704. Transfer learning 706 can then be applied, and fine-tuning 708 of the high-fidelity data can be performed to provide a penetration rate answer product prediction 710. Specifically, Figure 7 The workflow for determining permeability using either a BNN or Dropout method in two separate panels is illustrated. "Low-fidelity" data, appropriately sampled with heteroscedastic noise, can first be used to train the network. This operation allows the network to reach good local minima, which will be the starting point for the second step. In the second step, transfer learning can be used to initialize the parameters of the neural network for training on "high-fidelity" data, which is also sampled for noise in the features, to determine the final network used to predict permeability, denoted as the final MSE, along with the relevant uncertainty of the rock sample. For each method, each involves training on the low-fidelity data, passing weights, and training on the high-fidelity data to obtain the final model for permeability prediction. In this example, n can refer to a new training dataset obtained for each epoch. Specifically, the low-fidelity data can be trained using either an ANN or dropout. The weights of the BNN or dropout can then be passed to where the high-fidelity data can be trained. The final permeability, denoted as the final MSE, can then be predicted.
[0068] Now for reference Figures 8A-8D This demonstrates the application of the aforementioned penetration rate prediction technology. Specifically, Figure 8A A graph showing the true permeability (x-axis) versus the predicted permeability (y-axis) based on the model of Equation 1 is shown. The logarithmic MSE value of this graph is 0.55. Figure 8B A graph showing the true permeability (x-axis) versus the predicted permeability (y-axis) based on the Timur-Coates model according to Equation 3 is presented. The logarithmic MSE value of this graph is 0.49. Figure 8C Based on Figure 4 and 5 A graph of the true permeability (x-axis) versus the predicted permeability (y-axis) of the ANN model. The logarithmic MSE value of this graph is 0.16. Figure 8D Based on Figure 6 and 7 A graph showing the true permeability (x-axis) versus predicted permeability (y-axis) of the BNN model. The logarithmic MSE value of this graph is 0.15. Note that, compared to... Figure 8A The KSDR model shown Figure 8B The Timur-Coates model shown or Figure 8C Compared to the simple artificial neural network model shown, Figure 8D Bayesian neural networks (i.e., those with lower MSE values) show better penetration predictions and can also provide additional information on penetration uncertainty.
[0069] In some embodiments, an evaluation metric may be introduced to assess one or more values of the predicted uncertainty for permeability. The uncertainty calibration metric assesses whether the predicted standard deviation at each permeability point is significantly close to the true permeability value. These results are summarized in... Figure 9A and 9B In the figure, the quality of the permeability uncertainty prediction is assessed by observing the proximity between the two lines. Furthermore, Figure 9A and 9B This plot represents the % confidence intervals around the predictability of the prediction (x-axis) against the % predictions that include the true value (y-axis) within the confidence intervals. For continuous z... The [0, 100] value is used to calculate one percentage point, predicting that the surrounding z% confidence interval contains the true penetration value.
[0070] In some embodiments according to this disclosure, multiple datasets may include data derived from NMR measurements of rock strata samples and mineralogical data corresponding to the rock strata samples. Rock strata samples may include, but are not limited to, rock cuttings, rock cores, drill cuttings, rock outcrops, or rock strata surrounding boreholes and coal.
[0071] In some embodiments according to this disclosure, the supervised learning process 10 introduces a probabilistic programming-based supervised machine learning workflow for regression problems in industrial applications. This workflow addresses various challenges faced in petroleum physics applications, such as the interpretation of subsurface multiphysics measurements.
[0072] In some embodiments according to this disclosure, the supervised learning process 10 may include obtaining multiple datasets representing samples. The method for supervised learning of rock physical parameters of the formation may further include providing a neural network with one or more discards. Low-fidelity datasets associated with the multiple datasets can be used to train the computational model.
[0073] In some embodiments, the fidelity of the measurement can be captured in its known probability density function, typically computed during the calibration of the relevant hardware. For example, when the probability density function is a multivariable Gaussian function, this fidelity can be adequately captured in the covariance matrix. To respect the noise information present in the input and output, sampling can be performed from known probability density functions in the input feature space and output variable space. One approach, referred to in this paper as the “sampling” method, is to directly train a neural network model dynamically (i.e., at each training epoch) on this sampled dataset, with the neural network backpropagating across different noisy implementations of the training sample dataset. Another approach, referred to in this paper as the “autoencoder” method, involves training an autoencoder on a dataset where the output can be a pure dataset and the input can be a dynamically generated, noise-corrupted dataset. Once training for this autoencoder has converged and the noise model of the data has been captured by the autoencoder, at least one parameter associated with the first or second autoencoder can be frozen, and the encoder can be used as a denoiser for any input-output pair, which can then be used to train a neural network model to learn the mapping from the input to the output. However, once noise is removed, there is no need to train the neural network model with dynamically sampled data, as mentioned in the "sampling" method.
[0074] In some embodiments according to this disclosure, the supervised learning process 10 may utilize a computational model, including a BNN or a standard neural network with a dropout computation module. In this example, the cost function minimized during training may be designed for a specific application. The distribution of the resulting likelihood can form the basis for selecting the cost function, based on intelligent guesses provided for the posterior distribution of the dataset based on domain expertise. For example, some of the most common cost functions are generated by Gaussian and Laplace likelihood assumptions (i.e., MSE and mean absolute error). Furthermore, new cost functions can be generated, such as those assuming heteroscedasticity in the standard Gaussian likelihood might cause heteroscedasticity loss. Currently, for neural networks with Concrete Dropout, heteroscedasticity loss can be used to obtain well-calibrated predictive uncertainty.
[0075] In some embodiments, the supervised learning process 10 may assume high-fidelity and low-fidelity datasets respectively. )and( To handle imbalanced multifidelity datasets, one can first train (…). ANN / BNN on ) . Additionally, using ( The weights of the obtained neural network can be fine-tuned.
[0076] Regarding the "sampling" method mentioned above, a computational model can be trained on a low-fidelity dataset using a standard neural network with one or more discarded parameters, and then fine-tuned using a high-fidelity dataset. When using a BNN as the neural network, a standard neural network model can be trained on a low-fidelity dataset using the "sampling" method, the learned parameters can be passed to a Bayesian architecture, and then the Bayesian model can be fine-tuned using a high-fidelity dataset. In both methods, the computational model can be updated by dynamically using different noise levels in the input data at each layer of the neural network.
[0077] In some embodiments, and also referring to Figure 10 The supervised learning process 10 can utilize the "autoencoder" method. For example, a first autoencoder can be trained using a low-fidelity dataset. Furthermore, a second autoencoder can be trained using a high-fidelity dataset. For instance, when using a standard neural network with discarding, two different autoencoders can be trained on the low-fidelity and high-fidelity datasets respectively, as... Figure 10 As shown. Low-fidelity input data can be trained using a low-fidelity encoder and used to train the computational model. The computational model can be further fine-tuned using denoised high-fidelity data obtained from a high-fidelity encoder. The neural network within the autoencoder can also be Dropout or a BNN, or it can be replaced by an application-based variable autoencoder. When using a BNN, the supervised learning process 10 remains the same as when using a standard neural network with dropout. The only difference is that the standard neural network is trained with denoised low-fidelity input, the learned parameters are fed to a Bayesian model, and the Bayesian model can be fine-tuned with denoised high-fidelity input. See again. Figure 10 For BNN and Dropout, the first step of training may include using low-fidelity data and a frozen low-fidelity encoder. Similarly, the second and final step of training may include using high-fidelity data and a frozen high-fidelity encoder to fine-tune the weights of the BNN or Dropout model. The final penetration rate can then be predicted, denoted as the final MSE and the final MSE, respectively.
[0078] Furthermore, regarding methods for measuring model robustness, accuracy can be used. And MSE, which depends on the application domain. For uncertainty, negative log-likelihood can be used. In some embodiments according to this disclosure, uncertainty can be evaluated in two ways. First, the robustness of the absolute value of one or more uncertainties can be determined. This can be done by determining... Does the test case actually exist on the ground? The uncertainty is measured within a confidence interval. Furthermore, how good one or more uncertainty values are can be determined based on their relative relative importance. Ideally, predictions generated from the in-distribution (ID) dataset might require low uncertainty, while those for the OOD dataset would require higher uncertainty. For this purpose, an AUC score can be used. When using an AUC score, the supervised learning process 10 can be modified to work in a regression setting. For example, one or more prediction uncertainties can be used as an anomaly detector to distinguish between good and bad predictions, with good predictions typically appearing on the ID dataset and bad predictions typically appearing on the OOD dataset. To evaluate the AUC, one or more ground truth binary labels and the classifier's score might be needed. In a regression setting, the ground truth binary label could be 1 for all ID data points and 0 for all OOD data points. Furthermore, the classifier's score can be the uncertainty of the model's prediction output for each ID and OOD data point. Moreover, the use of autoencoder techniques and their integration with uncertainty metrics can make this workflow highly relevant to industrial problems.
[0079] This document has described and illustrated several embodiments of methods and systems for learning and applying computational models that map a set of input features to estimates of the permeability of rock formation samples. While specific embodiments of the invention have been described, this does not imply that the invention is limited thereto, as this implies that the invention is broad within the scope permitted in the art, and the specification is equally applicable. Therefore, while specific neural network architectures and workflows have been disclosed, it should be understood that other specific neural network architectures and workflows may also be used. Consequently, those skilled in the art will understand that other modifications can be made to the provided invention without departing from the spirit and scope of the claims.
[0080] Some of the methods and processes described above can be implemented as computer program logic used with a computer processor. Computer program logic can be manifested in various forms, including source code or computer-executable form. Source code can include various programming languages (e.g., object code, assembly language, or languages such as C, C++, etc.). ++ A computer program instruction (or a high-level language like Java) is a set of computer program instructions. These computer instructions can be stored in a non-transitory computer-readable medium (such as memory) and executed by a computer processor. The computer instructions can be distributed in any form as a removable storage medium with accompanying printed or electronic documentation (such as shrink-wrapping software), pre-loaded onto a computer system (e.g., on a system ROM or fixed disk), or distributed from a server or electronic bulletin board via a communication system (e.g., the Internet or the World Wide Web).
[0081] Alternatively or additionally, the processor may include discrete electronic components coupled to a printed circuit board, an integrated circuit (e.g., an application-specific integrated circuit (ASIC)), and / or a programmable logic device (e.g., a field-programmable gate array (FPGA)). Any of the methods and processes described above can be implemented using such a logic device.
[0082] In one respect, any or any part or all of the steps or operations of the methods and processes described above can be performed by a processor. The term "processor" should not be construed as limiting the embodiments disclosed herein to any particular type of device or system. A processor may include a computer system. A computer system may also include a computer processor (e.g., a microprocessor, microcontroller, digital signal processor, or general-purpose computer) for performing any of the methods and processes described above.
[0083] The computer system may also include memory, such as semiconductor memory devices (e.g., RAM, ROM, PROM, EEPROM, or flash programmable RAM), magnetic memory devices (e.g., disks or fixed disks), optical memory devices (e.g., CD-ROMs), PC cards (e.g., PCMCIA cards), or other storage devices. The memory may be used to store any or all datasets from the methods and processes described above, such as datasets containing NMR-derived input features of rock samples, T2 distributions of rock samples, and mineralogical data of rock samples.
[0084] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems and methods according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, including one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions mentioned in the blocks may not occur in the order shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each block illustrated in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a system based on dedicated hardware or a combination of dedicated hardware and computer instructions that performs the specified function or action.
[0085] The terminology used herein is for the purpose of describing particular embodiments and not for the purpose of limiting this disclosure. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that, when used in this specification, the terms “comprising” and / or “including” specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.
[0086] The equivalents of the corresponding structures, materials, actions, and means or steps plus functional elements in the following claims are intended to include any structure, material, or action for performing a function in conjunction with other claimed elements of the specific claim. The description of this disclosure has been presented for purposes of illustration and description, but is not intended to be exhaustive or limited to the disclosure in its disclosed form. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of this disclosure. This embodiment was chosen and described in order to best explain the principles and practical application of this disclosure and to enable others skilled in the art to understand the disclosure of various embodiments with various modifications suitable for the particular intended use.
[0087] Although several exemplary embodiments have been described in detail above, those skilled in the art will readily understand that many modifications are possible in the exemplary embodiments without substantially departing from the scope of this disclosure as described herein. Therefore, such modifications are intended to be included within the scope of this disclosure as defined in the following claims. In the claims, the means plus function clause is intended to cover the structures described herein that perform the functions, and not only structural equivalents, but equivalent structures. Thus, although nails and screws may not be structural equivalents, since nails have a cylindrical surface for securing wooden parts together and screws have a helical surface, in the context of fastening wooden parts, nails and screws may be equivalent structures. The applicant’s explicit intention is not to invoke 35 USC § 112 6 to impose any limitation on any claim in this application, unless the claim expressly uses the term “means for…” and the associated function.
[0088] Therefore, the disclosure has been described in detail with reference to embodiments of this application, and it will be apparent that modifications and variations are possible without departing from the scope of the disclosure as defined in the appended claims.
Claims
1. A method for characterizing rock strata samples, comprising: Obtain multiple datasets characterizing rock strata samples; A neural network is trained to generate a computational model to predict the permeability of a rock formation using a low-fidelity training dataset in order to determine initial parameters, where permeability includes the flow rate of fluid through the rock formation. Transmit initial parameters to initialize training using a high-fidelity training dataset; Fine-tuning the neural network using a high-fidelity training dataset; The processor uses a singular value decomposition (SVD)-based encoder to encode the T2 distribution of the lithostratigraphic samples in the first-dimensional space to determine the encoded representation of the T2 distribution in the second-dimensional space. The encoded representation is then mapped to a simplified T2 feature set representing the lithostratigraphic samples, where the second-dimensional space is smaller than the first-dimensional space. Multiple input datasets are used as inputs to a computational model to determine the permeability and permeability uncertainty of a rock formation. The multiple input datasets include nuclear magnetic resonance (NMR) features of rock formation samples, a simplified T2 feature set representing the rock formation samples, and mineralogical data corresponding to the rock formation samples. The NMR features of the rock formations have a first fidelity, the mineralogical data have a second fidelity different from the first fidelity, and the computational model derives the permeability uncertainty based at least in part on the first fidelity and the second fidelity.
2. The method according to claim 1, wherein the computational model is based on a trained artificial neural network.
3. The method according to claim 1, wherein the computational model is based on a trained Bayesian neural network.
4. The method of claim 1, wherein the computational model is based on a trained artificial neural network, the artificial neural network employing discarded Bayesian inference.
5. The method of claim 1, wherein the plurality of input datasets comprises data obtained from NMR measurements of rock formation samples.
6. The method of claim 1, wherein the plurality of input datasets includes elemental data corresponding to the rock strata sample.
7. The method according to claim 1, wherein the rock formation sample is selected from the group consisting of rock cuttings, rock cores, rock drill cuttings, rock outcrops or rock strata surrounding boreholes and coal.
8. The method according to claim 1, wherein the second dimension space comprises 6 dimensions.
9. The method of claim 8, wherein the T2 distribution of the rock strata sample is encoded in a dimensional space comprising 64 dimensions.
10. The method of claim 1, wherein the first fidelity comprises well logging data, and the second fidelity comprises actual ground measurements.
11. The method of claim 1, wherein the first fidelity is lower than the second fidelity.
12. The method of claim 1, wherein the second dimension is at least 9% smaller than the first dimension.
13. The method according to claim 1, wherein the second dimension is no more than 10% smaller than the first dimension.
14. The method according to claim 1, wherein the first dimension space comprises 64 dimensions and the second dimension space comprises 6 dimensions.
15. The method of claim 1, wherein the low-fidelity training dataset comprises data having a lower signal-to-noise ratio than the high-fidelity training dataset.
16. The method of claim 1, wherein the high-fidelity training dataset comprises at least one of helium or nitrogen flow measurements of the core sample.
17. The method of claim 1, wherein the low-fidelity training dataset comprises downhole logging data.
18. A system for characterizing rock strata samples, comprising: processor; A non-volatile memory, connected to the processor, stores multiple datasets characterizing rock strata samples; and Instructions stored in the non-volatile memory, when executed by the processor, cause the system to: A neural network is trained to generate a computational model to predict the permeability of a rock formation using a low-fidelity training dataset in order to determine initial parameters, where permeability includes the flow rate of fluid through the rock formation. Transmit initial parameters to initialize training using a high-fidelity training dataset; Fine-tuning the neural network using a high-fidelity training dataset; The processor encodes the T2 distribution of the rock stratigraphic samples in the first-dimensional space using a singular value decomposition (SVD)-based encoder to determine the encoded representation of the T2 distribution in the second-dimensional space. This encoded representation is then mapped to a simplified T2 feature set representing the rock stratigraphic samples, where the second-dimensional space is smaller than the first-dimensional space. Multiple input datasets are used as inputs to a computational model to determine the permeability and permeability uncertainty of a rock formation. The multiple input datasets include nuclear magnetic resonance (NMR) features of rock formation samples, a simplified T2 feature set representing the rock formation samples, and mineralogical data corresponding to the rock formation samples. The NMR features of the rock formations have a first fidelity, the mineralogical data have a second fidelity different from the first fidelity, and the computational model derives the permeability uncertainty based at least in part on the first fidelity and the second fidelity.
19. The system of claim 18, wherein the computational model is based on at least one of a trained artificial neural network or a Bayesian neural network.
20. The system of claim 18, wherein the computational model is based on a trained artificial neural network employing discarded Bayesian inference.