Core Data Augmentation for Petrophysical Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The limited availability and high cost of core data for training petrophysical interpretation models, particularly in geospatial mapping and hydrocarbon exploration, hinder the generation of accurate and functional geospatial maps and models due to sparse data distribution across large areas.
Innovation Solution
The method employs Radial Basis Mapping Function (RBF) and Principal Component Analysis (PCA) to augment core sample datasets by generating synthetic data, which is integrated into the original dataset, enhancing the dataset's geographic coverage and complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If core data is collected from more wells to improve model training, then the quantity of training data increases, but the cost and complexity of data acquisition increases
Solution Approach 1:
The patent creates synthetic core data copies through data augmentation techniques including noise injection, scaling, and transformation operations. These synthetic copies simulate additional core samples without requiring physical collection from new wells, thereby increasing training data quantity while avoiding the complexity and cost of additional field operations
Solution Approach 2:
The patent transforms existing core data by applying various parameter changes such as adding different types of noise (Gaussian, uniform, salt-and-pepper), scaling amplitude and time axes, and applying frequency domain transformations. These parameter modifications create diverse synthetic variants from limited original samples, effectively multiplying the training dataset without additional data collection efforts
2Quantity of substance
If core samples are collected sparsely to reduce cost, then the cost of data acquisition decreases, but the geographic coverage and representativeness of the dataset decreases
Solution Approach 1:
The patent extends the limited spatial dataset into additional dimensions through synthetic data generation. By creating augmented samples that preserve the statistical and spatial characteristics of the original sparse data, the method effectively expands the geographic coverage and representativeness without requiring physical samples from every location
Solution Approach 2:
The patent performs preliminary data augmentation on the sparse core dataset before model training. By pre-processing the limited available data to create a more comprehensive synthetic dataset that captures the full range of geological variability, the method prepares the training data in advance to compensate for the sparse sampling strategy
3Reliability
If more core data is collected to avoid model overfitting, then the model generalization improves, but the time and resources required for data collection increases
Solution Approach 1:
The patent generates synthetic copies of core data through augmentation techniques that preserve the underlying geological patterns and statistical properties. These synthetic copies provide the model with diverse training examples needed for generalization without requiring time-consuming field collection of additional physical core samples
Solution Approach 2:
The patent applies parameter transformations including noise addition, scaling, and frequency domain modifications to create varied training samples. These parameter changes expose the model to different data variations during training, improving generalization capability while avoiding the time investment required to collect and process additional physical core data from multiple wells
Data Source
AI summary
A method for training a model. The method may include forming a data set from one or more measurements of core samples, selecting one or more parameters from the data set, inputting the one or more parameters into a kernel estimation function, determining a kernel density estimation from the kernel estimation function based at least in part on the one or more parameters, and selecting an input value based at least in part on the kernel density estimation. The method may further include creating a corresponding synthetic target value based at least in part on the input value, augmenting the data set with the corresponding synthetic target value and input value to form a synthetic data set, and training a petrophysical interpretation machine learning model from the data set and the synthetic data set.


