Machine learning-based automated well log quality check and reconstruction
Patent Information
- Authority / Receiving Office
- US · United States
- Patent Type
- Applications(United States)
- Current Assignee / Owner
- Filing Date
- 2026-02-05
- Publication Date
- 2026-08-13
AI Technical Summary
The quality assurance of well log datasets, however, remains difficult to perform at scale.
Smart Images

Figure US20260235019A1-D00000_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 757,413 filed on Feb. 12, 2025, the entirety of which is incorporated herein by reference to the extent consistent with the present disclosure.BACKGROUND
[0002] Well log data is widely used to characterize subsurface formations and support decisions in drilling, completion, and reservoir evaluation. The quality assurance of well log datasets, however, remains difficult to perform at scale. In many workflows, determining whether a log is usable, identifying gaps or flaws, and assessing consistency across runs or wells involves manual review by trained operators. This manual review is often labor-intensive, subjective, and prone to variability, particularly when datasets are large, sourced from multiple vendors, or collected under differing acquisition conditions. Normalization and standardization of well log data, which can improve comparability across wells and enhance downstream analysis, also commonly depend on manual or semi-manual procedures. For example, practitioners may manually select reference intervals, apply ad hoc scaling or shifting, and reconcile naming or unit inconsistencies across datasets. These approaches can be time-consuming and difficult to reproduce, and may introduce errors that propagate into interpretation, modeling, and automated analytics.
[0003] What is needed, then, are methods for processing well log data in a manner that is scalable and consistent, and that provides reliable quality assurance and normalization.SUMMARY
[0004] A method for processing well log data is disclosed. The method may include obtaining input data including the well log data. The well log data may include a plurality of datasets. Each dataset of the plurality of datasets may include one or more log curves and may be associated with a respective well of one or more wells. The method may also include harmonizing the plurality of datasets using a curated dictionary to produce harmonized datasets. The method may further include reconstructing the one or more log curves of each harmonized dataset of the harmonized datasets to produce reconstructed logs using a supervised machine-learning (ML) model. The method may also include generating normalized datasets based on the harmonized datasets and using log normalization. The method may also include generating an output based on the reconstructed logs and the normalized datasets.
[0005] A computing system is also disclosed. The computing system includes one or more processors and a memory system. The memory system includes one or more non-transitory computer-readable media storing instructions that, when executed by at least one of the one or more processors, cause the computing system to perform operations for processing well log data. The operations may include obtaining input data including the well log data. The well log data may include a plurality of datasets. Each dataset of the plurality of datasets may include one or more log curves and may be associated with a respective well of one or more wells. The operations may also include harmonizing the plurality of datasets using a curated dictionary to produce harmonized datasets. The operations may further include reconstructing the one or more log curves of each harmonized dataset of the harmonized datasets to produce reconstructed logs using a supervised machine-learning (ML) model. The operations may also include generating normalized datasets based on the harmonized datasets and using log normalization. The operations may also include generating an output based on the reconstructed logs and the normalized datasets.
[0006] A non-transitory computer-readable medium is also disclosed. The medium stores instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations for processing well log data. The operations may include obtaining input data including the well log data. The well log data may include a plurality of datasets. Each dataset of the plurality of datasets may include one or more log curves and may be associated with a respective well of one or more wells. The operations may also include harmonizing the plurality of datasets using a curated dictionary to produce harmonized datasets. The operations may further include reconstructing the one or more log curves of each harmonized dataset of the harmonized datasets to produce reconstructed logs using a supervised machine-learning (ML) model. The operations may also include generating normalized datasets based on the harmonized datasets and using log normalization. The operations may also include generating an output based on the reconstructed logs and the normalized datasets.
[0007] It will be appreciated that this summary is intended merely to introduce some aspects of the present methods, systems, and media, which are more fully described and / or claimed below. Accordingly, this summary is not intended to be limiting.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present teachings and together with the description, serve to explain the principles of the present teachings. In the figures:
[0009] FIG. 1 illustrates an example of a system that includes various management components to manage various aspects of a geologic environment, according to an embodiment.
[0010] FIGS. 2A-2D illustrate a flow chart of an exemplary method for processing well log data associated with one or more wells, according to an embodiment.
[0011] FIG. 3 illustrates a flow chart of an exemplary method for processing well log data associated with one or more wells, according to an embodiment.
[0012] FIG. 4 illustrates a schematic view of a computing system for performing at least a portion of the method(s) described herein, according to an embodiment.DETAILED DESCRIPTION
[0013] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings and figures. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0014] It will also be understood that, although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first object or step could be termed a second object or step, and, similarly, a second object or step could be termed a first object or step, without departing from the scope of the present disclosure. The first object or step, and the second object or step, are both, objects or steps, respectively, but they are not to be considered the same object or step.
[0015] The terminology used in the description herein is for the purpose of describing particular embodiments and is not intended to be limiting. As used in this description and the appended claims, the singular forms “a,”“an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any possible combinations of one or more of the associated listed items. It will be further understood that the terms “includes,”“including,”“comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Further, as used herein, the term “if” may be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” depending on the context.
[0016] Attention is now directed to processing procedures, methods, techniques, and workflows that are in accordance with some embodiments. Some operations in the processing procedures, methods, techniques, and workflows disclosed herein may be combined and / or the order of some operations may be changed.System Overview
[0017] FIG. 1 illustrates an example of a system 100 that includes various management components 110 to manage various aspects of a geologic environment 150 (e.g., an environment that includes a sedimentary basin, a reservoir 151, one or more faults 153-1, one or more geobodies 153-2, etc.). For example, the management components 110 may allow for direct or indirect management of sensing, drilling, injecting, extracting, etc., with respect to the geologic environment 150. In turn, further information about the geologic environment 150 may become available as feedback 160 (e.g., optionally as input to one or more of the management components 110).
[0018] In the example of FIG. 1, the management components 110 include a seismic data component 112, an additional information component 114 (e.g., well / logging data), a processing component 116, a simulation component 120, an attribute component 130, an analysis / visualization component 142 and a workflow component 144. In operation, seismic data and other information provided per the components 112 and 114 may be input to the simulation component 120.
[0019] In an example embodiment, the simulation component 120 may rely on entities 122. Entities 122 may include earth entities or geological objects such as wells, surfaces, bodies, reservoirs, etc. In the system 100, the entities 122 can include virtual representations of actual physical entities that are reconstructed for purposes of simulation. The entities 122 may include entities based on data acquired via sensing, observation, etc. (e.g., the seismic data 112 and other information 114). An entity may be characterized by one or more properties (e.g., a geometrical pillar grid entity of an earth model may be characterized by a porosity property). Such properties may represent one or more measurements (e.g., acquired data), calculations, etc.
[0020] In an example embodiment, the simulation component 120 may operate in conjunction with a software framework such as an object-based framework. In such a framework, entities may include entities based on pre-defined classes to facilitate modeling and simulation. A commercially available example of an object-based framework is the MICROSOFT® NET® framework (Redmond, Washington), which provides a set of extensible object classes. In the NET® framework, an object class encapsulates a module of reusable code and associated data structures. Object classes can be used to instantiate object instances for use in by a program, script, etc. For example, borehole classes may define objects for representing boreholes based on well data.
[0021] In the example of FIG. 1, the simulation component 120 may process information to conform to one or more attributes specified by the attribute component 130, which may include a library of attributes. Such processing may occur prior to input to the simulation component 120 (e.g., consider the processing component 116). As an example, the simulation component 120 may perform operations on input information based on one or more attributes specified by the attribute component 130. In an example embodiment, the simulation component 120 may construct one or more models of the geologic environment 150, which may be relied on to simulate behavior of the geologic environment 150 (e.g., responsive to one or more acts, whether natural or artificial). In the example of FIG. 1, the analysis / visualization component 142 may allow for interaction with a model or model-based results (e.g., simulation results, etc.). As an example, output from the simulation component 120 may be input to one or more other workflows, as indicated by a workflow component 144.
[0022] As an example, the simulation component 120 may include one or more features of a simulator such as the ECLIPSE™ reservoir simulator (SLB, Houston Texas), the INTERSECT™ reservoir simulator (SLB, Houston Texas), etc. As an example, a simulation component, a simulator, etc. may include features to implement one or more meshless techniques (e.g., to solve one or more equations, etc.). As an example, a reservoir or reservoirs may be simulated with respect to one or more enhanced recovery techniques (e.g., consider a thermal process such as SAGD, etc.).
[0023] In an example embodiment, the management components 110 may include features of a commercially available framework such as the PETREL© seismic to simulation software framework (SLB, Houston, Texas). The PETREL© framework provides components that allow for optimization of exploration and development operations. The PETREL© framework includes seismic to simulation software components that can output information for use in increasing reservoir performance, for example, by improving asset team productivity. Through use of such a framework, various professionals (e.g., geophysicists, geologists, and reservoir engineers) can develop collaborative workflows and integrate operations to streamline processes. Such a framework may be considered an application and may be considered a data-driven application (e.g., where data is input for purposes of modeling, simulating, etc.).
[0024] In an example embodiment, various aspects of the management components 110 may include add-ons or plug-ins that operate according to specifications of a framework environment. For example, a commercially available framework environment marketed as the OCEAN® framework environment (SLB, Houston, Texas) allows for integration of add-ons (or plug-ins) into a PETREL® framework workflow. The OCEAN® framework environment leverages .NET® tools (Microsoft Corporation, Redmond, Washington) and offers stable, user-friendly interfaces for efficient development. In an example embodiment, various components may be implemented as add-ons (or plug-ins) that conform to and operate according to specifications of a framework environment (e.g., according to application programming interface (API) specifications, etc.).
[0025] FIG. 1 also shows an example of a framework 170 that includes a model simulation layer 180 along with a framework services layer 190, a framework core layer 195 and a modules layer 175. The framework 170 may include the commercially available OCEAN® framework where the model simulation layer 180 is the commercially available PETREL® model-centric software package that hosts OCEAN® framework applications. In an example embodiment, the PETREL® software may be considered a data-driven application. The PETREL© software can include a framework for model building and visualization.
[0026] As an example, a framework may include features for implementing one or more mesh generation techniques. For example, a framework may include an input component for receipt of information from interpretation of seismic data, one or more attributes based at least in part on seismic data, log data, image data, etc. Such a framework may include a mesh generation component that processes input information, optionally in conjunction with other information, to generate a mesh.
[0027] In the example of FIG. 1, the model simulation layer 180 may provide domain objects 182, act as a data source 184, provide for rendering 186 and provide for various user interfaces 188. Rendering 186 may provide a graphical environment in which applications can display their data while the user interfaces 188 may provide a common look and feel for application user interface components.
[0028] As an example, the domain objects 182 can include entity objects, property objects and optionally other objects. Entity objects may be used to geometrically represent wells, surfaces, bodies, reservoirs, etc., while property objects may be used to provide property values as well as data versions and display parameters. For example, an entity object may represent a well where a property object provides log information as well as version information and display information (e.g., to display the well as part of a model).
[0029] In the example of FIG. 1, data may be stored in one or more data sources (or data stores, generally physical data storage devices), which may be at the same or different physical sites and accessible via one or more networks. The model simulation layer 180 may be configured to model projects. As such, a particular project may be stored where stored project information may include inputs, models, results and cases. Thus, upon completion of a modeling session, a user may store a project. At a later time, the project can be accessed and restored using the model simulation layer 180, which can recreate instances of the relevant domain objects.
[0030] In the example of FIG. 1, the geologic environment 150 may include layers (e.g., stratification) that include a reservoir 151 and one or more other features such as the fault 153-1, the geobody 153-2, etc. As an example, the geologic environment 150 may be outfitted with any of a variety of sensors, detectors, actuators, etc. For example, equipment 152 may include communication circuitry to receive and to transmit information with respect to one or more networks 155. Such information may include information associated with downhole equipment 154, which may be equipment to acquire information, to assist with resource recovery, etc. Other equipment 156 may be located remote from a well site and include sensing, detecting, emitting or other circuitry. Such equipment may include storage and communication circuitry to store and to communicate data, instructions, etc. As an example, one or more satellites may be provided for purposes of communications, data acquisition, etc. For example, FIG. 1 shows a satellite in communication with the network 155 that may be configured for communications, noting that the satellite may additionally or instead include circuitry for imagery (e.g., spatial, spectral, temporal, radiometric, etc.).
[0031] FIG. 1 also shows the geologic environment 150 as optionally including equipment 157 and 158 associated with a well that includes a substantially horizontal portion that may intersect with one or more fractures 159. For example, consider a well in a shale formation that may include natural fractures, artificial fractures (e.g., hydraulic fractures) or a combination of natural and artificial fractures. As an example, a well may be drilled for a reservoir that is laterally extensive. In such an example, lateral variations in properties, stresses, etc. may exist where an assessment of such variations may assist with planning, operations, etc. to develop a laterally extensive reservoir (e.g., via fracturing, injecting, extracting, etc.). As an example, the equipment 157 and / or 158 may include components, a system, systems, etc. for fracturing, seismic sensing, analysis of seismic data, assessment of one or more fractures, etc.
[0032] As mentioned, the system 100 may be used to perform one or more workflows. A workflow may be a process that includes a number of worksteps. A workstep may operate on data, for example, to create new data, to update existing data, etc. As an example, a may operate on one or more inputs and create one or more results, for example, based on one or more algorithms. As an example, a system may include a workflow editor for creation, editing, executing, etc. of a workflow. In such an example, the workflow editor may provide for selection of one or more pre-defined worksteps, one or more customized worksteps, etc. As an example, a workflow may be a workflow implementable in the PETREL© software, for example, that operates on seismic data, seismic attribute(s), etc. As an example, a workflow may be a process implementable in the OCEAN® framework. As an example, a workflow may include one or more worksteps that access a module such as a plug-in (e.g., external executable code, etc.).Machine Learning Based Automated Well Log Quality Check and Reconstruction
[0033] The present disclosure introduces a method, and a software tool for implementing the method (e.g., a UI-based plugin in Techlog®), that automates well log quality assurance tasks, such as for example, metadata checks, log inventory check, log homogenization, log normalization, outlier detection, depth matching, or a combination thereof. In some embodiments, the process or method may apply machine learning (ML) for one or more of outlier detection, log prediction, log normalization, depth matching, or a combination thereof, to ensure efficiency, consistency, and accuracy throughout the process.
[0034] Unlike conventional tools or manual workflows, in some embodiments, the method may utilize machine learning and automated algorithms, following a config-run-summary pattern that is user-friendly and highly customizable. Conventional tools often demand extensive manual intervention and trial-and-error processes to fine-tune hyperparameters, whereas the disclosed process eliminates the need for such adjustments, streamlining the entire workflow.
[0035] In some embodiments, the process may utilize a machine learning-based, automated well log quality check and reconstruction software tool (e.g., a UI-based plugin in Techlog®). This software tool may streamline data quality processes, from metadata and log inventory checks to outlier detection, log reconstruction, log normalization and depth matching, using advanced automation and visualization techniques. In some embodiments, the method may provide a standardized, scalable, and efficient solution for ensuring well log data integrity.
[0036] In some embodiments, the machine learning-powered software tool may transform well log quality assurance and reconstruction in well log applications (e.g., Techlog®) by automating complex processes such as metadata validation, log homogenization, and outlier detection. In some embodiments, the software tool may provide a user-friendly interface, comprehensive features, and robust machine learning algorithms to enhance productivity, reduce errors, and standardize workflows resulting in enhanced efficiency and improved data integrity for clients.Exemplary Method for Well Log Data Quality Assurance
[0037] FIG. 2 illustrates a flow chart of an exemplary method 200 for processing well log data associated with one or more wells, according to one or more embodiments. An illustrative order of the method 200 is discussed below; however, one or more portions of the method 200 may be performed in a different order, simultaneously, repeated, or omitted. At least a portion of the method 200 may be performed using a computing system.
[0038] The methods 200 may include receiving input data, as at 202. The methods 200 disclosed may include dataset selection to establish an initial processing set for downstream metadata enrichment, harmonization, quality control, and reconstruction operations or processes, as at 204. Inputs for data selection may include user-selected wells and datasets, which may include well name, dataset name, and available log curves, existing metadata (e.g., units, families, tool names, acquisition details if present, etc.), depth index and unit information, optional zonation tags from source systems, or any combination thereof.
[0039] The process for data selection may include querying the database for all available wells in the project. The process for data selection may also include retrieving, for each well, all continuous datasets. The process for data selection may also include filtering out system-generated datasets. The process for data selection may also include compiling a structured list of tuples for user selection. The tuples may include a well and a dataset. The process for data selection may also include ingesting selected datasets into the processing pipeline.
[0040] The output for data selection may include a structured dataset collection ready for zonation discovery, metadata checks, harmonization, quality assessment, and log reconstruction. The output for data selection may also include well-dataset mapping for downstream processing modules.
[0041] The methods 200 disclosed may also include zonation selection to determine common zonation across selected datasets for consistent normalization and comparison, as at 206. In at least one embodiment, the methods 200 disclosed may omit or skip zonation selection if no common zonation exists. The inputs for zonation selection may include the selected wells from data section, zonation definitions or tags per dataset from the database, zone interval bounds (top and bottom depths) per zone per dataset, or a combination thereof.
[0042] The process for zonation selection may include, for each well, querying all datasets that contain zonation information. The process for zonation selection may also include identifying zonation datasets that are common across all selected wells using set intersection. The process for zonation selection may also include verifying minimum depth overlap within same-named zones across wells. The process for zonation selection may also include storing store zone identifiers and interval bounds if common zonations are present. If no common zonation is identified or found a skip flag for zonation-dependent operations may be generated.
[0043] The output for zonation selection may include common zones list with interval bounds per dataset. The output for zonation selection may also include one or more diagnostics indicating no common zonation and a skip flag for zonation-dependent normalization.
[0044] The methods 200 disclosed may also include a metadata check and enrichment process, as at 208. The metadata check and enrichment may include an acquisition method imputation process, a log curve family inference process, a unit imputation via offset wells using quantile overlap process, a fallback unit imputation via family value range mapping process, a quantile computing and display process, a well path inference process, or any combination thereof, as at 210.
[0045] The acquisition method imputation may be for imputing or validating acquisition methods 200 (e.g., wireline vs logging while drilling [LWD]) from dataset naming conventions and tool indicators. The inputs for acquisition method imputation may include dataset names (string identifiers), a curated tool-to-acquisition-method dictionary mapping tool mnemonics to Wireline or LWD, or a combination thereof. The acquisition method imputation may include parsing dataset names using pattern matching and tokenization, extracting tool mnemonics (e.g., isolate the tool block following a run identifier), mapping extracted tokens to acquisition method using the curated dictionary, assigning methods as Wireline, LWD, or Unknown based on matching results, or a combination thereof. The output from the acquisition method imputation may include acquisition method labels for each dataset.
[0046] The log curve family inference may be for imputing or validating the family classification for each log curve (e.g., Gamma Ray, Bulk Density, Resistivity, Neutron Porosity, Sonic, etc.). The inputs for log curve family inference may include curve names (mnemonics) per dataset, a curated family dictionary mapping aliases to standard family names, or a combination thereof. The log curve family inference process may include normalizing curve names (remove special characters, standardize case), matching normalized names against the family dictionary, resolving to a standard family name via dictionary lookup, or a combination thereof. The outputs from the log curve family inference may include family assignments for each log curve in each dataset.
[0047] Unit imputation via offset wells using quantile overlap may impute missing units for log curves with known family by comparing quantile distributions with offset wells over overlapping depth intervals. The inputs for unit imputation via offset wells may include target dataset log curve values, P10 / P90 quantiles, depth index, unit, a family-specific unit dictionary with one or more candidate wells containing known units and precomputed quantiles, configurable thresholds for depth overlap (e.g., ≥80%), quantile overlap threshold, or the like, or any combination thereof. In one example, the family-specific unit dictionary may include at least one candidate well and up to 100 candidate wells, or more.
[0048] The process for unit imputation via offset wells may include selecting candidates (e.g., Select top N=5 offset datasets from the family's unit dictionary with known units, P10 / P90, and source identifiers), aligning depth and harmonizing the units, filtering depth overlap, comparing the quantile envelope, selecting and assigning a match unit, or a combination thereof. Aligning depth and harmonizing the units may include converting target and offset depth units into a common unit (e.g., feet), and computing depth spans. Filtering depth overlap may include determining an overlap percentage, and proceeding if a predetermined overlap percentage (e.g., >80%) is achieved. Comparing the quantile envelope may include determining an overlap of a target and offset, and identifying or selecting the highest overlap. Selecting and assigning the match unit may include assigning the identified match unit, recording the provenance (e.g., offset source, overlap metrics), or the like. In some examples, if no match qualifies then the method may utilize the fallback unit imputation. The output from the unit imputation via offset wells may include unit assignment for target variable (if matched), provenance information (e.g., chosen offset dataset, overlap percentages, etc.), or a combination thereof.
[0049] The fallback unit imputation via family value-range mapping may be utilized when the primary quantile-overlap method may not be applied. The fallback unit imputation via family value-range mapping may be fore inferring the unit by mapping the curve's median to curated unit ranges for the inferred family. The inputs for the fallback unit imputation via family value-range mapping may include target dataset and variable name, family assignments, a family-specific range table {Family, Unit, Min, Max}, or a combination thereof.
[0050] The process for the fallback unit imputation via family value-range mapping may include determining the median of the target curve (excluding invalid values), selecting candidate rows in the range table whose interval contains the median (Min≤median≤Max), applying one or more tie-breaking rules if multiple candidates exist, flagging user review if no candidate matches exist, or a combination thereof. The one or more tie-breaking rules may be or include, but are not limited to, preferring relatively narrower intervals, preferring higher containment overlap between target [P10, P90] and candidate [Min, Max], applying regional priority (if available), or any combination thereof. The outputs for the fallback unit imputation via family value-range mapping may include unit assignment via range mapping (or family default) or flags for manual review if no match is found.
[0051] The quantile computing and display process may determine and display distribution quantiles for each curve to support unit imputation and validation. The inputs may include the log curve values, excluding invalid values (e.g., −9999, NaN, etc.). The process for the quantile computing and display process may include filtering out invalid / sentinel values from each curve, computing the 10th percentile (P10), the median (P50), and the 90th percentile (P90), displaying quantiles alongside inferred family and unit for user validation, comparing quantile range to expected unit intervals, or a combination thereof. The output may include the P10, the P50, and the P90 values per log curve, and a distribution summary for transparency and validation.
[0052] The well path inference process may determine the well path classification (Vertical, Deviated, Highly Deviated, Horizontal, Snake) using borehole deviation measurements. The inputs for the well path inference process may include all borehole deviation family variables within a dataset, depth index, unit, or a combination thereof. The well path inference process may include identifying candidates, selecting variables (e.g., best variables), determining median, and mapping classifications. Identifying candidates may include gathering all deviation curve candidates under the borehole deviation family. Selecting variables may include determining non-missing count and coverage depth (span of valid depths) for each candidate, selecting the variable with the highest non-missing count, breaking ties using the largest coverage depth (optional tertiary: most recent acquisition or highest sampling density), recording the selection provenance (variable name, counts, coverage, tie-break method), or any combination thereof. The classifications may be mapped according to the following: 0°≤median<5°→Vertical, 5°≤median<450→Deviated, 45°≤median<800→Highly Deviated, 80°≤median<93°→Horizontal, median≥93°→Snake, and Else→Unknown. The outputs for the well path inference may include well path classification labels, provenance of the selection, or a combination thereof.
[0053] The methods 200 disclosed may also include a log inventory check, as at 212. The log inventory check may ensure mandatory families (e.g., True Vertical Depth (TVD), True Vertical Depth Sub Sea (TVDSS), X Offset, Y Offset) exist across datasets. The log inventory check may also determine, per-family depth, coverage against the measured-depth reference using unit-normalized interval extraction and merging. The log inventory check may also repair TVD / TVDSS gaps using regression-based extrapolation (endpoints) and linear interpolation (middle gaps). The inputs for the log inventory check may include log curve datasets with associated wells, log curve family assignments, a mandatory family list (e.g., TVD, TVDSS, X Offset, Y Offset), or a combination thereof.
[0054] The process for the log inventory check may include interval extraction and unit normalization (as at 214), interval merging (as at 216), determining a measured-depth (md) reference interval (as at 218), determining coverage fraction per main family (as at 220), aggregating results (as at 222), or a combination thereof. Interval extraction and unit normalization may include, for each (well, dataset, family) triple, retrieving all log curves under the family, for each log curve, obtaining the top and bottom depth with its unit, converting units to common or normalized units (e.g., meters), or a combination thereof. Determining coverage fraction per main family may include determining a total coverage length, a total possible length, and the coverage fraction based on the total coverage length and the total possible length. The outputs for the log inventory check may include a per-well inventory data structure with unit-normalized, merged intervals. The outputs for the log inventory check may also include coverage fractions per main family and average coverage. The outputs for the log inventory check may also include consistent feet canonical unit for all interval arithmetic. The outputs for the log inventory check may also include inputs to coverage dashboards and QC readiness metrics.
[0055] The methods 200 disclosed may also include data harmonization, as at 224. Data harmonization may enforce standardized aliases and units by family, trim datasets to reference extents, delete non-harmonized variables, and compute or compute conversion failure rates to quantify harmonization quality. The inputs for data harmonization may include datasets with metadata, a canonical family dictionary with standard aliases and units, a reference log family selection (e.g., Gamma Ray), unit conversion rules and factors, or a combination thereof.
[0056] The process for data harmonization may include canonical measured-depth (MD) establishment (as at 226), alias normalization and unit conversion (as at 228), dataset flagging for lineage and data governance (as at 230), depth trimming to a reference extent (as at 232), pruning and conversion metrics (as at 234), or a combination thereof. The canonical MD establishment process may include identifying the MD variable within each dataset with the highest non-missing sample count, harmonizing its unit to the canonical MD unit (e.g., feet or meters), renaming the MD variable to a canonical alias (e.g., “MD”), storing under the Measured Depth family, preserving sentinel values (e.g., −9999) and ignoring during conversion, or a combination thereof. The alias normalization and unit conversion process may include, for each non-depth curve: assigning the canonical alias for that curve's family (one alias per family per dataset), updating description metadata to preserve an audit trail to the original raw variable name, converting the curve's unit to the canonical family unit using pre-defined conversion factors (e.g., gamma ray→GAPI, bulk density→g / cc, sonic→μs / ft, Resistivity→ohm·m, etc.), or a combination thereof. In some examples, if a conversion cannot be resolved deterministically (e.g., missing unit metadata), a conversion error may be recorded, and the curve may be maintained with a low conversion confidence flag. Dataset flagging for lineage and data governance may include creating a dataset flag variable marking, per depth sample. Depth trimming to a reference extent may include, identifying a reference log and identifying the first and last valid samples for the reference log, reducing the dataset to that interval, and ensuring all harmonized variables share the same valid depth window. Pruning and conversion metrics may include removing any variables not in the harmonized set, removing duplicates and legacy variants that may cause ambiguity, determining conversion failure rates, determining well-level rate by averaging dataset rates, and generating metrics for quality dashboards and review workflows. The outputs for data harmonization may include a harmonized dataset containing one canonical alias per family with all in canonical units, a dataset trimmed to the reference log extent, one or more dataset flags for lineage tracking, per-dataset and / or per-well conversion failure metrics, family-level details (e.g., original variable, original unit, canonical alias, conversion error status), provenance artifacts documenting conversions, renames, family assignments, and trim decisions, or any combination thereof.
[0057] In at least one embodiment, data harmonization may also include clustering datasets through caliper readings. In another example, clustering datasets may be separated from the data harmonization. Clustering dataset may detect and label whether a single dataset is actually a merge of multiple runs by analyzing caliper consistency across depth using robust clustering. The resultant labels may become run tags to support per-run QC, normalization, and modeling. The inputs for clustering datasets may include caliper readings for the harmonized dataset, in canonical caliper unit (e.g., inches or mm), a smoothing window size (e.g., median filter window size 60) to reduce noise, an unsupervised clustering algorithm (e.g., K-means), an elbow detector (knee method) to determine optimal cluster count, a median-difference threshold for distinguishing distinct runs (e.g., >2 units), or a combination thereof.
[0058] The process for clustering datasets may include smoothing, an optimal cluster count, clustering and refining labels, run tagging, or a combination thereof. The smoothing may include applying a median filter with a configured window (e.g., 60 samples) to the caliper series to suppress short-term spikes while preserving long-interval shifts indicative of different runs. The optimal cluster count may include fitting a clustering model (e.g., K-means) across a candidate range (e.g., k=1.7), determining a sum of squared errors (SSE) for each k, utilizing a knee / elbow detector to identify the optimal k, reflecting the number of stable caliper levels, or a combination thereof. Clustering and refining labels may include clustering the smoothed caliper values with the chosen k, assigning a label per depth sample, reflecting groupings by characteristic caliper level, and determining median caliper per label. Run tagging may include converting final labels into run intervals by grouping contiguous depth segments of identical labels, and storing run tags. The output for clustering datasets may include run-level labels per depth sample, contiguous run intervals for the dataset, metadata enabling per-run QC, and normalization.
[0059] The methods 200 disclosed may also include a quality control process, as at 236. The quality control process may include one or more of an out-of-range detection process (as at 238), a linear segment detection process (as at 240), a bad hole flagging process (as at 242), or a combination thereof.
[0060] The out-of-range detection process may automatically detect and flag intervals where harmonized log curves violate user-defined bounds (minimum and maximum per log family). The out-of-range detection process may also produce per-curve error flags and a merged overall flag per dataset, and aggregate summary statistics for auditability and QC dashboards. The inputs for the out-of-range detection process may include harmonized datasets per well, including canonical aliases and units, user-provided cutoff configuration: a mapping per family to [min, max] bounds, example families: Gamma Ray, Resistivity, Density, Neutron Porosity, Sonic, etc., or a combination thereof.
[0061] The process for the out-of-range detection may include retrieving all harmonized curves for each well and dataset, checking each curve against the configured cutoffs for its family, flagging the depth samples that are out of range, or a combination thereof. In some examples, the process for the out-of-range detection may include adjusting cutoffs if units (e.g., fractions vs. percentage) vary. The process for the out-of-range detection may also include merging all curve-level flags into one dataset-level flag to indicate any out-of-range condition. The process for the out-of-range detection may also include summarizing results per well: total samples checked, out-of-range counts, and affected curves. The outputs for the out-of-range detection may include per-curve out-of-range flags, merged overall flag per dataset, summary tables for QC dashboards with statistics (total samples, flagged samples, percentage), or a combination thereof.
[0062] The linear segment detection process may identify log intervals that may be suspiciously constant or nearly linear over long depth ranges, which may indicate tool malfunction, data corruption, or synthetic fill. The inputs for the linear segment detection process may include harmonized datasets (canonical names and units), families to check (e.g., Gamma Ray, Resistivity, Density, Neutron Porosity, Sonic), a window size for a linear segment detection algorithm, variance or slope-based thresholds, or a combination thereof.
[0063] The process for the linear segment detection may include for each well and dataset, iterating through all harmonized curves in the selected families, applying a sliding window algorithm to analyze curve values, detecting long, flat, or near-linear segments using one or more of slope-based checks (e.g., consecutive samples with slope≈0) and / or variance-based checks (e.g., variance within window below threshold), flagging detected linear segments, merging all curve-level flags into one dataset-level flag for easy QC gating, summarizing results per well: total samples checked, count of flagged intervals, and affected curves, or a combination thereof. The outputs for the linear segment detection may include per-curve linear segment flags, a merged overall flag per dataset, summary tables for QC dashboards with statistics, or a combination thereof.
[0064] The bad hole flagging process may identify intervals where borehole conditions compromise log quality using up to three complementary physics-based rules. These identified intervals may be flagged to exclude them from training data or reconstruction. The physics-based rules may include one or more of a Caliper Rule, a Bulk Density Correction Rule, and a NPHI-RHOB-PEF Combined Rule. The inputs for the bad hole flagging process may include harmonized datasets with canonical units, caliper readings (borehole diameter), bulk density (RHOB) and bulk density correction (DRHO), neutron porosity (NPHI), photoelectric effect (PEF), out-of-range and linear segment flags, or a combination thereof. The process for bad hole flagging may include applying the rules and merging their respective results.
[0065] The caliper rule may flag intervals where the borehole is excessively enlarged or constricted based on caliper readings. The caliper rule may include determining median caliper per run separately, determining median caliper for the entire dataset, excluding samples already flagged as out-of-range or linear segments, defining thresholds (e.g., minimum and / or maximum), or a combination thereof. The caliper rule may output bad hole caliper flags per depth sample.
[0066] The bulk density correction rule may flag intervals where the bulk density correction (DRHO) exceeds acceptable limits, indicating poor pad contact or mudcake effects. The process for the bulk density correction rule may include retrieving DRHO (bulk density correction) values, excluding samples already flagged as out-of-range or linear segments, defining a threshold (e.g., DRHO>0.15 g / cc indicates bad hole), or a combination thereof. The bulk density correction rule may flag per depths samples having bad holes.
[0067] The NPHI-RHOB-PEF Combined Rule, also referred to as the combination of Neutron Porosity (NPHI), Bulk Density (RHOB), and Photoelectric Effect (PEF), may flag intervals exhibiting physically implausible combinations of Neutron Porosity (NPHI), Bulk Density (RHOB), and Photoelectric Effect (PEF), often due to washouts, mudcake, or gas effects. The process for the NPHI-RHOB-PEF Combined Rule may include retrieving NPHI, RHOB, and PEF values, excluding samples already flagged as out-of-range or linear segments, applying physics-based rules (configurable thresholds), and, for each depth sample, determine if any physics-based rule is violated and / or if any necessary logs are missing. The NPHI-RHOB-PEF Combined Rule may flag per depth samples.
[0068] The bad hold flagging process may include merging all the bad hole flags identified via the caliper rule, the DRHO rule, and the NPHI-RHOB-PEF combined rule. The process may include merging the flags and / or determining or generating summary statistics per dataset and per well. The outputs from the bad hole flagging process may include individual bad hole flags, merged bad hole flags per dataset, summary tables with statistics (e.g., total samples, bad hole counts, percentage affected, etc.), per-run statistics if run clustering is applied, or a combination thereof.
[0069] The methods 200 disclosed may include an unsupervised machine learning (ML) outlier detection process, as at 244. The unsupervised machine learning (ML) outlier detection process may identify multivariate outliers in the log data that rule-based methods may miss, using unsupervised machine learning algorithms that learn the normal data distribution. The inputs for the unsupervised machine learning (ML) outlier detection process may include the harmonized datasets, one or more selected log families for multivariate analysis (e.g., Gamma Ray, Bulk Density, Resistivity, Neutron Porosity), out-of-range, linear segment, and bad hole flags, one or more algorithms (e.g., One-Class SVM, Isolation Forest, Local Outlier Factor), one or more hyperparameters (e.g., contamination rate (e.g., 0.05=5%), n_neighbors, n_estimators, etc.), or a combination thereof.
[0070] The process for the unsupervised machine learning (ML) outlier detection may include data preparation, filtering, model training and prediction (e.g., per dataset), selecting algorithms or algorithm options, creating flags, generating or displaying a summary, or a combination thereof. The data preparation may include, for each (well, dataset), extracting values for selected log families, replacing values with NaN, removing values where all selected features are NaN, applying independent component analysis (ICA) for feature decorrelation, or a combination thereof. Filtering may include excluding samples already flagged by out-of-range, linear segment, or bad hole detection. Model training and prediction may include, per dataset, fitting the selected unsupervised algorithm (e.g., SVM, Isolation Forest) on the cleaned feature matrix, predicting outlier labels (e.g., 1=outlier, 0=normal, −9999=missing), or a combination thereof. In one example, the contamination parameter may control the expected proportion of outliers. The algorithms may include one or more of One-Class SVM, Isolation Forest, Local Outlier Factor (LOF), or a combination thereof. The flags may be created for outliers per dataset. The outlier flags may be merged with the other flags previously identified. The process for the unsupervised machine learning (ML) outlier detection may output one or more outlier flags per dataset, a merged summary of flags, a summary of statistics (e.g., total samples, outlier count, percentage, etc.), or a combination thereof. The process for the unsupervised machine learning (ML) outlier detection may also output outlier flags per depth sample per dataset, merged flags, summary tables for QC dashboards with outlier statistics, model diagnostics, or a combination thereof.
[0071] The methods 200 disclosed may also include ML-based log prediction and reconstruction, as at 246. The ML-based log prediction and reconstruction process may reconstruct missing log intervals and replace flagged (bad hole, outlier, out-of-range, linear) intervals using machine learning models trained on high-quality portions of the dataset. The ML-based log prediction and reconstruction process may support both dataset-level and well / field-level training. The inputs for the ML-based log prediction and reconstruction process may include the harmonized datasets, the flags (e.g., out-of-range, linear segment, bad hole, ML outlier, etc.), target families for prediction (e.g., Gamma Ray, Bulk Density, Resistivity, Neutron Porosity), one or more ML algorithms (e.g., LightGBM Regressor (gradient boosting)), one or more spatial features (X, Y coordinates for field-level models), or a combination thereof.
[0072] The process for the ML-based log prediction and reconstruction process may include data preparation and cleaning, feature combination generation, model training (e.g., per dataset), trajectory fitting for TVD / TVDSS, creating prediction flags, or any combination thereof.
[0073] The data preparation and cleaning may include aggregating flags, filtering training data, and consolidating features. Aggregating flags may include, identifying all flagged intervals for each dataset, preserving original values in variables for visualization comparison, or a combination thereof. Filtering training data may include excluding samples flagged with bad hole flags, excluding samples including cutoff or outlier flags for the target family, setting flagged target values to NaN, or a combination thereof. Consolidating features may include, for each main family (e.g., Bulk Density), if multiple sub-family curves exist, selecting the one with least missing values. Consolidating features may also include renaming a selected curve to the main family canonical name, dropping unselected sub-family curves, or a combination thereof.
[0074] Feature combination generation may include, for each target family to predict, generating all possible feature combinations using other available families. The feature combination generation process may start with the largest combination and progress to the minimum. For example, to predict RHOB, the combinations may be or include, but are not limited to, [GR, NPHI, DT, LLD], [GR, NPHI, DT], [GR, NPHI], [GR, LLD], etc.
[0075] The model training for the ML-based log prediction and reconstruction process may include, for each target family and / or each feature combination, one or more of depth trend extraction (e.g., physical pre-processing) (as at 248), a train-test split process (e.g., on detrended data), model training (e.g., on detrended data), ML prediction (e.g., on detrended data) (as at 250), trend restoration (e.g., physical post-processing) (as at 252), physical range clipping, a quality metrics process, or a combination thereof.
[0076] Depth Trend Extraction may include, before training, extracting and removing depth-dependent trends from log curves using curve fitting, fitting a logarithmic trend model for bulk density and resistivity, fitting negative logarithmic trends to enforce physical constraints for neutron porosity, removing trends from raw logs before training, or any combination thereof. The depth trend extraction may improve model generalization by removing strong depth-dependent components. The train-test split process may include a training set and a prediction set. The training set may include detrended samples where the target is valid (e.g., not NaN) and not flagged. The prediction set may include detrended samples where the target is NaN (e.g., missing or flagged). The model training may include training the LightGBM Regressor on the detrended training set. The model training may include one or more hyperparameters. The ML prediction may include predicting detrended target values for the prediction set using the trained model, storing the ML predictions as an intermediate output, or a combination thereof. Trend Restoration may include restoring the extracted trend back to predicted values to ensure the predicted logs maintain realistic depth-dependent trends. The physical range clipping may include clipping all predictions to family-specific physically plausible ranges. The quality metrics process may include determining an R2 score, RMSE on validation holdout (if available), selecting feature combinations based on validation metrics, determining confidence intervals, or a combination thereof. In some examples, if no validation set is present, all non-flagged samples may be used for training.
[0077] The trajectory fitting, as at 254, for true vertical depth (TVD) and true vertical depth subsea (TVDSS) may fill missing TVD / TVDSS values using geometric relationships with measured depth (MD) and borehole deviation. The process for trajectory fitting TVD / TVDSS may include, if TVD or TVDSS has gaps (not all missing), utilizing regression-based extrapolation for endpoint gaps, which may include fitting a polynomial or an exponential model to valid TVD vs. MD, extrapolating to missing endpoints, utilizing linear interpolation for middle gaps, which may include interpolating between valid samples. The process for trajectory fitting TVD / TVDSS may also include, if TVD or TVDSS is completely missing, determining from MD and deviation using geometric integration according to equation:TVDi=TVDi-1+(MDi-MDi-1)cos(θ1)where θi is the borehole deviation at depth index i. The process for trajectory fitting TVD / TVDSS may also include determining TVDSS from TVD according to equation:TVDSS=TVD-KBwhere TVD is true vertical depth referenced to the Kelly Bushing (KB), KB is the elevation of the Kelly Bushing above sea level, and TVDSS is the true vertical depth referenced to sea level. The process for trajectory fitting TVD / TVDSS may output filled TVD and TVDSS curves with provenance flags indicating filled vs original samples. Creating prediction flags for the ML-based log prediction and reconstruction process may include creating a prediction flag variable for one or more conditions (e.g., existing, missing, predicted), and storing the predicted values in the original family variable with overlayed flags / missing intervals.The methods 200 disclosed may include depth matching and splicing, as at 256. Depth matching and splicing may align multiple overlapping datasets from the same well by computing optimal depth shifts using cross-correlation, then splice them into a single composite dataset with maximum coverage and continuity. The inputs for depth matching and splicing may include the harmonized datasets per well, reconstructed logs, a reference log family for correlation (e.g., Gamma Ray), the quality metrics (e.g., coverage, non-missing sample count, acquisition method, etc.), or a combination thereof.The process for depth matching and splicing may include dataset prioritization, as at 258. Dataset prioritization may include prioritizing datasets to determine or establish an initial reference. Dataset prioritization may include coverage ranking, quality ranking, or a combination thereof. Coverage ranking may include determining depth coverage for each dataset (e.g., total valid depth range), ranking by coverage, or a combination thereof. Quality ranking may include determining or selecting a dataset among a plurality of datasets having similar or substantially similar coverage based on one or more qualities and / or properties of the dataset. For example, quality ranking may include selecting a dataset including a relatively higher number of non-missing sample counts in a reference log, a dataset acquired via wireline vs LWD, an acquisition date of the dataset (e.g., most recent), or a combination thereof. Data prioritization may be utilized to select an initial reference based on the coverage ranking, the quality ranking, or a combination thereof.The process for depth matching and splicing may include iterative depth matching based on the initial reference. Iterative depth matching may include preparing a reference and target logs. Preparing the reference and target logs may include extracting a reference log from the reference dataset (as at 260), extracting a reference log from one or more target datasets (as at 262), removing linear segments in overlap zones to avoid false correlations, or a combination thereof.
[0081] Iterative depth matching may also include autocorrelation based shift computation, as at 264, which may include checking overlaps (as at 266), resampling to a uniform grid (as at 268), cross-correlation and confidence (as at 270), or a combination thereof. Checking overlaps may include finding overlapping depth intervals and determining a minimum threshold for the overlap depth intervals. Resampling to the uniform grid may include resampling both the reference and target logs to a uniform depth grid (e.g., 0.5 ft spacing) within the overlap. Resampling may also include using linear interpolation. Cross-correlation and confidence may include determining normalized cross-correlation between the reference and target log over a depth range of depth shifts, shifting a range (e.g., ±20 ft to +50 ft), and determining a cross-correlation via the Pearson correlation coefficient. The Pearson correlation coefficient may identify a shift with a maximum correlation coefficient, and if the max correlation is below a threshold, generate a flag as low confidence.
[0082] Iterative depth matching may also include applying a shift and splice process, as at 272. The shift and splice may include shifting all logs, such as all depth values, in a target dataset, as at 274. The shift and splice may also include merging the datasets by appending the target data with non-overlap or keeping a reference data if there is overlap, as at 276. The shift and splice may also include a weighted blend to blend the reference and target using a quality-weighted average if there is an overlap, where weights may be based on the number of flags present. The shift and splice may also include selecting sample-by-sample based on quality flags, where unflagged samples are selected over flagged samples. Iterative depth matching may also include resampling merged datasets to a consistent depth grid (e.g., 0.5 ft spacing), interpolating, and preserving for missing intervals, as at 278. Iterative depth matching may further include updating the reference dataset by utilizing the merged dataset as the new reference, and repeating one or more of the previous steps discussed above, as at 280.
[0083] The depth matching and splicing process may also include generating a final merged dataset, as at 282. Generating the final merged dataset may include replacing all log curves with canonical names, writing provenance metadata, or a combination thereof. The output from the depth matching and splicing may include one or more of a single composite dataset per well with maximum coverage, a shift table documenting depth adjustments per dataset, such as a dataset name, a depth shift applied (feet or meters), a correlation coefficient, an overlap interval, a confidence flag, or a combination thereof. The output may also include an updated flag dataset indicating data provenance per sample, a summary of statistics, and / or a quality control plot (e.g., side-by-side comparison of original vs. shifted logs in overlap zones). The summary of statistics may include one or more of an original coverage per dataset, a final merged coverage, overlap statistics, correlation quality metrics, or a combination thereof.
[0084] The process or methods 200 disclosed herein may include coal flagging, as at 284, which may not be zonation dependent. Coal flagging may automatically identify coal intervals using a pretrained ML classifier (coal vs. non-coal) complemented by rule-based log response patterns (e.g., low density, high resistivity, characteristic GR / NPHI signatures, etc.). Coal flagging may be conducted per or for each dataset (pre- or post-splicing) regardless of zonation availability. The inputs for coal flagging may include one or more of a spliced / merged dataset, harmonized datasets, log families: Gamma Ray, Bulk Density, Resistivity, and Neutron Porosity, coal detection thresholds (e.g., regional calibration): —RHOB<2.0 g / cc; —Resistivity>20 ohm·m, —GR<60 GAPI (for clean coal; higher for shaly coal); —NPHI>0.30 (high apparent porosity), or a combination thereof.
[0085] Coal flagging may include a machine-learning (ML) based detection, a rule-based detection, fusion and interval consolidation, flag creation, or a combination thereof. The ML-based detection may include loading pretrained coal classifier (trained on labeled coal / non-coal intervals with RHOB, Resistivity, GR, NPHI features and optional derived ratios), scoring each depth sample, and flagging if above or below a threshold value. It should be appreciated that classifier parameters and thresholds are configurable based on regional geology. Rule-based detection may include applying a multi-log criteria at each depth sample. For example, if RHOB<2.0 and Resistivity>20 and GR<60 and NPHI>0.30, then a flag as coal may be generated. Fusion and interval consolidation may include merging ML-based flags and rule-based flags, consolidating flagged samples into coal intervals (contiguous flagged zones), and applying a minimum thickness filter (e.g., discard flagged intervals<2 ft), or a combination thereof. Flag creation may include creating a coal flag variable (e.g., 1=coal, 0=not coal, −9999=missing). The output from coal flagging may include one or more of coal flags per depth sample, a coal interval list (top, bottom, thickness), summary statistics (e.g., total coal thickness, number of coal seams per well, etc.), or a combination thereof.
[0086] The processes or methods 200 disclosed herein may include a zonation-based log normalization, as at 286. Zonation-based log normalization may normalize log responses (e.g., Gamma Ray) within common geological zones across wells to account for regional variations and improve cross-well consistency. It should be appreciated that this step may be conducted if common zonation was previously identified. The inputs for the zonation-based log normalization may include one or more of spliced / merged datasets, harmonized datasets, common zonation (e.g., zone names and depth intervals per well, etc.), target log families for normalization (e.g., Gamma Ray, etc.), a normalization method (e.g., quantile-based scaling or linear shift / scaling).
[0087] The zonation-based log normalization may include a quantile-based normalization (e.g., per zone, per family). For example, for each zone and each log family, zone quantiles per well may be determined, reference quantiles may be determined, normalization may be applied, and log values may be updated. Determining zone quantiles per well may include, for each well, extracting log values within the zone depth interval, determining P10, P50, and / or P90 for the zone, and excluding flagged samples and missing values. Determining reference quantiles may include determining a respective median of the P10, P50, and / or P90 across all wells in the zone. Applying normalization may include, for each well, applying piecewise linear scaling withing the zone. For example, mapping a well's P10 to a reference P10, mapping a well's P50 to a reference P50, mapping a well's P90 to a reference P90, or a combination thereof. Applying normalization may also include using linear interpolation between quantile anchors, extrapolating linearly outside P10 and / or P90, or a combination thereof. Updating log values may include replacing original log values with normalized values within the zone, preserving normalized values in a new variable for traceability, or a combination thereof.
[0088] In addition to or in substitution of the zonation-based log normalization, the process or method may include a linear shift / scale normalization. The linear shift / scale normalization may include, for each zone and each log family, determining a zone mean and standard deviation per well, determining a reference mean and standard deviation (median across wells), and applying a linear transformation. The output from normalization may include normalized log curves per zone per well, normalization parameters per zone per well (e.g., original quantiles (P10, P50, P90), reference quantiles, scaling factors, etc.), summary statistics: normalization quality metrics (variance reduction, cross-well consistency improvement), QC plots: before / after cross-plots and zone-wise histograms, or a combination thereof.
[0089] The processes and methods 200 disclosed herein may also include generating an output, as at 288. The output may be results, such as final results, a summary, based on any of the foregoing, or the like, or any combination thereof. The output may be high-quality logs. For example, the process and methods 200 disclosed herein may provide the consumable log set (e.g., reconstructed, depth-shifted / spliced, and, if applicable, zone-normalized) ready for downstream petrophysics or modeling. The logs may include canonical aliases in final units, prediction flags, dataset flag lineage, a shift table, normalization parameters, or a combination thereof. The output may also include a QC report, dashboards, or a combination thereof. The inputs for generating the output may be or include any one or more of the previously detailed outputs, such as quality flags, processing outputs, prediction statistics, depth matching results, normalization results, or the like, or any combination thereof.
[0090] Generating the output may include generating a summary for each dataset, which may include, for each dataset, compiling a metadata summary including well name, dataset name, and acquisition method, a depth range (e.g., top, bottom, unit), family inventory (e.g., available families, coverage per family, etc.). Generating the output may also include generating a summary for each dataset, which may include, for each dataset, compiling quality metrics including out-of-range flagged samples (e.g., count, percentage per family), linear segment flagged samples (e.g., count, percentage per family), bad hole flagged samples (e.g., count, percentage, breakdown by rule), ML outlier flagged samples (e.g., count, percentage), an overall flag summary (combined percentage flagged), or a combination thereof. Generating the output may also include generating a summary for each dataset, which may include, for each dataset, compiling prediction statistics including families predicted, samples existing vs. missing vs. predicted per family, a model performance (R2, RMSE) per family, or a combination thereof. Generating the output may also include generating a summary for each dataset, which may include, for each dataset, compiling a harmonization summary including original curve names→canonical aliases, unit conversions applied, conversion failure count, or a combination thereof. Generating the output may also include generating a summary for each dataset, which may include, for each dataset, compiling depth matching including depth shift applied (feet / meters), correlation coefficient, overlap interval with reference, or a combination thereof.
[0091] Generating the output may also include generating a per-well summary, which may include, for each well, compiling a dataset inventory including a number of datasets processed, a total depth coverage (merged), acquisition methods present, or a combination thereof. Generating the output may also include generating a per-well summary, which may include, for each well, compiling aggregated quality metrics, total samples across all datasets, total flagged samples (percentage), breakdown by flag type (out-of-range, linear, bad hole, outlier), or a combination thereof. Generating the output may also include generating a per-well summary, which may include, for each well, compiling a log inventor including families available after harmonization, average coverage per family across datasets, or a combination thereof. Generating the output may also include generating a per-well summary, which may include, for each well, compiling a splicing summary including a number of datasets spliced, a final merged depth range, a shift table summary, or a combination thereof. Generating the output may also include generating a per-well summary, which may include, for each well, compiling a normalization summary including zones normalized, normalization quality metrics, or a combination thereof.
[0092] The summary reports may be in one or more formats, including an interactive dashboard including tabular summaries with drill-down capability, traffic light indicators for quality thresholds, links to detailed plots, or a combination thereof. The summary reports may also be in the form of exported reports including excel spreadsheets with summary tables, pdf reports with key statistics and plots, or a combination thereof. The summary reports may also be in the form of database properties including summary statistics written back to well / dataset properties in the database for querying. The outputs from generating the output may be or include, but are not limited to, one or more of per-dataset summary tables (metadata, quality, prediction, harmonization), per-well summary tables (aggregated statistics, inventory, splicing, normalization), an overall project summary (cross-well statistics, processing completion status), a documentation of an audit trail (e.g., provenance, processing parameters, timestamps), QC plots including log tracks showing original vs. harmonized vs. predicted curves with flag overlays, cross-plots (e.g., RHOB vs. NPHI colored by flags), depth matching correlation plots, normalization before / after comparisons, or a combination thereof.
[0093] The present disclosure may include one or more key innovations or features, including, but not limited to, one or more of an integrated metadata imputation pipeline, including, a novel two-tier unit imputation: quantile-overlap method with fallback to value-range mapping, an automated acquisition method detection from naming conventions, a well path classification from deviation measurements, or a combination thereof, a robust multi-rule quality control, including, combining physics-based rules (out-of-range, linear segments, bad hole) with unsupervised ML outlier detection, a per-run bad hole detection using caliper clustering, hierarchical flag merging for comprehensive quality assessment, or a combination thereof, a ML-Based Log Reconstruction, including, dataset-level LightGBM models with automatic feature combination selection, integration of quality flags to exclude training on unreliable data, specialized trajectory fitting for TVD / TVDSS using geometric relationships, or a combination thereof the one or more key innovations or features may also include, but are not limited to, one or more of an auto-correlation depth matching & splicing, including, priority-based dataset ranking for optimal reference selection, cross-correlation based shift computation with confidence scoring, linear segment removal in overlap zones to improve correlation accuracy, provenance tracking via dataset flags throughout splicing, or a combination thereof, zonation-based normalization, including, quantile-anchored piecewise linear scaling for cross-well consistency, skip mechanism when no common zonation exists, or a combination thereof, a comprehensive provenance and audit trail, including, every processing step records metadata, original values, and transformation parameters, conversion failure metrics quantify harmonization quality, summary statistics enable data-driven QC decisions, or a combination thereof.
[0094] The present disclosure may include computer-implemented methods and systems for automated quality control and reconstruction of wireline logging data. The present disclosure may include a metadata enrichment module employing quantile-overlap based unit imputation with fallback to value-range mapping, enabling robust unit inference without manual input. The present disclosure may also include a multi-tiered quality control framework combining physics-based rules (out-of-range, linear segment, bad hole) with unsupervised machine learning outlier detection, providing comprehensive anomaly identification. The present disclosure may also include a caliper-based run clustering method using K-means with elbow detection and median-difference refinement, enabling per-run quality control in merged datasets. The present disclosure may also include a hybrid machine learning and physics-based reconstruction module that: extracts depth-dependent trends from logs before training (logarithmic fitting for Bulk Density and Resistivity, negative logarithmic for Neutron Porosity), trains LightGBM models on detrended data for improved generalization, applies trained models only to high-quality intervals (non-flagged samples), automatically selects optimal feature combinations from available log families, restores extracted physical trends to predicted values after model inference, applies hard physical range limits per log family (RHOB: 1.6-3.0 g / cc, NPHI: −0.15-0.95, etc.), generates confidence intervals (lower / upper quantile predictions) for prediction uncertainty quantification, or a combination thereof. The present disclosure may also include an autocorrelation based depth matching algorithm employing cross-correlation shift computation with linear segment exclusion in overlap zones, enabling accurate dataset alignment and splicing. The present disclosure may also include a zonation-based normalization module using quantile-anchored piecewise linear scaling for cross-well log consistency, applicable only when common zonation is detected. The present disclosure may also include a comprehensive provenance tracking system recording transformations, conversions, predictions, and dataset contributions at each depth sample, ensuring full auditability and reproducibility.Exemplary Method for Well Log Data Quality Assurance
[0095] FIG. 3 illustrates a flow chart of an exemplary method 300 for processing well log data associated with one or more wells, according to one or more embodiments. An illustrative order of the method 300 is provided below; however, one or more portions of the method 300 may be performed in a different order, simultaneously, repeated, or omitted. At least a portion of the method 300 may be performed using a computing system.
[0096] The method 300 may include receiving input data including well log data associated with the one or more wells, as at 302. The well log data may include a plurality of datasets associated with the one or more wells, wherein each dataset of the plurality of datasets is associated with a respective well of the one or more wells. Each dataset of the plurality of datasets may include one or more log curves, respective log curve values for each log curve of the one or more log curves, a depth index, zonation data, trajectory, metadata, or a combination thereof. The respective metadata of each dataset of the plurality of datasets may include one or more of a dataset name, one or more acquisition methods, a respective well name, a curve name for each log curve of the one or more log curves, a respective family for each log curve of the one or more log curves, a respective unit name for each log curve of the one or more log curves, or a combination thereof. In one example, a respective acquisition method of the one or more acquisition methods may be associated with each log curve of the one or more log curves. The respective acquisition method may be a wireline method, a logging while drilling method, or an unknown method. The depth index may include a true vertical depth (TVD) or a true vertical depth sub-sea (TVDSS). In one example, the trajectory may include a combination of an X-offset, a Y-offset, or a combination thereof. In another example, the trajectory may include a combination of measured depth (MD), inclination, and azimuth. The log curve values of each dataset may include a plurality of datapoints. The plurality of datasets may include one or more system-generated datasets. The respective zonation data for each dataset of the plurality of datasets may include one or more identified zones, a respective zonation tag for each identified zone of the one or more identified zones, respective zone interval bounds for each identified zone of the one or more identified zones, or a combination thereof. The respective zone interval bounds may include an upper bound and a lower bound.
[0097] The method 300 may include selecting one or more datasets from the plurality of datasets to generate a collection of initial processing datasets. The collection of initial processing datasets may be selected based on user-defined inputs. The collection of initial processing datasets may include the one or more selected datasets from the plurality of datasets.
[0098] The method 300 may also include generating a common zonation list based on the respective zonation data of each dataset of the collection of initial processing datasets. Generating the common zonation list may include identifying two or more datasets of the collection of initial processing datasets including the same respective zonation data using set intersection. Generating the common zonation list may also include generating the common zonation list based on the two or more datasets. The common zonation list may include the two or more datasets.
[0099] The method 300 may also include enriching the respective metadata of each dataset of the collection of initial processing datasets using a metadata check and enrichment process to produce enriched metadata. Enriching the respective metadata of each dataset of the collection of initial processing datasets may include generating a respective acquisition method label for each dataset of the collection of initial processing datasets. Generating the respective acquisition method label for each dataset of the collection of initial processing datasets may include parsing the respective dataset name of each dataset of the collection of initial processing datasets. Parsing the respective dataset name may include matching and tokenizing the respective dataset name of each dataset of the collection of initial processing datasets. Generating the respective acquisition method label for each dataset of the collection of initial processing datasets may also include extracting one or more tool mnemonics from each dataset of the collection of initial processing datasets to produce one or more extracted tool mnemonics. A respective tool mnemonic of the one or more tool mnemonics may be extracted from each log curve of the one or more log curves. Generating the respective acquisition method label for each dataset of the collection of initial processing datasets may also include associating each extracted tool mnemonic of the one or more extracted tool mnemonics with the respective acquisition method of the one or more acquisition methods of each dataset of the collection of initial processing datasets using a curated dictionary. Associating each extracted tool mnemonic of the one or more extracted tool mnemonics with the respective acquisition method may include mapping each extracted tool mnemonic of the one or more extracted tool mnemonics with the respective acquisition method of the one or more acquisition methods using the curated dictionary to produce one or more mapped extracted mnemonic tools. Associating each extracted tool mnemonic of the one or more extracted tool mnemonics with the respective acquisition method may also include assigning the one or more extracted tool mnemonics with the respective acquisition methods of each dataset of the collection of initial processing datasets based on the one or more mapped extracted mnemonic tools to produce one or more assigned tool mnemonics. Associating each extracted tool mnemonic of the one or more extracted tool mnemonics with the respective acquisition method may also include generating a respective acquisition method label for each dataset of the collection of initial processing datasets based on the one or more assigned tool mnemonics.
[0100] Enriching the respective metadata of each dataset of the collection of initial processing datasets may also include resolving the respective family associated with each log curve of the one or more log curves using the curated dictionary to map aliases of the respective family to a standard family. Resolving the respective family using the curated dictionary may include validating the respective family associated with the one or more log curves, imputing the respective family associated with the one or more log curves, or a combination thereof.
[0101] Enriching the respective metadata of each dataset of the collection of initial processing datasets may also include resolving the respective unit name associated with each log curve of the one or more log curves using quantile overlap or family value-range mapping. Resolving the respective unit name may include imputing missing units for the one or more log curves using the quantile overlap and based on the respective family of each log curve of the one or more log curves. Imputing the missing units may include comparing quantile distributions with offset wells using the quantile overlap. Resolving the respective units may include imputing the missing units for the one or more log curves using the family value-range mapping based on respective medians of each log curve of the one or more log curves. Enriching the respective metadata of each dataset of the collection of initial processing datasets may also include determining a well path classification for each well of the one or more wells based on the trajectory and the depth index thereof. The well path classification may be vertical, deviated, highly deviated, horizontal, snake, or unknown.
[0102] The method 300 may also include determining a respective coverage fraction of the family of each log curve of the one or more log curves of the collection of initial processing datasets using a log inventory process. The log inventory process may include extracting, for the respective family of each log curve of the collection of initial processing datasets, intervals to produce extracted intervals. The log inventory process may also include normalizing respective units of the extracted intervals to produce unit-normalized intervals. The log inventory process may also include merging the unit-normalized intervals to produce merged intervals. The log inventory process may also include determining measured-depth (MD) reference intervals based on the merged intervals. The log inventory process may also include determining the respective coverage fraction based on the measured-depth reference intervals and the merged intervals.
[0103] The method 300 may also include harmonizing the plurality of datasets using a curated dictionary to produce harmonized datasets, as at 304. Generating the harmonized datasets may include harmonizing the respective log curve names and the respective units of the datasets of the collection of initial processing datasets using the curated dictionary to produce the harmonized datasets.
[0104] The method 300 may also include conducting a quality control process based on the harmonized datasets to produce one or more flags, as at 306. Conducting the quality control process comprises one or more of generating one or more physics-based outlier flags based on the harmonized datasets, generating one or more ML outlier flags using an unsupervised machine-learning outlier detection process based on the harmonized datasets and the physics-based outlier flags, or a combination thereof.
[0105] Generating the one or more physics-based outlier flags may include conducting, on the harmonized datasets, an out-of-range detection process, a linear segment detection process, a bad hole flagging process, or a combination thereof. The out-of-range detection process may include identifying one or more out-of-range intervals within the harmonized datasets based on user-defined bounds. The out-of-range detection process may also include generating one or more out-of-range flags based on the one or more out-of-range intervals. The linear segment detection process may include identifying one or more linear intervals within the harmonized datasets using a sliding window and based on one or more statistical consistency checks. The one or more statistical consistency checks may include a slope-based check, a variance-based check, or a combination thereof. The linear segment detection process may also include generating one or more linear flags based on the one or more linear intervals. The bad hole flagging process may include identifying one or more bad hole intervals within the harmonized datasets based on a caliper rule, a bulk density correction rule, and a combined NPHI-RHOB-PEF rule. The combined NPHI-RHOB-PEF rule may be based on neutron porosity (NPHI), bulk density (RHOB), and photoelectric effect (PEF). The bad hole flagging process may also include generating one or more bad hole flags based on the one or more bad hole intervals. The one or more bad hole flags may indicate compromised borehole conditions. The one or more physics-based outlier flags may include the out-of-range flags, the linear flags, the bad hole flags, or a combination thereof.
[0106] Generating the one or more ML outlier flags may include training one or more unsupervised anomaly-detection models based on the harmonized datasets and the physics-based outlier flags. The one or more unsupervised anomaly-detection models may include one or more of one-class SVM, isolation forest, and local outlier factor (LOF). Generating the one or more ML outlier flags may also include identifying one or more multivariate intervals within the harmonized datasets using the one or more unsupervised anomaly-detection models. Generating the one or more ML outlier flags may also include generating the one or more ML outlier flags based on the multivariate intervals. The flags comprise the one or more physics-based outlier flags, the one or more ml outlier flags, or a combination thereof.
[0107] The method 300 may also include reconstructing respective log curves of the harmonized datasets using a supervised machine-learning reconstruction process and based on the flags to produce reconstructed logs, as at 308. The supervised machine-learning reconstruction process may include processing the harmonized datasets to generate processed datasets. Processing the harmonized datasets may include excluding respective intervals associated with the physics-based outlier flags and the ML outlier flags to generate the processed datasets. The supervised machine-learning reconstruction process may also include detrending the processed datasets using depth trend extraction to produce detrended datasets and an extracted depth trend. The supervised machine-learning reconstruction process may also include training one or more supervised machine-learning models based on the detrended datasets to produce one or more trained machine-learning models. The supervised machine-learning reconstruction process may also include generating initial ML predictions based on the detrended datasets and using the one or more trained machine-learning models. The initial ML predictions may exclude the extracted depth trend. The supervised machine-learning reconstruction process may also include combining or restoring the extracted depth trend with the initial ML predictions to produce intermediate ML predictions. The intermediate ML predictions include the extracted depth trend. The supervised machine-learning reconstruction process may also include processing the intermediate ML predictions based on user-defined thresholds to produce final ML predictions. The supervised machine-learning reconstruction process may also include performing a trajectory fitting for the respective true vertical depth (TVD) and the respective true vertical depth sub-sea (TVDSS) using the final ML predictions. The supervised machine-learning reconstruction process may also include reconstructing the respective log curves of the harmonized datasets based on the trajectory fitting.
[0108] The method 300 may also include depth matching and splicing two or more respective datasets of the harmonized datasets with one another to produce one or more matched datasets. The two or more respective datasets of the harmonized datasets may be associated with a single well of the one or more wells. Depth matching and splicing may include identifying an initial reference dataset within the harmonized datasets based on, for the respective family of each dataset of the harmonized datasets, the respective coverage fraction and respective quality ranking. The respective family may be Gamma Ray. The respective quality ranking may be based on the acquisition method, a non-missing sample count, the respective acquisition date, or a combination thereof. Depth matching and splicing may also include conducting the depth matching and splicing based on the initial reference dataset and the harmonized datasets to produce the one or more matched datasets.
[0109] The method 300 may also include generating normalized datasets based on the harmonized datasets, the one or more matched datasets, or a combination thereof using log normalization, as at 310. Generating the normalized datasets may include, for the respective family of each dataset of the harmonized dataset, the one or more matched datasets, or a combination thereof, conducting a quantile-based normalization, a linear scale normalization, or a combination thereof. The quantile-based normalization may be based on one or more median values of the respective family of each dataset of the harmonized datasets, the one or more matched datasets, or a combination thereof. The one or more median values of the respective family may include a P10 median value, a P50 median value, a P90 median value, or a combination thereof. The linear scale normalization may be based on a respective mean and a respective standard deviation of each dataset of the harmonized dataset, the one or more matched datasets, or a combination thereof. Generating the normalized datasets may further include performing an unsupervised agglomerative clustering on each dataset of the harmonized datasets and the one or more matched datasets to generate a first cluster of datasets and a second cluster of datasets. The first cluster of datasets may include relatively more datasets than then second cluster of datasets. The unsupervised agglomerative clustering may be based on a P10, a P50, a P90, of each dataset of the harmonized datasets and the one or more matched datasets. Generating the normalized datasets may further include generating the normalized datasets based on the first cluster of datasets and using the quantile-based normalization, the linear scale normalization, or a combination thereof. The quantile-based normalization may further be based on the respective identified zones of each dataset of the harmonized dataset, the one or more matched datasets, or a combination thereof.
[0110] The method 300 may further include generating an output, as at 312. The output may be based on one or more of the harmonized datasets, the matched datasets, the normalized datasets, the reconstructed logs, the physics-based outlier flags, the ML outlier flags, or a combination. The output may include a summary based on one or more of the reconstructed datasets, the matched datasets, the normalized datasets, the physics-based outlier flags, the ml outlier flags, or a combination. The output may include a processed well log data.
[0111] The method 300 may also include performing an action in response to generating the output. The action may include generating and / or transmitting a signal that recommends, instructs, or causes a physical action to occur. The physical action may include one or more of displaying the output, optimizing a trajectory of a wellbore drilling operation, conducting drilling operations, conducting an exploratory operation, designing a production strategy, designing a hydraulic fracturing strategy, conducting risk assessments, or any combination thereof.Log Quality Check Process and Data Selection
[0112] The methods disclosed herein may include various graphical user interfaces of a software tool implementing the log quality check process. In some embodiments, the user interface for data selection may include a plurality of datasets. For example, in some embodiments, each dataset may contain data corresponding to a different well (e.g., dataset 1 containing data corresponding to Well-1, dataset 2 containing data corresponding to Well-2, and so on). When in the user interface for data selection, the user may select which dataset of the plurality of datasets to be worked on.Zonation Selection
[0113] The methods disclosed herein may include user interfaces for zonation selection. In some embodiments, the user interface for zonation selection may allow a user to select a common zonation from a list of zones. In some embodiments, a user may sort the list and / or filter the list to make locating the desired zonation easier.Metadata Check
[0114] The methods disclosed herein may include user interfaces for a metadata check. In some embodiments, the metadata check module may verify whether the values for acquisition method, family, unit, and well path are defined for each dataset. For example, in some embodiments, in the case of Well 1, Dataset 1, if the family and unit attributes are properly defined, the user interface may highlight those attributes with a green color code or denote those attributes with the term “existing” or the like to indicate their presence. If a value is classified as “imputed” or yellow color coded in the user interface, it may mean that attribute is not explicitly defined in the project, but can be inferred using certain applied logic. For example, in some embodiments, the acquisition method may be determined by applying pattern matching to the dataset name. If a value is classified as “missing” or red color coded in the user interface, it may mean that the attributes may not be defined or may not be mapped using the applied logic.Auto-Processing and Summary
[0115] The methods disclosed herein may include user interfaces for auto-processing and summary. In some embodiments, the user interface for auto-processing and summary may illustrate steps of the automated well log quality check process and the percentage completion of each step. For example, in some embodiments, the user interface for auto-processing and summary may include, but is not be limited to, one or more of the steps of “Log Inventory Check,”“Data Harmonization,”“Out of Range Detection,”“Linear Segment Detection,”“Bad Hole Flagging,” Unsupervised Outlier Detection,”“Coal Flagging,”“Log Prediction,” Log Normalization,”“Depth Matching and Splicing,” or the like, or any combination thereof. In some embodiments, the percentage completion of each step may be illustrated visually both by a numeric percentage and a completion bar.Log Inventory Summary
[0116] The methods disclosed herein may include user interfaces for log inventory summary. In some embodiments, the user interface for log inventory summary may illustrate data completeness of various variables associated with one or more wells. In some embodiments, which variables to check data completeness for may be selected by the user. For example, in some embodiments, the user interface for log inventory summary may illustrate the level of data completeness for, but is not limited to, one or more of the variables of bulk density, gamma ray, measured depth, neutron porosity, resistivity, true vertical depth, or the like, or any combination thereof. In some embodiments, the level of data completeness may be illustrated as a numeric percentage and / or a completion bar. In some embodiments, the user interface for log inventory summary may also include an average coverage indicator (e.g., numeric percentage) configured to illustrate the overall average data completeness for all selected variables for a specific well.Data Harmonization
[0117] The methods disclosed herein may include user interfaces for data harmonization summary. In some embodiments, the Data Harmonization module may be designed to select the best curve for each required family and standardize the variables into consistent aliases and units for subsequent processes, such as log reconstruction and depth matching. The user interface may include a table having a “Curve Selected” column and a “Original Unit” column which indicate which curve from each family has been chosen and the original unit for user's reference. The user interface may include a percentage indicating if there are any issues with unit conversion. For example, a 0% indicator means there are no issues with unit conversion. If, however, there are issues, such as incompatible units for conversion or missing units, the user interface may display a percentage greater than zero indicating how many curves within a dataset or well could not be converted.Out-of-Range Summary
[0118] The methods disclosed herein may include user interfaces for out-of-range summary. In some embodiments, the user interface for out-of-range summary may list individual well datasets and a percentage of out-of-range values for each of one or more selected variables (e.g., bulk density, gamma ray, neutron porosity, resistivity, etc.) from each of the datasets. In some embodiments, the user interface may illustrate an overall out-of-range percentage for the selected variables in the dataset. In some embodiments, user interface for out-of-range summary may provide a visualization one or more of the out-of-range percentages such as, but not limited to, a bar or column graph, an X-Y plot, a scatter plot, or other visualization.Linear Segmentation Summary
[0119] The methods disclosed herein may include user interfaces for linear segmentation summary. In some embodiments, the user interface for linear segmentation summary may list individual datasets and a percentage of linear outliers for each of one or more selected variables (e.g., bulk density, gamma ray, neutron porosity, resistivity, etc.) from each of the datasets. In some embodiments, the user interface may illustrate an overall linear outlier percentage for the selected variables in the dataset. In some embodiments, user interface for linear segmentation summary may provide a visualization one or more of the linear outlier percentages such as, but not limited to, a bar or column graph, an X-Y plot, a scatter plot, or other visualization.Bad Hole Flagging Summary
[0120] The methods disclosed herein may include user interfaces for bad hole flagging summary. In some embodiments, the user interface for bad hole flagging summary may list individual well datasets and an indicator regarding whether the section of the borehole represented by the data is considered “bad hole,” meaning the section has irregularities such as uneven diameter, rough surfaces, or other conditions that may affect the accuracy of the logging measurements taken in that area. In some embodiments, the user interface for bad hole flagging summary may provide a visualization such as, but not limited to, a bar or column graph, an X-Y plot, a scatter plot, or other visualization to evaluate whether the bad hole predictions are reasonable and to check for any evident bias in the algorithm towards a particular variable (e.g., bulk density, neutron porosity, etc.).Unsupervised Outlier Detection Summary
[0121] The methods disclosed herein may include user interfaces for unsupervised outlier detection summary. In some embodiments, the user interface for unsupervised outlier detection summary may list individual well datasets and a percentage of unsupervised outlier values for each of the datasets. In some embodiments, the user interface may illustrate an overall unsupervised outlier percentage in the dataset(s). In some embodiments, user interface for unsupervised outlier detection summary may provide a visualization of the unsupervised outlier percentage such as, but not limited to, a bar or column graph, an X-Y plot, a scatter plot, or other visualization.Log Reconstruction Summary
[0122] The methods disclosed herein may include user interfaces for log reconstruction summary. In some embodiments, the user interface for log reconstruction summary may list individual well datasets and a percentage of data in the dataset that has been reconstructed. In some embodiments, the user interface may illustrate an overall total predicted percentage. In some embodiments, the user interface may illustrate a percentage of data that has been reconstructed (e.g., filled) such as bad hole filled, outlier fille, missing intervals filled, etc. In some embodiments, user interface may provide a visualization of the data reconstructed such as, but not limited to, a bar or column graph, an X-Y plot, a scatter plot, or other visualization.Log Normalization Summary
[0123] The methods disclosed herein may include user interfaces for log normalization summary. In some embodiments, the user interface for log normalization summary may illustrate the normalization of one or more datasets. For example, in some embodiments, the user interface for log normalization summary may allow the user to select one or more datasets or zones and a variable (e.g., gamma ray) from the one or more datasets to be displayed. In some embodiments, the user interface for log normalization summary may present a summary of the data, such as for example, in an output table. In some embodiments, the user interface for log normalization summary may present the data both before and after normalization. For example, in some embodiments, the user interface for log normalization summary may present a first histogram of the variable from the one or more datasets before normalization and a second histogram of the variable after normalization for comparison.Exemplary Method for Well Log Data Quality Assurance
[0124] The present disclosure may include a method for well log data quality assurance. The method may be configured in a variety of ways. An illustrative order of the method is provided below; however, one or more portions of the method may be performed in a different order, simultaneously, repeated, or omitted. At least a portion of the method may be performed with a computing system (described below).
[0125] In some embodiments, the example method may include selecting one or more well datasets to perform automated quality assurance on. In some embodiments, selecting one or more well datasets may include selecting one or more wells or one or more zones containing each containing one or more wells.
[0126] In some embodiments, the method may include performing a metadata check on the data in the selected data bases. In some embodiments, the metadata check may include verifying whether the values for, for example, acquisition method, family, unit, and well path are defined for each dataset
[0127] In some embodiments, the method may include performing an automated well log data quality check on the one or more selected data sets. The automated well log data quality check may be configured in a variety of ways. In some embodiments, the automated well log data quality check may include one or more of performing a log inventor check, performing data harmonization, performing out-of-range detection, performing linear segment detection, performing bad hole flagging, performing unsupervised outlier detection, performing log prediction, performing log normalization, performing log reconstruction, and performing depth matching and splicing.
[0128] In some embodiments, the method may include displaying via a user-interface a summary of an automated well log data quality check. In some embodiments, the summary may include displaying the data completeness (e.g., percentage complete) of various variables associated with one or more wells. In some embodiments, the summary may include displaying, via a user-interface, a summary of the well log data quality check by displaying at least one of the data completeness of one or more variables associated with one or more wells or an interactive visualization of the processed data. The method may include displaying an interactive visualization for each of the quality check steps or modules (e.g., data harmonization, out-of-range detection, linear segment detection, bad hole flagging, etc.).Exemplary Computing System
[0129] In some embodiments, the methods of the present disclosure may be executed by a computing system. FIG. 4 illustrates an example of such a computing system 400, in accordance with some embodiments. The computing system 400 may include a computer or computer system 401A, which may be an individual computer system 401A or an arrangement of distributed computer systems. The computer system 401A includes one or more analysis modules 402 that are configured to perform various tasks according to some embodiments, such as one or more methods disclosed herein. To perform these various tasks, the analysis module 402 executes independently, or in coordination with, one or more processors 404, which is (or are) connected to one or more storage media 406. The processor(s) 404 is (or are) also connected to a network interface 407 to allow the computer system 401A to communicate over a data network 409 with one or more additional computer systems and / or computing systems, such as 401B, 401C, and / or 401D (note that computer systems 401B, 401C and / or 401D may or may not share the same architecture as computer system 401A, and may be located in different physical locations, e.g., computer systems 401A and 401B may be located in a processing facility, while in communication with one or more computer systems such as 401C and / or 401D that are located in one or more data centers, and / or located in varying countries on different continents).
[0130] A processor may include a microprocessor, microcontroller, processor module or subsystem, programmable integrated circuit, programmable gate array, or another control or computing device.
[0131] The storage media 406 may be implemented as one or more computer-readable or machine-readable storage media. Note that while in the example embodiment of FIG. 4 storage media 406 is depicted as within computer system 401A, in some embodiments, storage media 406 may be distributed within and / or across multiple internal and / or external enclosures of computing system 401A and / or additional computing systems. Storage media 406 may include one or more different forms of memory including semiconductor memory devices such as dynamic or static random access memories (DRAMs or SRAMs), erasable and programmable read-only memories (EPROMs), electrically erasable and programmable read-only memories (EEPROMs) and flash memories, magnetic disks such as fixed, floppy and removable disks, other magnetic media including tape, optical media such as compact disks (CDs) or digital video disks (DVDs), BLURAY® disks, or other types of optical storage, or other types of storage devices. Note that the instructions discussed above may be provided on one computer-readable or machine-readable storage medium, or may be provided on multiple computer-readable or machine-readable storage media distributed in a large system having possibly plural nodes. Such computer-readable or machine-readable storage medium or media is (are) considered to be part of an article (or article of manufacture). An article or article of manufacture may refer to any manufactured single component or multiple components. The storage medium or media may be located either in the machine running the machine-readable instructions, or located at a remote site from which machine-readable instructions may be downloaded over a network for execution.
[0132] In some embodiments, computing system 400 contains one or more method execution module(s) 408. In the example of computing system 400, computer system 401A includes the method execution module 408. In some embodiments, a single method execution module may be used to perform some aspects of one or more embodiments of the methods disclosed herein. In other embodiments, a plurality of method execution modules may be used to perform some aspects of methods herein.
[0133] It should be appreciated that computing system 400 is merely one example of a computing system, and that computing system 400 may have more or fewer components than shown, may combine additional components not depicted in the example embodiment of FIG. 4, and / or computing system 400 may have a different configuration or arrangement of the components depicted in FIG. 4. The various components shown in FIG. 4 may be implemented in hardware, software, or a combination of both hardware and software, including one or more signal processing and / or application specific integrated circuits.
[0134] Further, the steps in the processing methods described herein may be implemented by running one or more functional modules in information processing apparatus such as general-purpose processors or application specific chips, such as ASICs, FPGAs, PLDs, or other appropriate devices. These modules, combinations of these modules, and / or their combination with general hardware are included within the scope of the present disclosure.
[0135] Computational interpretations, models, and / or other interpretation aids may be refined in an iterative fashion; this concept is applicable to the methods discussed herein. This may include use of feedback loops executed on an algorithmic basis, such as at a computing device (e.g., computing system 400, FIG. 4), and / or through manual control by a user who may make determinations regarding whether a given step, action, template, model, or set of curves has become sufficiently accurate for the evaluation of the subsurface three-dimensional geologic formation under consideration.
[0136] The foregoing description, for purpose of explanation, has been described with reference to specific embodiments. However, the illustrative discussions above are not intended to be exhaustive or limiting to the precise forms disclosed. Many modifications and variations are possible in view of the above teachings. Moreover, the order in which the elements of the methods described herein are illustrated and described may be re-arranged, and / or two or more elements may occur simultaneously. The embodiments were chosen and described in order to best explain the principles of the disclosure and its practical applications, to thereby enable others skilled in the art to best utilize the disclosed embodiments and various embodiments with various modifications as are suited to the particular use contemplated.
Claims
1. A method for processing well log data, the method comprising:obtaining input data comprising the well log data, wherein the well log data comprises a plurality of datasets, wherein each dataset of the plurality of datasets comprises one or more log curves and is associated with a respective well of one or more wells;harmonizing the plurality of datasets using a curated dictionary to produce harmonized datasets;reconstructing the one or more log curves of each harmonized dataset of the harmonized datasets to produce reconstructed logs using a supervised machine-learning (ML) model;generating normalized datasets based on the harmonized datasets and using log normalization;generating an output based on the reconstructed logs and the normalized datasets.
2. The method of claim 1, further comprising conducting a quality control process based on the harmonized datasets to produce one or more flags, wherein the output is generated based on the reconstructed logs, the normalized datasets, and the one or more flags.
3. The method of claim 2, wherein conducting the quality control process comprises generating one or more physics-based outlier flags based on the harmonized datasets, and wherein generating the one or more physics-based outlier flags comprises using one or more of an out-of-range detection process, a linear segment detection process, a bad hole flagging process, or a combination thereof.
4. The method of claim 3, wherein conducting the quality control process further comprises generating one or more ML outlier flags using an unsupervised ML outlier detection process and based on the harmonized datasets and the physics-based outlier flags.
5. The method of claim 4, wherein generating the one or more ML outlier flags using the unsupervised ML outlier detection process comprises:training an unsupervised anomaly-detection model based on the harmonized datasets and the one or more physics-based outlier flags;identifying one or more multivariate intervals within the harmonized datasets using the unsupervised anomaly-detection model; andgenerating the one or more ML outlier flags based on the one or more multivariate intervals.
6. The method of claim 2, wherein reconstructing the one or more log curves of each harmonized dataset of the harmonized datasets comprises:excluding respective intervals of the harmonized datasets associated with the one or more flags to generate processed datasets;detrending the processed datasets using depth trend extraction to produce detrended datasets and an extracted depth trend;training the supervised ML model based on the detrended dataset to produce a trained ML model;generating initial ML predictions based on the detrended datasets and using the trained ML model;combining the extracted depth trend with the initial ML predictions to produce intermediate ML predictions;processing the intermediate ML predictions based on user-defined thresholds to produce final ML predictions;performing a trajectory fitting for each harmonized dataset of the harmonized datasets using the final ML predictions; andreconstructing the one or more log curves of each harmonized dataset of the harmonized datasets based on the respective trajectory fitting thereof.
7. The method of claim 1, wherein generating the normalized datasets comprises, conducting, for each harmonized dataset of the harmonized datasets, one or more of a quantile-based normalization, a linear scale normalization, or a combination thereof.
8. The method of claim 1, wherein each dataset of the plurality of datasets further comprises metadata, and wherein the method further comprises enriching the respective metadata of each dataset of the plurality of datasets.
9. The method of claim 1, wherein the output comprises a processed well log data, a summary based on the reconstructed logs and the harmonized datasets, or a combination thereof.
10. The method of claim 1, further comprising performing an action in response to generating the output, wherein the action comprises generating or transmitting a signal that recommends, instructs, or causes a physical action to occur, wherein the physical action comprises one or more of displaying the output, optimizing a trajectory of a wellbore drilling operation, conducting drilling operations, conducting an exploratory operation, designing a production strategy, designing a hydraulic fracturing strategy, conducting risk assessments, or any combination thereof.
11. A computing system, comprising:one or more processors; anda memory system comprising one or more non-transitory computer-readable media storing instructions that, when executed by at least one of the one or more processors, cause the computing system to perform operations for processing well log data, the operations comprising:obtaining input data comprising the well log data, wherein the well log data comprises a plurality of datasets, wherein each dataset of the plurality of datasets comprises one or more log curves and is associated with a respective well of one or more wells;harmonizing the plurality of datasets using a curated dictionary to produce harmonized datasets;reconstructing the one or more log curves of each harmonized dataset of the harmonized datasets to produce reconstructed logs using a supervised machine-learning (ML) model;generating normalized datasets based on the harmonized datasets and using log normalization;generating an output based on the reconstructed logs and the normalized datasets.
12. The computing system of claim 11, further comprising:determining a respective coverage fraction of each log curve of the one or more log curves of each dataset of the plurality of datasets using a log inventory process;depth matching and splicing two or more harmonized datasets of the harmonized datasets with one another based on the respective coverage fraction of each log curve of the one or more log curves thereof to produce one or more matched datasets; andgenerating the output based on the one or more matched datasets, the normalized datasets, and the reconstructed logs.
13. The computing system of claim 11, further comprising conducting a quality control process based on the harmonized datasets to produce one or more flags, wherein conducting the quality control process comprises:generating one or more physics-based outlier flags based on the harmonized datasets; andgenerating one or more ML outlier flags using an unsupervised ML outlier detection process and based on the harmonized datasets and the one or more physics-based outlier flags,wherein the output is generated based on the reconstructed logs, the normalized datasets, and the one or more flags.
14. The computing system of claim 13, wherein generating the one or more physics-based outlier flags comprises using one or more of an out-of-range detection process, a linear segment detection process, a bad hole flagging process, or a combination thereof, wherein:the out-of-range detection process comprises:identifying one or more out-of-range intervals within the harmonized dataset based on user-defined bounds; andgenerating one or more out-of-range flags based on the one or more out-of-range intervals;the linear segment detection process comprising:identifying one or more linear intervals within the harmonized datasets based on one or more statistical consistency checks, wherein the one or more statistical consistency checks comprise a slope-based check, a variance-based check, or a combination thereof, andgenerating one or more linear flags based on the one or more linear intervals; andthe bad hole flagging process comprises:identifying one or more bad hole intervals within the harmonized datasets based on a Caliper Rule, a Bulk Density Correction Rule, and a combined NPHI-RHOB-PEF rule, wherein the combined NPHI-RHOB-PEF rule is based on Neutron Porosity (NPHI), Bulk Density (RHOB), and Photoelectric Effect (PEF) of the one or more log curves; andgenerating one or more bad hole flags based on the one or more bad hole intervals.
15. The computing system of claim 14, wherein generating the one or more ML outlier flags using the unsupervised ML outlier detection process comprises:training an unsupervised anomaly-detection model based on the harmonized datasets and the physics-based outlier flags;identifying one or more multivariate intervals within the harmonized datasets using the unsupervised anomaly-detection model; andgenerating the one or more ML outlier flags based on the multivariate intervals.
16. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations for processing well log data, the operations comprising:obtaining input data comprising the well log data, wherein the well log data comprises a plurality of datasets, wherein each dataset of the plurality of datasets comprises one or more log curves and is associated with a respective well of one or more wells;harmonizing the plurality of datasets using a curated dictionary to produce harmonized datasets;reconstructing the one or more log curves of each harmonized dataset of the harmonized datasets to produce reconstructed logs using a supervised machine-learning (ML) model;generating normalized datasets based on the harmonized datasets and using log normalization;generating an output based on the reconstructed logs and the normalized datasets.
17. The non-transitory computer-readable medium of claim 16, further comprising:generating one or more coal flags and one or more non-coal flags using a pretrained ML model and based on the harmonized datasets; andgenerating the output based on the one or more coal flags, the one or more non-coal flags, the reconstructed logs, and the normalized datasets.
18. The non-transitory computer-readable medium of claim 17, wherein the pretrained ML model is trained on labeled coal intervals, labeled non-coal intervals, and a combination of gamma ray, bulk density, resistivity, and neutron porosity.
19. The non-transitory computer-readable medium of claim 16, wherein generating the normalized datasets comprises, conducting, for each harmonized dataset of the harmonized datasets, one or more of a quantile-based normalization, a linear scale normalization, or a combination thereof.
20. The non-transitory computer-readable medium of claim 16, wherein generating the normalized datasets further comprises performing an unsupervised agglomerative clustering on the harmonized datasets to generate a first cluster of harmonized datasets and a second cluster of harmonized datasets, wherein a number of harmonized datasets in the first cluster of harmonized datasets is greater than a number of harmonized datasets in the second cluster of harmonized datasets, and wherein the normalized datasets are generated based on the first cluster of harmonized datasets.