Automated source rock characteristics and class prediction
A computing platform with a graph database and machine learning engine transforms and filters geochemical data to accurately characterize and classify source rocks, addressing inefficiencies in existing methods and improving energy development workflows.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2026-03-12
AI Technical Summary
Source rock geochemical data is sparse and of variable quality, making it difficult to interpret without expert knowledge, and existing methods are time-consuming and prone to human and system biases, leading to inefficiencies in subsurface geological structure analysis for energy development.
A computing platform comprising a graph database, data processing system, and machine learning engine is used to transform and filter geochemical data, resolve discrepancies, and train a subterranean model for accurate source rock characterization and classification, leveraging curated bulk pyrolysis and maceral data.
Enables rapid and consistent source rock characterization and classification with high accuracy, reducing interpreter bias and enhancing energy development workflows by providing timely and consistent mechanisms for subsurface characterization.
Smart Images

Figure US2024045462_12032026_PF_FP_ABST
Abstract
Description
Attorney Docket No.: IS24.0141-WO-PCTAUTOMATED SOURCE ROCK CHARACTERISTICS AND CLASS PREDICTIONINTRODUCTION
[0001] The disclosed technology is directed to a modeling solution for determining rock characteristics and class prediction associated with a resource site.BACKGROUND
[0002] Because source rock geochemical data may be sparse and of variable quality, they may be difficult to interpret without expert knowledge and / or without expert systems associated with hydrocarbon geochemistry. In addition, geochemical data interpretation can be time consuming and subject to system and / or human biases and / or errors.
[0003] There is therefore a need to provide timely and / or efficient and / or consistent mechanisms to characterize and / or classify source rocks based on curated bulk geochemical data including bulk pyrolysis data and maceral data. In addition, there is a need to enhance or otherwise optimize energy development systems and / or workflows that are inefficient at processing and / or analyzing subsurface geological structures (e.g., rocks) in order to generate accurate model predictions and / or minimize uncertainties sin risk of charge evaluation associated with subsurface characterizations for energy development.SUMMARY
[0004] Disclosed are methods, systems, and computer programs that characterize and / or classify rocks at a resource site for energy development. According to an embodiment, a method for characterizing and / or classifying source rocks at a resource site comprises determining a computing platform for modeling source rocks, the computing platform including a database system, a data processing system, and a machine learning engine.
[0005] The method further comprises generating, using the database system, analyzed graph data. According to one embodiment, generating the analyzed graph data comprises: transforming a data schema indicating node type data and link type data into a graph database, the node type data and link type data being parameterized using geochemical data associated with a first resource site; creating, using the graph database, a graph data representationAttorney Docket No.: IS24.0141-WO-PCT indicating characterized or classified subsurface structures associated with the first resource site, and conditioning the graph database to generate the analyzed graph data.
[0006] The method further comprises: filtering, using the data processing system, the analyzed graph data based on vitrinite reflectance data and thereby generate trainable data; resolving, using the data processing system, data discrepancies within the trainable data and thereby generate resolved data; holistically enhancing, using the data processing system, the resolved data to be compatible with a plurality of subterranean structures and thereby generate training data; applying, using the machine learning engine, the training data to train a subterranean model and thereby generate a trained subterranean model, the subterranean model characterizing or classifying rock characteristics associated with the first resource site; and testing, using the machine learning engine, the trained subterranean model and thereby generate a prediction report indicating rock characteristics and classification of a source rock for the first resource site or a second resource site that is similar to, or distinct from the first resource site.
[0007] In other embodiments, a system and a computer program can include or execute the method described above. These and other implementations may each optionally include one or more of the following features.
[0008] The node type data comprises at least one of: a borehole node that stores borehole data including a borehole identifier data, borehole name data, borehole location data, and borehole province data; a sample node that stores source rock sample data including hydrocarbon shows data, pyrolysis data, and maceral group geochemistry data; an environment node that stores environment interpretation data of intervals along a well trajectory including depositional environment explanation data, geological basin data, and interval depth data; a lithology node that stores lithology interpretation data of intervals along the well trajectory including lithology explanation data and interval depth data; and a formation node that stores geological information data of intervals along the well trajectory including formation name data, geological age data, and interval depth data.
[0009] In some embodiments, the link type data comprises at least one of: a first computing structure that stores link type data between a first borehole relative to a sample derived from the first borehole; a second computing structure indicating a link between the first borehole and a second borehole and which stores data including geographical distance data between the first borehole and the second borehole; a third computing structure that links borehole data andAttomev Docket No.: IS24.0141-WO-PCT depositional environment data; a fourth computing structure that links borehole data and lithology data; and a fifth computing structure that links borehole data and formation data.
[0010] Transforming the data schema into a graph database comprises formatting datasets derived from the data schema into a table such that data rows of the table comprise different samples or boreholes while data columns of the table comprise features of each sample according to some embodiments.
[0011] Moreover, each feature comprised in the features of each sample can indicate geochemical data including one of: pyrolysis data including total organic carbon (TOC) data; SI parameter data; S2 parameter data; S3 parameter data; and Tmax parameter data.
[0012] In addition, each feature comprised in the features of each sample can indicate geochemical data including one of: inertinite data; liptinite data; and vitrinite data.
[0013] It is appreciated that nodes and links comprised in the graph database facilitate querying or filtering the graph database using graph query sentences aligned with semantic or syntactic query structures within a query language associated with the graph database.
[0014] It is further appreciated that the above method further comprises deriving the analyzed graph data from a quality control process including conditioning the graph database such that the graph database is used to generate a graph data representation. Moreover, conditioning the graph database can comprise determining data interactions of the graph representation to gain quick insights or establish data relationships between neighbor wells at the first resource site.
[0015] According to some embodiments, the quality control process may be used to validate or confirm an accuracy of the graph representation.
[0016] In some implementations, the trainable data may be generated based on filtering out core sample data indicating low thermal maturity levels within the analyzed graph data.
[0017] In addition, holistically enhancing the resolved data can comprise converting local feature data associated with the resolved data to global feature data.
[0018] Moreover, holistically enhancing the resolved data can comprise converting a plurality of data elements comprised in the resolved data or imputed data into data modes of the training data that drive properties of the subterranean model
[0019] According to one embodiment, converting the plurality of data elements comprised in the resolved data comprises: converting geological basin identifier data to basin type data including at least one of a rift basin type, a passive basin type, a foreland basin type, or a back-Attorney Docket No.: IS24.0141-WO-PCT arc basin type; converting stratigraphy identifier data to geological age data; and converting hydrocarbon shows data from string data to sequential integer data.
[0020] Furthermore, applying the training data to train the subterranean model comprises one or more of: grid search training or optimization computing operations on the subterranean model using the training data; dimension reduction training or optimization computing operations on the subterranean model using the training data; category data encoding training or optimization computing operations on the subterranean model using the training data; missing data imputation training or optimization computing operations on the subterranean model using the training data; and hyperparameter fine-tune training or optimization operations on the subterranean model using the training data.
[0021] In some cases, testing the trained subterranean model comprises: receiving geochemical data associated with the first resource site or the second resource site; applying the geochemical data to one or more parameters of the trained subterranean model; and generating the prediction report in response to applying the geochemical data.
[0022] According to one embodiment, the geochemical data is derived from lab analysis of core data extracted at the first resource site or the second resource site.
[0023] In some implementations, the one or more parameters of the trained subterranean model comprises at least one of: a total organic carbon (TOC) parameter associated with the geochemical data; a bulk pyrolysis parameter associated with the geochemical data, the bulk pyrolysis parameter comprising one of an SI parameter, an S2 parameter, an S3 parameter, and Tmax parameter; and a maceral composition parameter.
[0024] Furthermore, the prediction report may be analyzed by the machine learning engine to: confirm probability data associated with prediction report or associated with result data of the prediction report and thereby ensure that the prediction report is within probabilistic thresholds associated the first resource site or the second resource site; and analyze the result data by mapping or correlating the result data to data structures comprised in the graph representation.
[0025] The above method further comprises configuring equipment based on the prediction report.Attorney Docket No.: IS24.0141-WO-PCTBRIEF DESCRIPTION OF THE DRAWINGS
[0026] The disclosure is illustrated by way of example, and not by way of limitation in the figures of the accompanying drawings in which like reference numerals are used to refer to similar elements. It is emphasized that various features may not be drawn to scale and the dimensions of various features may be arbitrarily increased or reduced for clarity of discussion.
[0027] FIG. 1A depicts an exemplary computing platform used to implement the disclosed technology.
[0028] FIG. IB depicts an exemplary graph representation that is generated by a graph query subsystem to visualize a depositional environment distribution in neighbor wells at a resource site.
[0029] FIG. 1C depicts an exemplary visualization indicating data conversions associated with a subterranean model.
[0030] FIG. 2 depicts a cross-sectional view of a resource site for which the process of FIGS. 4 and 5 may be executed.
[0031] FIG. 3 depicts a networked system illustrating a communicative coupling of devices or systems associated with the resource site of FIG. 2.
[0032] FIG. 4 depicts an exemplary detailed workflow for characterizing and classifying source rocks at a resource site.
[0033] FIG. 5 depicts an exemplary flowchart for generating analyzed graph data.DETAILED DESCRIPTION
[0034] Reference will now be made in detail to embodiments, examples of which are illustrated in the accompanying drawings and figures. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the disclosed subject-matter. However, it will be apparent to one of ordinary skill in the art that the solutions disclosed may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0035] The disclosed systems and methods may be accomplished using interconnected devices and systems that obtain a plurality of data associated with various parameters of interest at a resource site. The workfl ows / flowcharts described in this disclosure, according to someAttorney Docket No.: IS24.0141-WO-PCT embodiments, implicate a new processing approach (e.g., hardware, special purpose processors, and specially programmed general-purpose processors) because such analyses are too complex and cannot be done by a person in the time available or at all. Thus, the described systems and methods are directed to tangible implementations or solutions to specific technological problems in developing natural resources such as oil, gas, water well industries, and other mineral exploration operations. More specifically, the systems and methods presently disclosed may be applicable to operations associated with stratigraphic analysis associated with a resource site.
[0036] Attention is now directed to methods, techniques, infrastructure, and workflows for operations that may be carried out at a resource site. Some operations in the processing procedures, methods, techniques, and workflows disclosed herein may be combined while the order of some operations may be changed. Some embodiments include an iterative refinement of one or more data models associated with the resource site via feedback loops executed by one or more computing device processors and / or through other control devices / mechanisms that make determinations regarding whether a given action, template, or resource data, etc., is sufficiently accurate.Overview
[0037] Geochemical characterizations and classifications of source rocks play a key role in hydrocarbon exploration and production. In particular, it is necessary to evaluate charge risks as well as prediction of hydrocarbon types using gravity data and / or oil-gas ratio data in combination with basin distribution data associated with a resource site. According to some embodiments, the foregoing data may be used to further support field development and production operations to inform the determination of fluid property distribution data.
[0038] Furthermore, geochemical data interpretation may be used to characterize and / or classify source rocks based on organic matter type data associated with the source rocks, depositional environment data associated with the source rocks, facies data associated with the source rocks, age data associated with the source rocks, thermal maturity data associated with the source rocks, and level of subsurface alteration data associated with the source rocks.
[0039] According to one embodiment, bulk pyrolysis data and maceral data measured from rock sample data (e.g., core sample data, cuttings sample data, sidewall core sample data,Attorney Docket No.: IS24.0141-WO-PCT outcrop sample data) can be used to characterize and / or classify hydrocarbon source rocks at a resource site. It is appreciated that the rock sample data may be derived from multiple rock samples per well such that the pyrolysis and / or maceral data is determined using the multiple rock samples instead of molecular data of the multiple rock samples according to some embodiments. Furthermore, the pyrolysis data and / or the maceral data may be used in energy or hydrocarbon exploration workflows (e.g., basin modeling workflows).
[0040] The assessment of source rock characteristics and / or classification is needed to inform decisions in hydrocarbon exploration (including risk of charge evaluation) as well as support hydrocarbon production operations based on reservoir continuity data, asphaltenes presence data, and hydrocarbon distribution data. However, some methods to achieve this are hampered by data volumes, time constraints, and / or resource intensive limitations. Through the application of machine learning and data science techniques, source rock analysis may be automated, as disclosed, to allow rapid and consistent source rock characterization and classification.
[0041] According to one embodiment, curated source rock geochemical data (e.g., bulk pyrolysis data and maceral data) may be used to build a machine learning system that enables the prediction of critical parameters including depositional environment parameters, source rock potential parameters, kerogen type parameters, hydrogen index parameters, etc. Furthermore, it is appreciated that each of the foregoing parameters may be configured to have at least an 80% accuracy, or at least 85% accuracy, or at least 90% accuracy, or at least 95% accuracy, or at least 99% accuracy depending on the implementation.
[0042] According to one embodiment, the disclosed solution can be implemented using a combination of: a graph database system, a data processing system, and a machine learning engine. These aspects are discussed in conjunction with FIG. 1A.Platform Environment
[0043] FIG. 1A shows an exemplary computing platform 100a used to implement the disclosed technology. The computing platform 100a includes a graph database system 150, a data processing system 160, and a machine learning engine 170, all of which are communicatively coupled to each other via wired or wireless data and / or signal lines.Attorney Docket No.: IS24.0141-WO-PCT
[0044] According to one embodiment, the database system 150 may be configured or structured to store a graph database based on a defined schema. The defined schema may include a plurality of different nodes and links. In some implementations, the defined schema includes one or more of the following data elements or features:• Node Type 1 computing structure: a borehole node that stores borehole data including a borehole identifier (ID) data, borehole name data, borehole location data, and borehole province data.• Node Type 2 computing structure: a sample node that stores source rock sample data including hydrocarbon shows data, pyrolysis data, and maceral group geochemistry data.• Node Type 3 computing structure: an environment node that stores environment interpretation data of intervals along a well trajectory including depositional environment explanation data, geological basin data, and interval depth data which can be derived from a report associated with a well at the resource site.• Node Type 4 computing structure: a lithology node that stores lithology interpretation data of intervals along the well traj ectory including lithology explanation data and interval depth data.• Node Type 5 computing structure: a formation node that stores geological information data of intervals along the well trajectory including formation name data, geological age data, and interval depth data.• Link Type 1 computing structure: stores link data between a borehole (e g., first borehole) relative to a sample (e.g., a core sample derived from the first borehole); in exemplary embodiments, 1 borehole can contain or be associated with a plurality of samples; a link may be created by matching a node’s borehole ID data with the sample ID of the node in question.• Link Type 2 computing structure: a computing structure indicating a link between a first borehole and a second borehole and stores data including the geographical distance between the first borehole and the second borehole; this structure is created by matching identifier data between the first borehole and the second borehole.• Link Type 3 computing structure: a computing structure that links borehole data and depositional environment data; exemplary embodiments includes 1 borehole being mappedAttorney Docket No.: IS24.0141-WO-PCT to one or more depositional environments; this structure is created by matching a node’s borehole ID with its environment ID.• Link Type 4 computing structure: a computing structure that links borehole data and lithology data; exemplary embodiments include 1 borehole including one or more lithologies; this structure can be created by matching a node’s borehole ID with its lithology ID• Link Type 5 computing structure: a computing structure that links borehole data and formation data; exemplary implementations have 1 borehole containing a plurality of formations; the structure may be created by matching a node’s borehole ID with its formation ID.
[0045] According to one embodiment, the graph database system 150 can comprise a schema design subsystem 102 that can be used to create or update schema 103. A data import subsystem 104 of the database system 150 can be used to import schema 103 as well as transform schema 103 into a graph database 105. According to one embodiment, the data schema 103 can be used by the data import subsystem 104 to create nodes and links during the data import process. Furthermore, datasets (e.g., derived from the data schema 103) associated with the import process may be tabularly formatted into a table such that data rows of the table comprise different samples or boreholes while data columns of the table comprise features of each sample. According to one embodiment, each feature indicates a geochemical data type including: pyrolysis data including total organic carbon (TOC) data, an SI parameter data, an S2 parameter data, an S3 parameter data, a Tmax parameter data, and related parameter data; and maceral composition data including inertinite data, liptinite data, and vitrinite data. Once the data is successfully imported, the graph database is populated with nodes and links.
[0046] It is appreciated that the SI parameter data indicates an amount of free volatile hydrocarbons (e.g., gas and / or oil) thermally flushed from a rock sample at a first temperature (e.g., about 300°C) with units in milligrams of hydrocarbon per gram of rock or mg HC / g rock. The amount of free hydrocarbons obtained by this thermal extraction with or without any significant decomposition of kerogen can be measured by a flame ionization detector (FID) to generate the SI parameter data. The S2 parameter data indicates an amount of hydrocarbons generated through thermal cracking of nonvolatile organic matter during pyrolysis temperaturesAttorney Docket No.: IS24.0141-WO-PCT between a second and third temperature (e.g., between 300°C to 600°C) with units in milligrams of hydrocarbon per gram of rock or mg HC / g rock. In other embodiments, the S2 parameter data indicates an amount of hydrocarbons generated through thermal cracking of nonvolatile organic matter during pyrolysis temperatures between a fourth and fifth temperature (e.g., between 300°C to 850°C) with units in milligrams of hydrocarbon per gram of rock or mg HC / g rock using rock evaluation instrument. Furthermore, all the kerogen (e.g., complex waxy mixture of hydrocarbon compounds which forms the primary organic component of oil shale) capable of generating petroleum may be converted into hydrocarbons, which are quantified by the energy decisions to give the S2 parameter data. In particular, the S2 parameter data can provide an indication of the quantity of hydrocarbons that a given subterranean structure (e.g., subsurface rocks) has the potential of producing should burial and maturation continue for said subterranean structure.
[0047] According to one embodiment, the S3 parameter data indicates an amount of carbon dioxide (e.g., measured in milligrams CO2 per gram of rock) produced during pyrolysis of rock samples between a sixth temperature and a seventh temperature (e.g., between 300°C and 600°C) or between an eighth temperature and a ninth temperature (e.g., between 300°C and 850°C) using a rock evaluation instrument such as a thermal conductivity detector (TCD). It is appreciated that the S3 parameter data may provide an indication of the amount of oxygen in kerogen associated with a resource site under consideration and can be used to calculate or otherwise determine an oxygen index.
[0048] It is further appreciated that the Tmax parameter data can indicate temperature information (e.g., in °C) at which the maximum release of hydrocarbons from cracking of kerogen occurs during pyrolysis (e.g., top of S2 peak). In particular, the Tmax parameter data can provide an indication of the stage of maturation of organic matter in a subsurface.
[0049] Furthermore, a graph query subsystem 106 of the database system 150 can operate on the graph database 105 to generate or create a graph representation 107 (e.g., a graph data representation) associated with characterized and / or classified subsurface or subterranean structures such as rocks. In particular, the created nodes and links inside the graph database can be queried or filtered (e.g., filtering or querying the graph database) by the graph query subsystem 106 using graph query sentences aligned with a required grammar (e.g., semantic and / or syntactic query structures within a query language associated with the graph database)Attorney Docket No.: IS24.0141-WO-PCT of the graph database. After successful execution of a graph query, the filtered nodes and links can be indicated as the graph representation 107. The created graph representation 107 of source rock data can enable analysis of the entire dataset using the analysis subsystem 108. In particular, the analysis subsystem 108 of the graph database system 150 can be used to receive the graph representation 108 for further analysis (e.g., visual analysis, datapoint analysis, data location analysis, data threshold analysis, etc.). The biggest benefit of the graph representation 107 is that neighbor nodes data can be directly visualized and used as input for missing data and thereby enrich data associated with a given analysis of subsurface.
[0050] In the analysis stage, the quality of the graph database 105 can be checked using a quality control process. In particular, the analysis (e.g., visual analysis, datapoint analysis, data location analysis, data threshold analysis, etc.) of the graph representation 107 can comprise a data conditioning computing operation that is based on the graph query subsystem determining data interactions of the graph representation (e.g., a graph data representation derived from the graph database) to gain quick insights or establish data relationships between neighbor wells of a resource site associated with the graph representation. For example, if a data relationship is to be determined between a neighbor borehole and its depositional environment, a graph query can be executed using the graph query subsystem 106 to generate the graph representation 107. FIG. IB illustrates an exemplary graph representation 100b that is generated by the graph query subsystem 106 to visualize a depositional environment distribution in neighbor wells at a resource site. As seen in this figure, the graph representation 100b includes a plurality of nodes 182 interconnected by a plurality of branches 184 (e.g., links). According to one embodiment, the analysis executed by the analysis subsystem 108 may comprise a quality control process used to validate or otherwise confirm the accuracy of the graph representation 107. Once the quality control process is completed, data from the graph database system 150 is transmitted to other systems of the computing platform 100 such as the data processing system and the machine learning engine.
[0051] Turning back to FIG. 1A, the data processing subsystem 160 can facilitate data filtering and / or generating training data 113 for the machine learning engine 170. For example, the data processing subsystem 160 can receive, using the data filter subsystem 110, analyzed data (e.g., analyzed graph data) associated with the graph representation 107 via the analysis subsystem 108. The data filter subsystem 110 can generate or update trainable data 111 using the analyzedAttorney Docket No.: IS24.0141-WO-PCT data from the analysis subsystem 108. Furthermore, the trainable data can be used by the data imputation subsystem 112 to impute / fill in data gaps or other missing data associated with the trainable data 111 and thereby feed the feature generalization subsystem 114 with data approximations used to generate the aforementioned training data 113. According to one embodiment, the generated training data 113 may be used by the machine learning engine 170 to develop models that characterize the subsurface of a resource site.
[0052] According to one embodiment, the data filter subsystem 110 is used to filter the data generated by the analysis subsystem 108 and thereby generate filtered data associated with immature or early mature subterranean formations corresponding to, for example, a vitrinite reflectance (e.g., based on vitrinite reflectance data of a resource site associated with the graph database system 150) equivalent less than about 0.6% Ro. In particular, reliable filtered data can be generated using immature up to early mature core samples (e.g., as defined by the vitrinite reflectance value of 0.6% Ro). This filtered data can be used in developing a model (e.g., subterranean model) that predicts source rock characteristics and classifications. In some cases, using higher maturity samples (e.g., greater than 0.6% Ro) derived from mature subterranean formations leads to unreliable model prediction results. This limitation can be due to loss of chemical feature of the subterranean formation because of increasing burial data and / or temperature data and / or maturity data associated with the mature subterranean formation. Hence, for source rock classifications, core samples (e.g., core sample data comprised in the analyzed graph data) showing or indicating low thermal maturity levels are filtered out by, for example, the data filter subsystem 110, to generate the trainable data 111. Furthermore, machine learning training inputs and targets features can be kept or maintained or in other words relevant features in addition to the maturity level can be kept as part of the trainable data.
[0053] To reiterate, after generating the trainable data 111, missing data imputation may be executed on said trainable data by the data imputation subsystem 112. This missing data imputation may include:(i) copying data from neighboring samples (e.g., core samples) in the same borehole;(ii) determining that samples within the same stratigraphy formation unit share geochemistry data; andAttorney Docket No.: IS24.0141-WO-PCT(iii) copying or interpolating data from surrounding boreholes’ samples within a given geography distance.
[0054] For example, because geochemical information needed for source rock characterization may be missing from a set of samples (e.g., core samples) collected from the same depth interval and showing similar features in terms of lithology, it is necessary to copy such data using points (i)-(iii). In particular, geochemical data from one or more neighboring samples with similar geochemical properties relative a core sample under consideration may copied into or otherwise inserted into missing or faulty data points within the sample under consideration.
[0055] According to some embodiments, if samples (e.g., core samples) are in the same well and within a given geological zone, the missing data can be imputed directly by the data imputation subsystem 112 and thereby generate imputed data from which the training data 113 is derived. Specifically, distance relationships between core samples can serve as criteria that allow data imputation where one or more samples from the same well and within the same geological zone can be used to fill data gaps associated with one or more samples derived from the same well and / or geological zone. The geological zone can refer to a geological age interval between two or more core samples associated with a subsurface region according to some embodiments.
[0056] If samples are in different wells and different geological zones, additional rules or logic are applied to decide the workflow for copying or imputing the data. It is appreciated that the data imputation subsystem 112 is configured to resolve data discrepancies within the trainable data and thereby generate resolved data, which is holistically enhanced or generalized, using the feature generalization subsystem 114, to be compatible with a plurality of subterranean structures. In particular, the holistically enhanced or generalized resolved data constitutes the training data.
[0057] The feature generalization subsystem 114 can convert local feature data comprised in imputed data generated by the data imputation subsystem 112 into global feature data. In particular, a local feature refers to a feature which is available only for some of the samples in the training dataset whereas the global feature is defined (e.g., has a value) for all the samples in said dataset. For instance, because the graph database 105 is built or otherwise developed using source rock data in one or more basins associated with a resource site, the resulting modelAttorney Docket No.: IS24.0141-WO-PCT generated by the machine learning engine 170 can be used for a basin with specific features at the resource site or basins at a different resource site with features similar to the basins at the resource site from which the model was developed. In particular, after running the feature generalization process using the feature generalization subsystem 114, the model generated therefrom can be used to make predictions associated with subsurface structures at a similar and / or dissimilar resource site relative to the resource site from which the model was developed. According to one embodiment, the feature generalization subsystem can convert a plurality of imputed data (e.g., data elements comprised in resolved data or imputed data) obtained from the data imputation subsystem 112 into data modes of the training data 113 that drive properties of the model developed by the machine learning engine 170 using the training data 113. This conversion can include one or more of:• converting geological basin identifier data (e g., basin name data) to basin type data including at least one of a rift basin type, a passive basin type, a foreland basin type, or a back-arc basin type;• converting stratigraphy identifier data (e.g., stratigraphy name data) to geological age data; and• converting hydrocarbon shows data from string data to sequential integer data.
[0058] For example, the conversion process can be used to convert basin identifier data for a given basin to a specific basin type as indicated in FIG. 1C. In particular, FIG. 1C shows an exemplary visualization 100c indicating the conversion of basin identifier data 192 into basin type data 194. This figure also shows the conversion of stratigraphy identifier data 196 to geological age data 198.
[0059] Turning back to FIG. 1A, the model training engine 116 may receive the training data 113 and apply same to a subterranean model 117 which characterizes and / or classifies rock characteristics associated with a resource site. According to one embodiment, training / testing and / or optimization operations may be leveraged by the model training engine 116 in training the subterranean model. For example, the model training engine 116 may apply, based on the training data, one or more of grid search training and / or optimization computing operations, dimension reduction training and / or optimization computing operations, category data encoding training and / or optimization computing operations, missing data imputation trainingAttorney Docket No.: IS24.0141-WO-PCT and / or optimization computing operations, and hyperparameter fine-tune training and / or optimization operations on the subterranean model 117. The output of this process is a generated trained subterranean model with a holistic evaluation data including model metrics data, confusion matrix data, variable importance data, and receiver operating characteristic (ROC) curve data, etc. For example, these model metrics form part of the machine learning aspects of the disclosed method and facilitate understanding how well the model (e.g., subterranean model) is performing. In particular, the goal here is to build a machine learning model that optimally generalizes and / or customizes subsurface features of interest and can perform well into the future using known or unknown data (e.g., training and / or testing data). To achieve generating an optimal model for such a purpose, there is a need to evaluate the model to gain insight into how the model is performing and thereby adapt or otherwise refine the training data or testing data and / or refine the model parameters to facilitate or otherwise approximate configurations that lead to optimal model performance. As such metrics such as confusion matrices, variable importance metrics, and ROC curve data metrics can be leveraged in assessing performance data of a developed model using the disclosed techniques.
[0060] According to one embodiment, the trained subterranean model may be deployed by the model deployment engine 118 for usage in a prediction application 119. In particular, the prediction application 119 may comprise a user interface within which is imbedded the trained subterranean model. When looking for source rock characteristics and classifications, raw geochemical data comprising laboratory measurement data including total organic carbon (TOC), bulk pyrolysis data derived from bulk pyrolysis, such as, SI, S2, S3 and Tmax) and / or maceral composition data may be entered into input fields of the user interface associated with the prediction application 119. It is appreciated that in response to receiving the raw geochemical data, the prediction application 119 may be used by the model prediction engine 120 to predict the source rock characteristics and classification of a source rock or other subterranean structural properties associated with a resource site from which the geochemical data is derived. In response to the model prediction engine predicting the source rock characteristics and / or classification, the prediction application 119 may generate a prediction report including textual data and / or image data including multi-dimensional image data comprised in a prediction report. In particular, the prediction report indicates rockAttorney Docket No.: IS24.0141-WO-PCT characteristics and classification of a source rock for a resource site similar to or distinct from the resource site from which the raw geochemical data is derived.
[0061] According to one embodiment, the raw geochemical data may be derived from core samples and / or other sensor measurements at the resource site. In other embodiments, the raw geochemical data may be automatically analyzed to generate the laboratory measurement data which in turn may be automatically fed to the prediction application 119 seamlessly for usage by the model prediction engine 120. The prediction analysis engine 122 may then analyze the prediction results generated by the model prediction engine 120 by, for example, confirming or checking probability data associated with the prediction (e.g., result data comprised in a prediction report) to ensure that the predictions are within probabilistic thresholds associated the resource site in question. The prediction analysis engine 122 may also analyze the prediction results generated by the model prediction engine 120 by, for example, mapping and / or correlating the prediction results to data structures comprised in the graph representation 107. If the results (e.g., result data in the prediction report) of the prediction are incorrect or show low accuracy, the schema design subsystem 102 may automatically redesign the database schema 103 to enrich features and / or continuously grow the dataset required to make the subterranean model 117 optimal or performant.
[0062] It is appreciated that the above processes and systems or engines beneficially enable quickly and / or efficiently determining and validating information associated with:• source rock geochemical data exploitation for geoscientists without extensive petroleum geochemistry expertise and experience;• consistent source rock characteristics and classification (also referred to as class) assessment thereby avoiding interpreter bias due to the use of machine learning, training, and validation computing processors;• interpreting incomplete data sets using a data imputation subsystem; and• quickly processing of large datasets based on the disclosed automatic computing operations.
[0063] It is further appreciated that the disclosed solution can be used by upstream hydrocarbon development operations to reduce the risk of charge in exploration by improving the source rock analysis in areas (e.g., mature and / or immature areas) of a resource site with available data.Attorney Docket No.: IS24.0141-WO-PCTIn a sedimentary basin, hydrocarbons may be generated by organic rich rocks (e.g., also called source rocks). The identification of the presence and characteristics of these source rocks allow evaluating the chance of finding hydrocarbons can be referred to as “risk of charge.” The “charge” in this case comprises hydrocarbons generated and expelled from the source rocks to the reservoir rock.
[0064] In frontier areas where well data is scarce or absent, the disclosed technology can also be used to infer or otherwise predict source rock characteristics and class from analogues. Field development requiring a good understanding of vertical and lateral fluid property variations for continuity assessment or production allocation purposes may also beneficiate from a deeper understanding of the source rock characteristics derived using the disclosed techniques. Laboratories producing geochemical data from, for example, core samples from a resource site may also use this solution to provide interpretation data associated with source rock characteristics and class on top of the raw data from the core samples.Resource Site
[0065] FIG. 2 shows a cross-sectional view of a resource site 200 for which the process of FIG. 1 may be executed. While the illustrated resource site 200 represents a subterranean formation, the resource site, according to some embodiments, may be below water bodies such as oceans, seas, lakes, ponds, wetlands, rivers, or other marine environments.
[0066] According to one embodiment, various measurement tools capable of sensing one or more resource site data such as seismic two-way travel time, density, resistivity, production rate, etc., of a subterranean formation and / or geological formations may be provided at the resource site. As an example, wireline tools may be used to obtain measurement information related to geological attributes (e.g., geological attributes of a wellbore and / or reservoir) including geophysical and / or chemical information. For example, the chemical information may include chemical information associated with the subsurface and / or chemical information associated with the surface / above ground areas of the resource site 200.
[0067] In some embodiments, various sensors may be located at various locations around the resource site 200 to monitor and collect data and / or core samples for executing the process of FIGS. 4Attorney Docket No.: IS24.0141-WO-PCT
[0068] Part, or all, of the resource site 200 may be on land, on water, or below water. In addition, while a resource site 200 is depicted, the technology described herein may be used with any combination of one or more resource sites (e.g., multiple oil fields or multiple wellsites, one or more saline aquifers, one or more depleted oil / gas fields, etc ), one or more processing facilities, etc. As can be seen in FIG. 2, the resource site 200 may have data acquisition tools 202a, 202b, 202c, and 202d positioned at various locations within the resource site 200. The subterranean structure 204 may have a plurality of geological formations 206a- 206d. As shown, this structure may have several formations or layers, including a shale layer 206a, a carbonate layer 206b, a shale layer 206c, and a sand layer 206d. A fault 207 may extend through the shale layer 206a and the carbonate layer 206b. The data acquisition tools, for example, may be adapted to take measurements and detect geophysical and / or chemical characteristics of the various formations shown.
[0069] While a specific subterranean formation with specific geological structures is depicted, it is appreciated that the resource site 200 may contain a variety of geological structures and / or formations, sometimes having extreme complexity. In some locations of a given geological structure, for example below a water line (e.g., aquifer) relative to the given geological structure, fluid may occupy pore spaces of the formations. Each of the measurement devices may be used to measure properties of the formations and / or other geological features. While each data acquisition tool is shown as being in specific locations in FIG. 2, it is appreciated that one or more types of measurement may be taken at one or more locations across one or more sources of the resource site 200 or other locations for comparison and / or analysis.
[0070] The data collected from various sources at the resource site 200 may be processed and / or evaluated and / or used as training data, and or used to generate high resolution result sets for characterizing a resource at the resource site, and / or used for generating resource models, etc. In one embodiment, the core sample data and / or data collected by a set of sensors at the resource site may include data associated with the number of wells of a first reservoir or second reservoir at the resource site, data associated with the number of grid cells of the first or second reservoir, data associated with the average permeability of the first or second reservoir, data associated with the production duration history (e.g., number of years of production) of the first reservoir or second, etc.Attorney Docket No.: IS24.0141-WO-PCT
[0071] Data acquisition tool 202a is illustrated as a measurement truck, which may comprise devices or sensors that take measurements of the subsurface through sound vibrations such as, but not limited to, seismic measurements. Drilling tool 202b may include a downhole sensor adapted to perform logging while drilling (LWD) data collection. The wireline tool 202c may include a downhole sensor deployed in a wellbore or borehole. Production tool 202d may be deployed from a production unit or Christmas tree into a completed wellbore. Examples of resource site data that may be measured include weight on bit, torque on bit, subterranean pressures e.g., underground fluid pressure), temperatures, flow rates, compositions, rotary speed, particle count, voltages, currents, and / or other parameters of operations as further discussed below.
[0072] Sensors may be positioned about the resource site to collect data relating to various resource site operations, such as sensors deployed by the data acquisition tools 202. The sensor may include any type of sensor such as a metrology sensor (e.g., temperature, humidity), an automation enabling sensor, an operational sensor (e.g., pressure sensor, H2S sensor, thermometer, depth, tension), evaluation sensors, which can be used for acquiring data regarding the formation, wellbore, formation fluid / gas, wellbore fluid, gas / oil / water comprised in the formation / wellbore fluid, or any other suitable sensor. For example, the sensors may include accelerometers, flow rate sensors, pressure transducers, electromagnetic sensors, acoustic sensors, temperature sensors, chemical agent detection sensors, nuclear sensor, and / or any additional suitable sensors.
[0073] In one embodiment, the data captured by the one or sensors may be used to characterize, or otherwise generate one or more parameter values for a high-resolution result set used to, for example, label or configure a machine learning (ML) engine, a resource model as the case may require. In other embodiments, test data or synthetic data may also be used in developing the ML engine or resource model (e.g., a subterranean model) via one or more parameterization / labeling operations such as those discussed in association with FIG. 4.
[0074] Evaluation sensors may be featured in downhole tools such as tools 202b-202d and may include for instance electromagnetic, acoustic, nuclear, and optic sensors. Examples of tools including evaluation sensors that can be used in the framework of the current method include electromagnetic tools including imaging sensors such as FMI™ or QuantaGeo™ (mark of SLB, Houston, TX); induction sensors such as Rt Scanner™ (mark of SLB, Houston, TX),Attorney Docket No.: IS24.0141-WO-PCT multifrequency dielectric dispersion sensor such as Dielectric Scanner™ (mark of SLB, Houston, TX); acoustic tools including sonic sensors, such as Sonic Scanner™ (mark of SLB, Houston, TX) or ultrasonic sensors, such as pulse-echo sensor as in UBI™ or PowerEcho™ (marks of SLB, Houston, TX) or flexural sensors PowerFlex™ (mark of SLB, Houston, TX); nuclear sensors such as Litho Scanner™ (mark of SLB, Houston, TX) or nuclear magnetic resonance sensors; fluid sampling tools including fluid analysis sensors such as InSitu Fluid Analyzer ™ (mark of SLB, Houston, TX); distributed sensors including fiber optic. Such evaluation sensors may be used in particular for evaluating the formation in which the well is formed (z.e., determining petrophysical or geological properties of the formation), for verifying the integrity of the well (such as casing or cement properties) and / or analyzing the produced fluid (flow, type of fluid, etc.).
[0075] As shown, data acquisition tools 202a-202d may generate data plots or measurements 208a-208d, respectively. These data plots are depicted within the resource site 200 to demonstrate that data generated by some of the operations executed at the resource site 200. Data plots 208a-208c are examples of static data plots that may be generated by data acquisition tools 202a-202c, respectively. However, it is herein contemplated that data plots 208a-208c may also be data plots that may be generated and updated in real time. These measurements may be analyzed to better define properties of the formation(s) and / or determine the accuracy of the measurements and / or check for and compensate for measurement errors. The plots of each of the respective measurements may be aligned and / or scaled for comparison and verification purposes. In some embodiments, base data associated with the plots may be incorporated into site planning, modeling a test at the resource site 200. The respective measurements that can be taken may be any of the above.
[0076] Other data may also be collected, such as historical data of the resource site 200 and / or sites similar to the resource site 200, user inputs, information (e.g., economic information) associated with the resource site 200 and / or sites similar to the resource site 200, and / or other measurement data and other parameters of interest. Similar measurements may also be used to measure changes in formation aspects over time.
[0077] Computer facilities such as those discussed in association with FIG. 3 may be positioned at various locations about the resource site 200 (e.g., a surface unit) and / or at remote locations. A surface unit (e.g., one or more terminals 320) may be used to communicate withAttorney Docket No.: IS24.0141-WO-PCT the onsite tools and / or offsite operations, as well as with other surface or downhole sensors. The surface unit may be capable of sending commands to the oil field equipment / systems, and receiving data therefrom. The surface unit may also collect data generated during production operations and can produce output data, which may be stored or transmitted for further processing.
[0078] The data collected by sensors may be used alone or in combination with other data. The data may be collected in one or more databases and / or transmitted on or offsite. The data may be historical data, real time data, or combinations thereof. The real time data may be used in real time, or stored for later use. The data may also be combined with historical data or other inputs for further analysis or for modeling purposes to optimize production processes at the resource site 200. In one embodiment, the data is stored in separate databases, or combined into a single database.Network System
[0079] FIG. 3 shows a high-level networked system diagram illustrating a communicative coupling of devices or systems associated with the resource site 200 as described in FIG. 2. The system shown in the figure may include a set of processors 302a, 302b, and 302c for executing one or more processes discussed herein. The set of processors 302 may be electrically coupled to one or more servers (eg., computing systems) including memory 306a, 306b, and 306c that may store for example, program data, databases, and other forms of data. Each server of the one or more servers may also include one or more communication devices 308a, 308b, and 308c. The set of servers may provide a cloud-computing platform 310. In one embodiment, the set of servers includes different computing devices that are situated in different locations and may be scalable based on the needs and workflows associated with the resource site 200. The communication devices of each server may enable the servers to communicate with each other through a local or global network such as an Internet network. In some embodiments, the servers may be arranged as a town 312, which may provide a private or local cloud service for users. A town may be advantageous in remote locations with poor connectivity. Additionally, a town may be beneficial in scenarios with large networks where security may be of concern. A town in such large network embodiments can facilitate implementation of a private network within such large networks. The town may interface with other towns or a larger cloud network,Attorney Docket No.: IS24.0141-WO-PCT which may also communicate over public communication links. Note that cloud-computing platform 310 may include a private network and / or portions of public networks. In some cases, a cloud-computing platform 310 may include remote storage and / or other application processing capabilities.
[0080] The system of FIG. 3 may also include one or more user terminals 314a and 314b each including at least a processor to execute programs, a memory (e.g., 316a and 316b) for storing data, a communication device and one or more user interfaces and devices that enable the user to receive, view, and transmit information. In one embodiment, the user terminals 314a and 314b is a computing system having interfaces and devices including keyboards, touchscreens, display screens, speakers, microphones, a mouse, styluses, etc. The user terminals 314 may be communicatively coupled to the one or more servers of the cloud-computing platform 310. The user terminals 314 may be client terminals or expert terminals, enabling collaboration between clients and experts through the system of FIG. 3.
[0081] The system of FIG. 3 may also include at least one or more resource sites 200 having, for example, a set of terminals 320, each including at least a processor, a memory, and a communication device for communicating with other devices communicatively coupled to the cloud-computing platform 310. The resource site 200 may also have a set of sensors (e.g., one or more sensors described in association with FIG. 2) or sensor interfaces 322a and 322b communicatively coupled to the set of terminals 320 and / or directly coupled to the cloudcomputing platform 310. In some embodiments, data collected by the set of sensors / sensor interfaces 322a and 322b may be processed to generate a one or more resource models (e.g., reservoir models) or one or more resolved data sets used to generate the resource model which may be displayed on a user interface associated with the set of terminals 320, and / or displayed on user interfaces associated with the set of servers of the cloud computing platform 310, and / or displayed on user interfaces of the user terminals 314. Furthermore, various equipment / devices discussed in association with the resource site 200 may also be communicatively coupled to the set of terminals 320 and or communicatively coupled directly to the cloud-computing platform 310. The equipment and sensors may also include one or more communication device(s) that may communicate with the set of terminals 320 to receive orders / instructions locally and / or remotely from the resource site 200 and also send statuses / updates to other terminals such as the user terminals 314.Attorney Docket No.: IS24.0141-WO-PCT
[0082] The system of FIG. 3 may also include one or more client servers 324 including a processor, memory, and communication device. For communication purposes, the client servers 324 may be communicatively coupled to the cloud-computing platform 310, and / or to the user terminals 314a and 314b, and / or to the set of terminals 320 at the resource site 200 and / or to sensors at the oil field, and / or to other equipment at the resource site 200.
[0083] A processor, as discussed with reference to the system of FIG. 3, may include a microprocessor, a graphical processing unit (GPU), a microcontroller, a processor module or subsystem, a programmable integrated circuit, a programmable gate array, or another control or computing device.
[0084] The memory / storage media discussed above in association with FIG. 3 can be implemented as one or more computer-readable or machine-readable storage media that are non-transitory. In some embodiments, storage media may be distributed within and / or across multiple internal and / or external enclosures of a computing system and / or additional computing systems. Storage media may include one or more different forms of memory including semiconductor memory devices such as dynamic or static random access memories (DRAMs or SRAMs), erasable and programmable read-only memories (EPROMs), electrically erasable and programmable read-only memories (EEPROMs) and flash memories; magnetic disks such as fixed, floppy and removable disks; other magnetic media including tape; optical media such as compact disks (CDs) or digital video disks (DVDs), BluRays or any other type of optical media; or other types of storage devices. “Non-transitory” computer readable medium refers to the medium itself (i.e., tangible, not a signal) and not data storage persistency (e.g., RAM vs. ROM).
[0085] Note that instructions can be provided on one computer-readable or machine-readable storage medium, or alternatively, can be provided on multiple computer-readable or machine- readable storage media distributed in a large system having possibly plural nodes and / or non- transitory storage means. Such computer-readable or machine-readable storage medium or media is (are) considered to be part of an article (or article of manufacture). The storage medium or media can be located either in a computer system running the machine-readable instructions, or located at a remote site from which machine-readable instructions can be downloaded over a network for execution.Attorney Docket No.: IS24.0141-WO-PCT
[0086] It is appreciated that the described system of FIG. 3 is an example that may have more or fewer components than shown, may combine additional components, and / or may have a different configuration or arrangement of the components. The various components shown may be implemented in hardware, software, or a combination of both, hardware, and software, including one or more data processing and / or application specific integrated circuits.
[0087] Further, the steps in FIG. 4 described below may be implemented by running one or more functional modules in an information processing apparatus such as general-purpose processors or application specific chips, such as ASICs, FPGAs, PLDs, GPUs or other appropriate devices associated with the system of FIG. 3. For example, the flowchart of FIG. 4 below may be executed using a data engine or a data processing module (e.g., computing module) stored in memory 306a, 306b, or 306c such that the data engine / data processing module includes instructions that are executed by the one or more processors such as processors 302a, 302b, or 302c as the case may be. The various modules of FIG. 3, combinations of these modules, and / or their combination with general hardware are included within the scope of protection of the disclosure. While one or more computing processors e.g., processors 302a, 302b, or 302c) may be described as executing steps associated with, for example, FIG. 4, the one or more computing device processors may be associated with the cloud-based computing platform 310 and may be located at one location or distributed across multiple locations. In one embodiment, the one or more computing device processors may also be associated with other systems of FIG. 3 other than the cloud-computing platform 310.
[0088] In some embodiments, a computing system is provided that includes at least one processor, at least one memory, and one or more programs stored in the at least one memory, such that the programs comprise instructions, which when executed by the at least one processor, are configured to perform any method disclosed herein.
[0089] In some embodiments, a computer readable storage medium is provided, which has stored therein one or more programs, the one or more programs including instructions, which when executed by a processor, cause the processor to perform any method disclosed herein. In some embodiments, a computing system is provided that includes at least one processor, at least one memory, and one or more programs stored in the at least one memory for performing any method disclosed herein. In some embodiments, an information processing apparatus for use in a computing system is provided for performing any method disclosed herein.Attorney Docket No.: IS24.0141-WO-PCTExemplary Flowcharts
[0090] FIG. 4 shows an exemplary detailed workflow 400 for characterizing and classifying source rocks at a resource site. It is appreciated that a data engine stored in a memory device may cause a computer processor to execute the various processing stages of the workflow 400. For example, the disclosed techniques may be implemented as a data engine of a computing platform associated with a geological software tool such that the data engine enables optimally executing modeling subterranean structures at a resource site.
[0091] At block 402, the data engine determines a computing platform for modeling source rocks. The computing platform, according to one embodiment, includes a database system, a data processing system, and a machine learning engine.
[0092] At block 404, the data engine generates, using the database system, analyzed graph data. The process of generating the analyzed graph data is further discussed in association with FIG.5.
[0093] At block 406, the data engine filters, using the data processing system, the analyzed graph data based on vitrinite reflectance data and thereby generate trainable data.
[0094] The data engine may further resolve, using the data processing system at block 408, data discrepancies within the trainable data and thereby generate resolved data.
[0095] In one embodiment, the data engine holistically enhances, using the data processing system at block 410, the resolved data to be compatible with a plurality of subterranean structures and thereby generate training data.
[0096] Turing to block 408, the data engine may apply, using the machine learning engine, the training data to train a subterranean model and thereby generate a trained subterranean model. According to one embodiment, the subterranean model characterizes or classifies rock characteristics associated with the first resource site.
[0097] The data engine may be further used to test, using the machine learning engine at block 410, the trained subterranean model and thereby generate a prediction report indicating rock characteristics and classification of a source rock for the first resource site or a second resource site that is similar to, or distinct from the first resource site.
[0098] FIG. 5 shows an exemplary flowchart 500 for generating analyzed graph data. At block 502, the data engine transforms a data schema indicating node type data and link type data intoAttorney Docket No.: IS24.0141-WO-PCT a graph database. The node type data and link type data can be parameterized using geochemical data associated with a first resource site.
[0099] At block 504, the data engine creates, using the graph database, a graph data representation indicating characterized or classified subsurface structures associated with the first resource site.
[0100] At block 506, the data engine conditions the graph database to generate the analyzed graph data.
[0101] In other embodiments, a system and a computer program can include or execute the method described above. These and other implementations may each optionally include one or more of the following features.
[0102] The node type data comprises at least one of a borehole node that stores borehole data including a borehole identifier data, borehole name data, borehole location data, and borehole province data; a sample node that stores source rock sample data including hydrocarbon shows data, pyrolysis data, and maceral group geochemistry data; an environment node that stores environment interpretation data of intervals along a well trajectory including depositional environment explanation data, geological basin data, and interval depth data; a lithology node that stores lithology interpretation data of intervals along the well trajectory including lithology explanation data and interval depth data; and a formation node that stores geological information data of intervals along the well trajectory including formation name data, geological age data, and interval depth data.
[0103] In some embodiments, the link type data comprises at least one of: a first computing structure that stores link type data between a first borehole relative to a sample derived from the first borehole; a second computing structure indicating a link between the first borehole and a second borehole and which stores data including geographical distance data between the first borehole and the second borehole; a third computing structure that links borehole data and depositional environment data; a fourth computing structure that links borehole data and lithology data; and a fifth computing structure that links borehole data and formation data.
[0104] Transforming the data schema into a graph database comprises formatting datasets derived from the data schema into a table such that data rows of the table comprise different samples or boreholes while data columns of the table comprise features of each sample according to some embodiments.Attorney Docket No.: IS24.0141-WO-PCT
[0105] Moreover, each feature comprised in the features of each sample can indicate geochemical data including one of: pyrolysis data including total organic carbon (TOC) data; SI parameter data; S2 parameter data; S3 parameter data; and Tmax parameter data.
[0106] In addition, each feature comprised in the features of each sample can indicate geochemical data including one of: inertinite data; liptinite data; and vitrinite data.
[0107] It is appreciated that nodes and links comprised in the graph database facilitate querying or filtering the graph database using graph query sentences aligned with semantic or syntactic query structures within a query language associated with the graph database.
[0108] That disclosed method may further comprise deriving the analyzed data from a quality control process including conditioning the graph database such that the graph database is used to generate a graph data representation. Moreover, conditioning the graph database can comprise determining data interactions of the graph representation to gain quick insights or establish data relationships between neighbor wells at the first resource site.
[0109] According to some embodiments, the quality control process may be used to validate or confirm an accuracy of the graph representation.
[0110] In some implementations, the trainable data may be generated based on filtering out core sample data indicating low thermal maturity levels within the analyzed graph data.
[0111] In addition, holistically enhancing the resolved data can comprise converting local feature data associated with the resolved data to global feature data.
[0112] Moreover, holistically enhancing the resolved data can comprise converting a plurality of data elements comprised in the resolved data or imputed data into data modes of the training data that drive properties of the subterranean model
[0113] According to one embodiment, converting the plurality of data elements comprised in the resolved data comprises: converting geological basin identifier data to basin type data including at least one of a rift basin type, a passive basin type, a foreland basin type, or a back- arc basin type; converting stratigraphy identifier data to geological age data; and converting hydrocarbon shows data from string data to sequential integer data.
[0114] Furthermore, applying the training data to train the subterranean model comprises one or more of: grid search training or optimization computing operations on the subterranean model using the training data; dimension reduction training or optimization computing operations on the subterranean model using the training data; category data encoding trainingAttorney Docket No.: IS24.0141-WO-PCT or optimization computing operations on the subterranean model using the training data; missing data imputation training or optimization computing operations on the subterranean model using the training data; and hyperparameter fine-tune training or optimization operations on the subterranean model using the training data.
[0115] In some cases, testing the trained subterranean model comprises: receiving geochemical data associated with the first resource site or the second resource site; applying the geochemical data to one or more parameters of the trained subterranean model; and generating the prediction report in response to applying the geochemical data.
[0116] According to one embodiment, the geochemical data is derived from lab analysis of core data extracted at the first resource site or the second resource site.
[0117] In some implementations, the one or more parameters of the trained subterranean model comprises at least one of: a total organic carbon (TOC) parameter associated with the geochemical data; a bulk pyrolysis parameter associated with the geochemical data, the bulk pyrolysis parameter comprising one of an SI parameter, an S2 parameter, an S3 parameter, and Tmax parameter; and a maceral composition parameter.
[0118] Furthermore, the prediction report may be analyzed by the machine learning engine to: confirm probability data associated with the prediction report or associated with result data of the prediction report and thereby ensure that the prediction report is within probabilistic thresholds associated with the first resource site or the second resource site; and analyze the result data by mapping or correlating the result data to data structures comprised in the graph representation.
[0119] According to one embodiment, the disclosed method comprises configuring equipment (e.g., equipment or systems associated with a resource site; equipment or systems associated with energy development) based on the prediction report. For example, the prediction report may be used for: well placement operations at the first resource site or the second resource site; equipment placement operations at the first resource site or the second resource site; surgically locating hydrocarbons at the first resource site or the second resource site; and configuring equipment associated with energy development at the first resource site or the second resource site. In particular, the prediction report can be automatically transmitted to stakeholders and / or to stakeholder systems and automatically used for at least one of: well placement operations at the first resource site or the second resource site; equipment placement operations at the firstAttorney Docket No.: IS24.0141-WO-PCT resource site or the second resource site; surgically locating hydrocarbons at the first resource site or the second resource site; and configuring equipment associated with energy development at the first resource site or the second resource site.
[0120] In some implementations, captured geological data (e g., using one or more logging equipment or tools referenced in association with FIG. 2 or other core sampling systems) comprising core samples extracted or otherwise retrieved from a resource site (e.g., the first resource site or the second resource site or a third resource site) can be analyzed to determine geochemical data used to train and / or test the subterranean model. For example, the core samples can comprise cylindrically shaped geological samples that are gathered or otherwise collected from a drilled hydrocarbon well using a coring tool at the lowermost end of a drill string associated with the coring tool. Once the core sample is retrieved, one or more geological samples from an inner part of the core sample may be used for geochemical analysis that avoids any contamination from a drilling fluid used during the coring process.
[0121] While any discussion of or citation to related art in this disclosure may or may not include some prior art references, such discussions are neither concessions nor acquiescence to the position that any given reference is prior art or analogous prior art.
[0122] The foregoing description, for purpose of explanation, has been described with reference to specific embodiments. However, the illustrative discussions above are not intended to be exhaustive or to limited to the precise forms disclosed. Many modifications and variations are possible in view of the above teachings. The embodiments were chosen and described in order to explain the principles and its practical applications, to thereby enable others skilled in the art to use various embodiments with various modifications as are suited to the particular use contemplated. It is appreciated that the term optimize / optimal and its variants (e.g., efficient or optimally) may simply indicate improving, rather than the ultimate form of 'perfection' or the like.
[0123] It will also be understood that, although the terms first, second, etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used to distinguish one element from another. For example, a first object or step could be termed a second object or step, and, similarly, a second object or step could be termed a first object or step, without departing from the scope. The first object or step, and the second objectAttorney Docket No.: IS24.0141-WO-PCT or step, are both objects or steps, respectively, but they are not to be considered the same object or step.
[0124] The terminology used in the description herein is for the purpose of describing particular embodiments and is not intended to be limiting. As used in the description and the appended claims, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and encompasses any possible combination of one or more of the associated listed items. It will be further understood that the terms “includes,” “including,” “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0125] As used herein, the term “if’ may be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” depending on the context.
[0126] Those with skill in the art will appreciate that while some terms in this disclosure may refer to absolutes, e.g., all source receiver traces, each of a plurality of objects, etc., the methods and techniques disclosed herein may also be performed on fewer than all of a given thing, e.g., performed on one or more components and / or performed on one or more source receiver traces. Accordingly, in instances in the disclosure where an absolute is used, the disclosure may also be interpreted to be referring to a subset.
Claims
Atorney Docket No.: IS24.0141-WO-PCTCLAIMSWhat is claimed is:
1. A method for characterizing and classifying source rocks at a resource site, the method comprising: determining a computing platform for modeling source rocks, the computing platform including a database system, a data processing system, and a machine learning engine; generating, using the database system, analyzed graph data by: transforming a data schema indicating node type data and link type data into a graph database, the node type data and link type data being parameterized using geochemical data associated with a first resource site; creating, using the graph database, a graph data representation indicating characterized or classified subsurface structures associated with the first resource site, and conditioning the graph database to generate the analyzed graph data; filtering, using the data processing system, the analyzed graph data based on vitrinite reflectance data and thereby generate trainable data; resolving, using the data processing system, data discrepancies within the trainable data and thereby generate resolved data; holistically enhancing, using the data processing system, the resolved data to be compatible with a plurality of subterranean structures and thereby generate training data; applying, using the machine learning engine, the training data to train a subterranean model and thereby generate a trained subterranean model, the subterranean model characterizing or classifying rock characteristics associated with the first resource site; and testing, using the machine learning engine, the trained subterranean model and thereby generate a prediction report indicating rock characteristics and classification of a source rock for the first resource site or a second resource site that is similar to, or distinct from the first resource site.
2. The method of Claim 1, wherein the node type data comprises at least one of:Attorney Docket No.: IS24.0141-WO-PCT a borehole node that stores borehole data including a borehole identifier data, borehole name data, borehole location data, and borehole province data; a sample node that stores source rock sample data including hydrocarbon shows data, pyrolysis data, and maceral group geochemistry data; an environment node that stores environment interpretation data of intervals along a well trajectory including depositional environment explanation data, geological basin data, and interval depth data; a lithology node that stores lithology interpretation data of intervals along the well trajectory including lithology explanation data and interval depth data; and a formation node that stores geological information data of intervals along the well trajectory including formation name data, geological age data, and interval depth data.
3. The method of Claim 1, wherein the link type data comprises at least one of a first computing structure that stores link type data between a first borehole relative to a sample derived from the first borehole; a second computing structure indicating a link between the first borehole and a second borehole and which stores data including geographical distance data between the first borehole and the second borehole; a third computing structure that links borehole data and depositional environment data; a fourth computing structure that links borehole data and lithology data; and a fifth computing structure that links borehole data and formation data.
4. The method of Claim 1, wherein transforming the data schema into a graph database comprises formatting datasets derived from the data schema into a table such that data rows of the table comprise different samples or boreholes while data columns of the table comprise features of each sample.
5. The method of Claim 4, wherein each feature comprised in the features of each sample indicates geochemical data including one of: pyrolysis data including total organic carbon (TOC) data;S 1 parameter data;Attorney Docket No.: IS24.0141-WO-PCT52 parameter data;53 parameter data; and Tmax parameter data.
6. The method of Claim 4, wherein each feature comprised in the features of each sample indicates geochemical data including one of: inertinite data; liptinite data; and vitrinite data.
7. The method of Claim 1, wherein nodes and links comprised in the graph database facilitate querying or fdtering the graph database using graph query sentences aligned with semantic or syntactic query structures within a query language associated with the graph database.
8. The method of Claim 1 , further comprising deriving the analyzed graph data from a quality control process including conditioning the graph database such that the graph database is used to generate a graph data representation, wherein conditioning the graph database comprises determining data interactions of the graph representation to gain quick insights or establish data relationships between neighbor wells at the first resource site.
9. The method of Claim 8, wherein the quality control process is used to validate or confirm an accuracy of the graph representation.
10. The method of Claim 1, wherein the trainable data is generated based on filtering out core sample data indicating low thermal maturity levels within the analyzed graph data.
11. The method of Claim 1, wherein holistically enhancing the resolved data comprises converting local feature data associated with the resolved data to global feature data.Attorney Docket No.: IS24.0141-WO-PCT12. The method of Claim 1, wherein holistically enhancing the resolved data comprises converting a plurality of data elements comprised in the resolved data or imputed data into data modes of the training data that drive properties of the subterranean model13. The method of Claim 12, wherein converting the plurality of data elements comprised in the resolved data comprise: converting geological basin identifier data to basin type data including at least one of a rift basin type, a passive basin type, a foreland basin type, or a back-arc basin type; converting stratigraphy identifier data to geological age data; and converting hydrocarbon shows data from string data to sequential integer data.
14. The method of Claim 1, wherein applying the training data to train the subterranean model comprises one or more of grid search training or optimization computing operations on the subterranean model using the training data; dimension reduction training or optimization computing operations on the subterranean model using the training data; category data encoding training or optimization computing operations on the subterranean model using the training data; missing data imputation training or optimization computing operations on the subterranean model using the training data; and hyperparameter fine-tune training or optimization operations on the subterranean model using the training data.
15. The method of Claim 1, wherein testing the trained subterranean model comprises: receiving geochemical data associated with the first resource site or the second resource site; applying the geochemical data to one or more parameters of the trained subterranean model; and generating the prediction report in response to applying the geochemical data.Attorney Docket No.: IS24.0141-WO-PCT16. The method of Claim 15, wherein the geochemical data is derived from lab analysis of core data extracted at the first resource site or the second resource site.
17. The method of Claim 15, wherein the one or more parameters of the trained subterranean model comprises at least one of a total organic carbon (TOC) parameter associated with the geochemical data; a bulk pyrolysis parameter associated with the geochemical data, the bulk pyrolysis parameter comprising one of an SI parameter, an S2 parameter, an S3 parameter, and Tmax parameter; and a maceral composition parameter.
18. The method of Claim 1, wherein the prediction report is analyzed by the machine learning engine to: confirm probability data associated with the prediction report or associated with result data of the prediction report and thereby ensure that the prediction report is within probabilistic thresholds associated with the first resource site or the second resource site; and analyze the result data by mapping or correlating the result data to data structures comprised in the graph representation.
19. The method of Claim 1, further comprising configuring equipment based on the prediction report.
20. A system for characterizing and classifying source rocks at a resource site, the system comprising: a computing platform comprising a database system, a data processing system, and a machine learning engine; a computer processor; and memory storing instructions which are executable by the computer processor to: generate, using the database system, analyzed graph data by:Attorney Docket No.: IS24.0141-WO-PCT transforming a data schema indicating node type data and link type data into a graph database, the node type data and link type data being parameterized using geochemical data associated with a first resource site; creating, using the graph database, a graph data representation indicating characterized or classified subsurface structures associated with the first resource site, and conditioning the graph database to generate the analyzed graph data; filter, using the data processing system, the analyzed graph data based on vitrinite reflectance data and thereby generate trainable data; resolve, using the data processing system, data discrepancies within the trainable data and thereby generate resolved data; holistically enhance, using the data processing system, the resolved data to be compatible with a plurality of subterranean structures and thereby generate training data; apply, using the machine learning engine, the training data to train a subterranean model and thereby generate a trained subterranean model, the subterranean model characterizing or classifying rock characteristics associated with the first resource site; and test, using the machine learning engine, the trained subterranean model and thereby generate a prediction report indicating rock characteristics and classification of a source rock for the first resource site or a second resource site that is similar to, or distinct from the first resource site.
Citation Information
Patent Citations
Detection of missing entities in a graph schema
US10789296B2
Probability distribution assessment for classifying subterranean formations using machine learning
US20220004919A1
Method and system for spectroscopic prediction of subsurface properties using machine learning
US20220351037A1
Method for categorizing a rock on the basis of at least one image
US20230008058A1
Method of Detecting at Least One Geological Constituent of a Rock Sample
US20230154208A1