A method for identifying the genesis of unconventional oil and gas in a continental basin based on AI technology
By constructing a multi-dimensional identification feature system and geological constraint rules, the problem of strong multiple interpretations in unconventional oil and gas exploration in continental basins has been solved. High-precision and geologically interpretable oil and gas genetic identification has been achieved, which can be adapted to multi-source heterogeneous data and output continuous genetic distribution profiles.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XI'AN PETROLEUM UNIVERSITY
- Filing Date
- 2026-04-29
- Publication Date
- 2026-07-21
AI Technical Summary
Existing hydrocarbon genesis identification technologies suffer from problems such as strong multiple interpretations and insufficient geological reliability. In particular, in unconventional hydrocarbon exploration in continental basins, traditional methods are difficult to adapt to multi-source heterogeneous data, resulting in inaccurate identification results.
A multi-dimensional identification feature system is constructed, including hydrocarbon source differences, source-reservoir adjacency and migration and retention characteristics. Combined with geological constraint rules, high-precision hydrocarbon genesis identification is achieved through hierarchical progressive identification and Bayesian fusion optimization.
It enables high-precision and geologically interpretable identification of hydrocarbon genesis in unconventional oil and gas exploration in continental basins, adapts to multi-source heterogeneous data, and outputs continuous genetic distribution profiles to meet exploration and development needs.
Smart Images

Figure CN122432697A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of hydrocarbon genesis identification technology, and specifically discloses a method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology. Background Technology
[0002] Unconventional oil and gas in continental basins is the core replacement area for oil and gas reserves and production. Controlled by multi-cycle tectonic evolution and sedimentary environment, continental basins generally develop multiple sets of superimposed source rocks, with rapid sedimentary facies changes and strong reservoir heterogeneity. The formation process of unconventional oil and gas is complex, with various genetic types such as single-source hydrocarbon supply, mixed-source hydrocarbon supply, and near-source stagnant accumulation. Accurate identification of the genesis of oil and gas is the core prerequisite for clarifying the accumulation law, optimizing exploration deployment, and improving development efficiency.
[0003] The technologies related to hydrocarbon genesis identification can be mainly divided into the following categories: The first category is conventional oil-source correlation techniques based on geochemical indicators. These techniques mainly use single or a few geochemical characteristic parameters such as biomarkers, carbon isotopes, natural gas components and isotopes to compare the affinity between oil and gas samples and candidate source rocks, thereby achieving a preliminary identification of the genesis of oil and gas.
[0004] The second category is the contribution decomposition technology of mixed-source oil and gas based on mathematical statistical methods. This type of technology mainly uses mathematical methods such as principal component analysis, endmember mixing model, and multiple linear regression to fit and calculate the contribution ratio of source rocks of mixed-source oil and gas.
[0005] The third category is oil and gas reservoir identification technology based on artificial intelligence. In recent years, AI technologies such as machine learning and deep learning have been gradually applied to the field of oil and gas exploration. Some existing technologies attempt to achieve functions such as oil and gas layer identification, source rock evaluation, and prediction of favorable reservoir areas through AI models.
[0006] However, existing hydrocarbon origin identification technologies still have the following shortcomings: Conventional source-source correlation techniques based on geochemical indicators rely on single or a few characteristic parameters, resulting in multiple interpretations and difficulty in effectively distinguishing complex origin types with similar geochemical responses. Mixed-source decomposition techniques based on mathematical statistics generally emphasize mathematical fitting while neglecting geological constraints, easily leading to unreasonable results that are out of sync with actual hydrocarbon accumulation patterns. Existing AI-related identification technologies mostly use general black-box models, failing to construct a dedicated feature system and interpretable identification logic based on continental hydrocarbon accumulation mechanisms. This makes them difficult to adapt to the processing needs of multi-source heterogeneous data in exploration sites, resulting in insufficient geological reliability and exploration applicability of the identification results.
[0007] This invention provides a method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology, in order to solve the above-mentioned problems. Summary of the Invention
[0008] The purpose of this invention is to provide a method for identifying the genesis of unconventional hydrocarbons in continental basins that can integrate multi-source exploration data, fuse hydrocarbon accumulation geological mechanisms, and has high identification accuracy and strong geological interpretability.
[0009] To achieve the above objectives, the basic solution of this invention provides a method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology, comprising the following steps: Step A1: Obtain multi-source basic data within the target well area, divide the object to be identified into well sections to form target samples, and establish a reference feature set of candidate source rocks; Step A2: Perform deep correction, interval aggregation, standardization, and missing value completion on the multi-source basic data to construct a basic sample set; Step A3: Based on the basic sample set, construct a genetic identification feature system from three dimensions: source difference, source-reservoir adjacency, and migration and retention. Among them, the source difference feature represents the geochemical affinity between the target sample and each candidate source rock, the source-reservoir adjacency feature represents the spatial configuration relationship between the source rock and the reservoir, and the migration and retention feature represents the migration and retention capacity of oil and gas. Step A4: Based on the genetic identification feature system, perform hierarchical progressive genetic identification. First-level identification distinguishes between single-source hydrocarbon supply and multi-source hydrocarbon supply candidate samples. Second-level identification further distinguishes between mixed-source hydrocarbon supply type and near-source stagnation type hydrocarbon accumulation for multi-source hydrocarbon supply candidate samples. Contribution decomposition processing is performed on samples identified as mixed-source hydrocarbon supply type. Step A5: Combine the preset geological constraint rules to perform consistency correction on the identification results, optimize the probability of each cause type through Bayesian fusion, and output the final cause identification result.
[0010] Furthermore, in step A2, the standardization adopts a robust standardization method, and the standardized parameters... It can be represented as: ; In the formula, This represents the median of the j-th parameter in the entire sample. This represents the interquartile range of the j-th parameter across all samples. This indicates the smallest positive number whose denominator is zero. The j-th term parameter represents the interval value of the i-th target sample; Missing values are filled using weighted imputation of similar samples from neighboring well sections, adjacent well locations, or the same sedimentary unit. It can be expressed as follows: ; In the formula, This represents the planar distance between the i-th target sample and the r-th candidate sample. This indicates the difference in strata elevation between the two. This represents the difference in sedimentary facies, taking the value 0 when the sedimentary facies are the same and 1 when they are different. , , These represent the attenuation control parameters for planar distance, stratigraphic elevation difference, and sedimentary facies difference, respectively.
[0011] Furthermore, in step A3, the process of constructing the hydrocarbon source difference characteristics includes: The difference distance between the i-th target sample and the k-th candidate source rock is calculated using Mahalanobis distance. ; In the formula, This represents the geochemical feature vector of the i-th target sample. This represents the reference feature center of the k-th candidate source rock. The inverse matrix of the covariance matrix of the k-th candidate source rock is represented. Convert the difference distance into normalized response similarity. : ; In the formula, Indicates the temperature regulation coefficient. This indicates the total number of candidate source rocks. For summation index variables; Constructing multi-source response entropy based on similarity Separation between primary and secondary sources : ; In the formula, Represents the natural logarithm function. Represents a very small positive number; ; In the formula, and Let represent the first and second largest similarity values of the i-th target sample after sorting.
[0012] Furthermore, in step A3, the process of constructing the source-storage adjacency features includes: Calculate the normalized source-storage distance : ; In the formula, This represents the planar distance between the i-th target sample and the k-th candidate source rock. Indicates vertical distance. and These represent the scale factors for planar distance and vertical distance, respectively; The source-storage distance and source-storage interlayer frequency normalization indices are: The contact relationship index is Thickness matching index is The index of broken connectivity is Integrate into source-storage adjacency index : ; In the formula, Indicates the distance attenuation coefficient. .
[0013] Furthermore, in step A3, the process of constructing the transport and retention characteristics includes: Constructing the transport capability index and retention index : ; In the formula, This represents the standardized porosity of the i-th target sample. Indicates standardized penetration rate, Indicates a standardized crack development index. Indicates the standardized transport connectivity index. ; ; In the formula, This represents the standardized oil saturation of the i-th target sample. Indicates standardized capping capability indicators. Indicates standardized storage condition indicators. This indicates the standardization pressure to maintain the indicator. ; A comparative index of transport and retention was constructed based on the transport capacity index and the retention index. : .
[0014] Furthermore, in step A4, the first-level identification process includes: Constructing a multi-source hydrocarbon supply rating system : ; In the formula, Represents the multi-source response entropy. Indicates the separation degree between primary and secondary sources. This represents the maximum similarity corresponding to the i-th target sample. This represents the normalized transport capacity index. ; The multi-source hydrocarbon supply score is converted into an initial probability for first-level identification using the Sigmoid function: ; ; In the formula, This represents the initial probability that the i-th target sample belongs to the multi-source hydrocarbon supply candidate sample. This represents the initial probability that the i-th target sample belongs to the single-source hydrocarbon supply candidate sample. This represents the first-level kurtosis parameter. This represents the first-level identification threshold.
[0015] Furthermore, in step A4, the secondary identification process includes: Constructing a mixed-source hydrocarbon supply scoring system and near-source retention score : ; ; In the formula, , , , Weighting of mixed-source hydrocarbon supply for scoring. , , , As the weight for the near-source retention score, and These are the source-storage adjacency indices, representing the first and second largest similarities after sorting the similarities of the i-th target sample, respectively. Based on the difference between the two scores, the conditional probability of the second-level identification is obtained through the Sigmoid function: ; ; In the formula, Let represent the conditional probability that the i-th target sample belongs to the mixed-source hydrocarbon supply type, given that it has already been identified as a multi-source hydrocarbon supply candidate. Let represent the conditional probability that the i-th target sample, given that it has been identified as a multi-source hydrocarbon supply candidate, belongs to the near-source retention type hydrocarbon accumulation. This represents the second-order kurtosis parameter. This indicates the threshold for secondary identification.
[0016] Furthermore, in step A4, the contribution decomposition process includes: Construct an optimization objective function with dual geological constraints: ; Simultaneously satisfy ; In the formula, Let represent the objective function for the contribution decomposition of the i-th target sample. This represents the geochemical feature vector of the i-th target sample. This represents a matrix composed of reference feature centers from various candidate source rocks. This represents a vector composed of the weights of each prior contribution. Represents the balance coefficient. This represents the final contribution ratio of the k-th candidate source rock to the i-th target sample.
[0017] Furthermore, in step A5, the preset geological constraint rules include at least maturity window matching rules, source-reservoir space configuration rules, and fault conduction condition rules, and a comprehensive geological consistency coefficient is constructed by weighted summation. The corrected final probability is obtained by performing Bayesian fusion using the following formula: ; In the formula, This represents the final probability of the i-th target sample after geological correction under candidate causal type c. This represents the initial probability of the i-th target sample under candidate cause type c. A variable for traversing all candidate cause types.
[0018] Furthermore, step A5 also includes smoothly combining the identification results of adjacent target samples along the well depth direction to form a continuous longitudinal causal profile. The well section-level probability profile is represented as follows: ; In the formula, This represents the well segment-level probability that well w belongs to candidate causal type c at depth z. This represents the distance weight of the i-th target sample with respect to depth z. The local sample window corresponding to well w at depth z.
[0019] The principle and effect of this solution are as follows: 1. Compared with existing technologies, this invention constructs a standardized processing system for multi-source heterogeneous data throughout the entire process. Through steps such as depth correction, interval aggregation, robust standardization, and completion of missing geological constraint values, it achieves standardized integration of multi-source and multi-dimensional exploration data. At the same time, it completes the construction of refined sample units at the well section scale through well section division, avoiding the errors of traditional general analysis of the entire well section, and laying a high-quality data foundation for causal identification.
[0020] 2. Compared with existing technologies, this invention is based on the core mechanism of unconventional hydrocarbon accumulation in terrestrial facies and constructs a unique identification feature system with three dimensions: hydrocarbon source difference, source-reservoir adjacency, and migration and retention. It comprehensively covers the control elements of the entire chain of hydrocarbon generation, migration, and accumulation. The accompanying hierarchical and progressive AI identification logic can accurately distinguish complex genetic types such as single-source hydrocarbon supply, mixed-source hydrocarbon supply, and near-source retention accumulation. It solves the problems of strong ambiguity in traditional technology identification and lack of geological interpretability in AI models due to their black box nature.
[0021] 3. Compared with existing technologies, this invention constructs a mixed-source contribution decomposition model with dual-objective geological constraints. The dual optimization objectives are minimizing the fitting error of geochemical characteristics and minimizing the deviation from the prior weights of hydrocarbon accumulation geology. This model can output a quantitative hydrocarbon supply ratio that combines high accuracy and geological rationality, avoiding the problem of traditional pure mathematical fitting results being disconnected from geological laws. Simultaneously, through consistency correction of multi-dimensional geological rules such as maturity matching, source-reservoir configuration, and fracture transport, the initial identification results are optimized using Bayesian fusion, further improving the geological reliability of the identification results. It can output a continuous genetic distribution profile along well depth, fully adapting to the refined needs of actual exploration and development. This achieves a deep integration of AI technology and hydrocarbon accumulation geological mechanisms, possessing high identification accuracy, strong geological interpretability, and wide-area generalization ability. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 The flowchart illustrates a method for identifying the genesis of unconventional hydrocarbons in terrestrial basins based on AI technology, as proposed in an embodiment of this application. Detailed Implementation
[0024] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0025] A method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology, implementing, for example... Figure 1 As shown, it includes the following steps: Step A1: Obtain multi-source basic data within the target well area, divide the object to be identified into well sections to form target samples, and establish a reference feature set of candidate source rocks.
[0026] Acquiring multi-source basic data within the target well area specifically refers to multi-source basic data related to the identification of unconventional hydrocarbon genesis. This multi-source basic data includes at least geochemical data, well logging data, source rock parameters, reservoir parameters, and sedimentary facies and structural information corresponding to the target stratigraphic interval.
[0027] Specifically, geochemical data includes total organic carbon content, pyrolysis parameters, vitrinite reflectance, biomarker parameters, carbon isotope parameters, and natural gas composition and isotope parameters. Well logging data includes natural gamma rays, resistivity, sonic transit time, density, and neutron porosity. Source rock parameters include source rock thickness, organic matter type, maturity, hydrocarbon generation potential, stratigraphic position, and planar distribution characteristics. Reservoir parameters include porosity, permeability, oil saturation, fracture development, mineral composition, and brittleness characteristics. Sedimentary facies and structural data include sedimentary microfacies, stratigraphic distribution, fault development, burial depth, structural location, caprock conditions, and preservation conditions.
[0028] Furthermore, in step A1, in order to achieve accurate identification of well section scale and avoid errors caused by general analysis of the entire well section, the objects to be identified within the target well area are first divided into well sections, and the continuous well profile is split into multiple independent target sample units. The well section corresponding to each target sample for: ; In the formula, This represents the top boundary depth of the i-th target sample well segment. This represents the bottom boundary depth of the i-th target sample well section.
[0029] To improve the comparability and stability of subsequent genetic identification, a reference feature set of candidate source rocks is preferably established in step A1. Let there be K sets of candidate source rocks in the study area, and the number of reference samples corresponding to the k-th candidate source rock is... The geochemical eigenvector of its r-th reference sample is denoted as To obtain a representative geochemical response benchmark for each source rock set, the reference characteristic center of the k-th candidate source rock set was calculated by arithmetic mean. , is represented as: .
[0030] Furthermore, the covariance matrix of the k-th candidate source rock is calculated using the following formula. This is to characterize the joint fluctuation relationship among various geochemical parameters within the same set of candidate source rocks, and to eliminate the interference of parameter correlation on subsequent distance calculations. .
[0031] Step A2: Perform deep correction, interval aggregation, standardization, and missing value completion on the multi-source basic data to construct a basic sample set.
[0032] Following the multi-source basic data obtained in step A1, since the data from different sources have significant differences in depth benchmarks, layer correspondence, scale, and data completeness, they cannot be directly used for subsequent feature calculation and cause identification. Therefore, the multi-source data undergoes full-process standardization processing to eliminate systematic errors between data and construct a basic sample set with unified format, consistent dimensions, and complete data.
[0033] Specifically, depth correction and layer unification are first performed on data from different sources. Let the original depth of the j-th original parameter in the i-th target sample be... Its corrected depth The following formula represents the unification of core samples, cuttings, logging, well logging, experimental analysis, and geological interpretation results under the same depth and stratigraphic benchmark: ; In the formula, Indicates the well depth correction amount, This represents the layer-level contrast correction amount, where i is the target sample index and j is the parameter index.
[0034] For continuous logging curves, in the target sample well section Perform interval aggregation within the range to obtain the interval representative value of the j-th parameter in the i-th target sample. This enables the matching of continuous curves with discrete well section samples. ; In the formula, This indicates that the j-th continuous parameter is at the depth point The value at that location, Indicates the target sample well section The number of sampling points within, This is the index for the sampling points.
[0035] Furthermore, due to the significant differences in the dimensions and numerical ranges of different geological parameters, and to reduce the impact of common outliers on parameter scales during exploration, this embodiment preferably employs a robust standardization method to process the sample parameters. The standardized parameters... It can be represented as: ; In the formula, This represents the median of the j-th parameter in the entire sample. This represents the interquartile range of the j-th parameter across all samples. This indicates the avoidance of extremely small positive numbers with a denominator of zero.
[0036] To address the common problem of missing parameters in real-world exploration scenarios, for parameters with missing values, weighted completion is performed using similar samples from neighboring well sections, adjacent well locations, or the same sedimentary unit. The completed value of the j-th missing parameter in the i-th target sample is... It can be represented as: ; In the formula, This represents the set of candidate samples that are adjacent to or similar to the i-th target sample. This represents the value of the r-th candidate sample on the j-th parameter. It represents the completion weight of the r-th candidate sample on the i-th target sample.
[0037] Complete the weights It can be expressed as follows: ; In the formula, This represents the planar distance between the i-th target sample and the r-th candidate sample. This indicates the difference in strata elevation between the two. This represents the difference in sedimentary facies, taking the value 0 when the sedimentary facies are the same and 1 when they are different. , , These represent the attenuation control parameters for planar distance, stratigraphic elevation difference, and sedimentary facies difference, respectively.
[0038] After depth correction, interval aggregation, standardization, and missing value imputation, the base sample vector of the i-th target sample is constructed. : ; In the formula, Represents a vector of geochemical parameters. Represents a well logging parameter vector. Represents the source rock parameter vector. Represents a reservoir parameter vector. This represents a vector of sedimentary facies and tectonic parameters.
[0039] Step A3: Based on the basic sample set, construct a genetic identification feature system from three dimensions: source difference, source-reservoir adjacency, and migration and retention. Among them, the source difference feature represents the geochemical affinity between the target sample and each candidate source rock, the source-reservoir adjacency feature represents the spatial configuration relationship between the source rock and the reservoir, and the migration and retention feature represents the migration and retention capacity of oil and gas.
[0040] Building upon the basic sample set constructed in step A2, starting from the core control elements of unconventional hydrocarbon accumulation in continental basins, we construct an identification feature system that can accurately distinguish different hydrocarbon genesis types. This system is constructed around three core dimensions: hydrocarbon source difference, source-reservoir adjacency, and migration and retention, comprehensively covering the key features of the entire hydrocarbon chain from generation and migration to accumulation, avoiding the limitations of traditional single geochemical indicators.
[0041] First, a feature vector characterizing the source rock differences is constructed to represent the geochemical affinity between the target sample and different candidate source rocks. Let the geochemical feature vector of the i-th target sample be... The reference feature center of the k-th candidate source rock is The inverse matrix of the covariance matrix of the k-th candidate source rock is The Mahalanobis distance is used to calculate the comprehensive difference between the target sample and each set of source rocks, and the difference distance between the i-th target sample and the k-th candidate source rock set is calculated. The following formula is used to comprehensively characterize the combined differences among multiple geochemical parameters: .
[0042] After obtaining the difference distance, in order to more intuitively reflect the affinity between the target sample and each source rock, it is converted into the normalized response similarity of the target sample to different candidate source rocks. : ; In the formula, Indicates the temperature regulation coefficient. This indicates the total number of candidate source rocks. This is used as the summation index variable. A higher similarity indicates that the i-th target sample is more likely to be controlled by the hydrocarbon supply from the k-th candidate source rock.
[0043] Furthermore, to reflect the dispersion of the target sample's response to multiple candidate source rocks, a multi-source response entropy was constructed. : ; In the formula, Represents the natural logarithm function. It represents a very small positive number. The larger the value, the more dispersed the response of the i-th target sample to multiple candidate source rocks, and the more obvious the multi-source hydrocarbon supply characteristics.
[0044] Meanwhile, in order to quantify whether the target sample contains a single dominant source rock, a primary-secondary source separation degree is constructed based on the similarity ranking results. Let the first and second largest similarity values of the i-th target sample after sorting be respectively... and Then the separation degree between primary and secondary sources It can be represented as: .
[0045] The greater the separation between primary and secondary sources, the more obvious the single dominant candidate source rock is; the smaller the separation between primary and secondary sources, the more similar the responses of multiple candidate source rocks are, and the higher the possibility of multiple sources supplying hydrocarbons.
[0046] After completing the construction of source rock difference characteristics, source-reservoir adjacency characteristics are further constructed to characterize the spatial configuration relationship between source rocks and reservoirs.
[0047] First, the normalized source-reservoir distance is calculated by combining planar and vertical distances. Let the planar distance between the i-th target sample and the k-th candidate source rock be . The vertical distance is Then the normalized source-reservoir distance between the i-th target sample and the k-th candidate source rock is... It can be represented as: ; In the formula, and These represent the scale factors for planar distance and vertical distance, respectively.
[0048] Furthermore, in order to comprehensively quantify the favorable spatial configuration between source and reservoir, five core spatial elements—source-reservoir distance, mutual layer frequency, contact relationship, thickness matching degree, and fracture connectivity—are integrated to construct a comprehensive source-reservoir adjacency index.
[0049] Let the normalized index of the source-reservoir frequency of the i-th target sample relative to the k-th candidate source rock be . The contact relationship index is Thickness matching index is The index of broken connectivity is Then the source-reservoir adjacency index between the i-th target sample and the k-th candidate source rock is... It can be represented as: ; In the formula, Indicates the distance attenuation coefficient. .
[0050] Among them, thickness matching index The expression used to characterize the degree of matching between the thickness of the source rock and the thickness of the reservoir is as follows: ; In the formula, Indicates the thickness of the k-th candidate source rock. This indicates the reservoir thickness corresponding to the i-th target sample. The closer this index is to 1, the higher the matching degree of source and reservoir thickness.
[0051] Disconnectivity index The specific expression used to characterize the effect of fractures on the transport and connectivity between source and reservoir is as follows: ; In the formula, This represents the effective connectivity fracture length connecting the k-th candidate source rock and the i-th target sample reservoir. Indicates the total length of the relevant fracture. This represents the shortest distance between the target sample and the control fracture. This represents the fracture distance attenuation coefficient. The closer this index is to 1, the stronger the fracture's conduction and connectivity effect on the source and reservoir.
[0052] In the above manner, the multi-dimensional source-storage space configuration characteristics are uniformly represented as the source-storage adjacency index. The larger the index, the more likely the reservoir containing the k-th candidate source rock and the i-th target sample has near-source hydrocarbon supply and short-distance migration conditions, and the higher the geological probability of hydrocarbon supply.
[0053] After constructing the hydrocarbon source differences and source-reservoir adjacency characteristics, migration and retention characteristics are further constructed to characterize the migration and retention capabilities of oil and gas after generation. First, a migration capacity index is constructed to quantify the ease of oil and gas migration within the reservoir. Let the standardized porosity of the i-th target sample be . Standardization penetration rate Standardized crack development index is Standardized transport connectivity index is The transport capacity index It can be represented as: ; In the formula, .
[0054] in, This represents the value obtained by taking the natural logarithm of the permeability parameter and then standardizing it. This reduces the impact of large differences in permeability across multiple orders of magnitude. The larger the value of this index, the stronger the migration ability of oil and gas in the reservoir, and the easier it is for long-distance migration to occur.
[0055] Simultaneously, a retention and preservation index is constructed to quantify the retention and preservation capacity of oil and gas in the reservoir. Let the standardized oil saturation of the i-th target sample be . Standardized capping capability indicators are Standardized storage condition indicators are: Standardized pressure maintenance index is The retention index It can be represented as: ; In the formula, The higher the retention index, the stronger the retention capacity of oil and gas in the reservoir, and the easier it is to form near-source retention-type reservoirs.
[0056] To more intuitively reflect the tendency of the target sample's accumulation type, this embodiment constructs a migration and retention comparison index based on the migration capacity index and the retention and preservation index. The specific expression is: .
[0057] The larger the retention contrast index, the more likely the i-th target sample is to form a near-source retention type of hydrocarbon accumulation. The smaller the retention contrast index, the more likely it is to undergo a stronger migration process and is more likely to be an exogenous migration hydrocarbon accumulation.
[0058] Through the feature construction process described above, the core features from the three dimensions are integrated to construct the causal identification feature vector for the i-th target sample. The specific expression is: ; In the formula, This represents the sedimentary facies and structural auxiliary feature vector, which comprehensively covers the core control elements of hydrocarbon genesis identification and provides complete feature input for subsequent hierarchical genetic identification.
[0059] Step A4: Based on the genetic identification feature system, perform hierarchical progressive genetic identification. First-level identification distinguishes between single-source and multi-source hydrocarbon candidate samples. Second-level identification further distinguishes between mixed-source hydrocarbon candidate samples and near-source stagnant hydrocarbon accumulation. Contribution decomposition is performed on samples identified as mixed-source hydrocarbon type.
[0060] Following the genetic identification feature system constructed in step A3, this step adopts a hierarchical and progressive identification logic to address the complex characteristics of hydrocarbon genesis in continental basins. First, a first-level identification distinguishes between single-source and multi-source hydrocarbon-supplying candidate samples. Then, a second-level identification further distinguishes between mixed-source hydrocarbon-supplying and near-source stagnant hydrocarbon accumulation in multi-source candidate samples. Ultimately, this achieves accurate classification of the three genetic types. At the same time, for samples identified as mixed-source hydrocarbon-supplying, the hydrocarbon contribution of each set of source rocks is decomposed, providing quantitative evidence for hydrocarbon accumulation research.
[0061] First, a primary identification process is performed to quickly screen samples with the potential for multi-source hydrocarbon supply, achieving an initial distinction between single-source and multi-source hydrocarbon supply. Then, a multi-source hydrocarbon supply scoring system is constructed by integrating four core indicators: multi-source response entropy, primary and secondary source separation, maximum similarity, and transport capacity. This comprehensively reflects the multi-source hydrocarbon supply potential of the sample, and the specific expression is as follows: ; In the formula, Represents the multi-source response entropy. Indicates the separation degree between primary and secondary sources. This represents the maximum similarity corresponding to the i-th target sample. This represents the normalized transport capacity index. The higher the score, the more significant the multi-source hydrocarbon supply characteristics of the sample. .
[0062] To convert the scores into statistically significant identification probabilities, the multi-source hydrocarbon supply scores are transformed into initial probabilities for first-level identification using the Sigmoid function. The specific expression is as follows: ; ; In the formula, This represents the initial probability that the i-th target sample belongs to the multi-source hydrocarbon supply candidate sample. This represents the initial probability that the i-th target sample belongs to the single-source hydrocarbon supply candidate sample. This represents the first-level kurtosis parameter, used to control the rate of change of the probability curve. This represents the first-level identification threshold.
[0063] To ensure the threshold has geostatistical significance, the optimal threshold is preferably determined by maximizing the Youden index using a known genetic sample set. The specific expression is as follows: ; In the formula, Indicates threshold Sensitivity at the lower level Indicates threshold The specificity of the threshold determined by this method can distinguish between single-source and multi-source hydrocarbon supply samples to the greatest extent. After the first-level identification is completed, all target samples can be divided into two categories: single-source hydrocarbon supply candidate samples and multi-source hydrocarbon supply candidate samples.
[0064] For samples identified as candidates for multi-source hydrocarbon supply in the first-level identification, the second-level identification is carried out to further distinguish whether they are mixed-source hydrocarbon supply type with multiple sets of source rocks jointly supplying hydrocarbons, or near-source retention type hydrocarbon accumulation formed by a single near-source rock, thus solving the problem that the two types of hydrocarbon accumulation have similar geochemical responses and are difficult to distinguish using traditional methods.
[0065] In this embodiment, mixed-source hydrocarbon supply scores and near-source retention scores are constructed respectively to quantify the matching degree of the two etiological types from different dimensions. Let the source-storage adjacency indices corresponding to the first two similarities be respectively. and Then the mixed-source hydrocarbon supply score and near-source retention score They are represented as follows: ; ; In the formula, , , , Weighting of mixed-source hydrocarbon supply for scoring. , , , The weights for the near-source retention score are as follows: the mixed-source hydrocarbon supply score focuses on the dispersion of multi-source responses, the spatial matching of dual-source hydrocarbon supply conditions and migration capacity, while the near-source retention score focuses on the adjacency conditions of dominant source rocks, retention and preservation capacity and the comparison characteristics of migration and retention.
[0066] Based on the difference between the two scores, the conditional probability of the second-level identification is obtained again through the Sigmoid function, the specific expression of which is: ; ; In the formula, Let represent the conditional probability that the i-th target sample belongs to the mixed-source hydrocarbon supply type, given that it has already been identified as a multi-source hydrocarbon supply candidate. Let represent the conditional probability that the i-th target sample, given that it has been identified as a multi-source hydrocarbon supply candidate, belongs to the near-source retention type hydrocarbon accumulation. This represents the second-order kurtosis parameter. This indicates the threshold for secondary identification.
[0067] Combining the probability results of primary and secondary identification, the initial probabilities of three genetic types—single-source hydrocarbon supply, mixed-source hydrocarbon supply, and near-source retention—are finally obtained. The specific expressions are as follows: ; ; ; In the formula, This represents the initial probability that the i-th target sample belongs to the mixed-source hydrocarbon supply type. This represents the initial probability that the i-th target sample belongs to the near-source retention type of hydrocarbon accumulation, and the sum of the above three initial probabilities is 1.
[0068] For target samples ultimately determined to be of the mixed-source hydrocarbon-supply type, further contribution decomposition processing is performed to accurately determine the hydrocarbon supply ratio of different candidate source rocks. First, based on the similarity and source-reservoir adjacency index obtained in step A3, a priori contribution weights are constructed, taking into account both geochemical affinity and the favorable nature of space hydrocarbon supply, ensuring that the priori weights conform to basic geological laws. The specific expression is as follows: ; In the formula, Let represent the prior contribution weight of the k-th candidate source rock in the i-th target sample. The sum of all weights is 1, which satisfies the normalization requirement.
[0069] To obtain more stable contribution ratio results that better conform to geological laws, an optimization objective function with dual geological constraints is constructed. On the one hand, it minimizes the fitting error between the weighted combination of the reference features of each source rock and the actual geochemical features of the target sample, ensuring the matching of geochemical features. On the other hand, it minimizes the deviation between the final contribution ratio and the prior contribution weight, ensuring that the results conform to geological laws and avoiding geological anomalies caused by pure mathematical fitting.
[0070] Let the contribution ratio vector corresponding to the i-th target sample be . ,in Let represent the final contribution ratio of the k-th candidate source rock set to the i-th target sample. Let the reference source rock matrix, composed of the reference feature centers of each source rock set, be . Then the objective function of contribution decomposition can be expressed as: ; Simultaneously satisfying the nonnegativity constraint and normalization constraint of the contribution ratio: ; In the formula, Let represent the objective function for the contribution decomposition of the i-th target sample. This represents the geochemical feature vector of the i-th target sample. This represents a matrix composed of reference feature centers from various candidate source rocks. This represents a vector composed of the weights of each prior contribution. This represents the balance coefficient, which is used to control the balance between the fitting error and the prior constraints.
[0071] In this embodiment, the projection gradient method is preferably used to solve the above-mentioned constrained optimization problem. First, the gradient of the objective function with respect to the contribution ratio vector is calculated, and the specific expression is as follows: ; In the formula, This represents the gradient of the objective function with respect to the contribution ratio vector.
[0072] Based on the gradient calculation results, perform iterative update to solve the problem, let the i-th... The contribution ratio vector at the next iteration is First, gradient descent is performed to update the vector, resulting in the unprojected intermediate update vector. ; In the formula, This represents the unprojected intermediate update vector. This indicates the iteration step size.
[0073] Since the intermediate vector updated by gradient descent may not satisfy the nonnegativity and normalization constraints, it needs to be projected onto the probabilistic simplex constraint set. The above yields the update vector that satisfies the constraints: ; In the formula, This represents the projection onto the probabilistic simplex constraint set. The projection operator.
[0074] The iterative process continues until the convergence stopping condition is met: ; In the formula, This represents the convergence threshold. Once the convergence condition is met, the iteration stops, and the final hydrocarbon supply ratio of each set of candidate source rocks to the i-th target sample can be obtained, thus completing the contribution decomposition of the mixed source hydrocarbon supply sample.
[0075] Step A5: Combine the preset geological constraint rules to perform consistency correction on the identification results, optimize the probability of each cause type through Bayesian fusion, and output the final cause identification result.
[0076] Following the initial identification results obtained in step A4, and combining them with the basic geological laws of hydrocarbon accumulation in continental basins, the initial results are corrected for consistency. This eliminates the deviations caused by pure mathematical calculations that are inconsistent with geological understanding, improves the geological reliability of the identification results, and finally outputs standardized and applicable genetic identification results.
[0077] The preset geological constraint rules include at least maturity window matching rules, source-reservoir space configuration rules, sedimentary facies constraint rules, fault conduction condition rules, and preservation condition constraint rules. For the i-th target sample and a certain candidate genetic type c, the consistency score of each geological rule is first calculated to quantify the degree of matching between the initial identification results and geological laws.
[0078] First, the maturity matching score is calculated. The maturity of the oil and gas must match the maturity window of the source rock supplying hydrocarbons. This is a fundamental premise for source rock correlation. The maturity matching score... It can be represented as: ; In the formula, Let represent the equivalent maturity of the i-th target sample. This indicates the reference maturity level corresponding to candidate causal type c. This represents the maturity tolerance parameter. The closer the score is to 1, the higher the degree of maturity matching.
[0079] Secondly, the source-reservoir space configuration score is calculated. Source rocks with a higher hydrocarbon supply ratio must possess more favorable source-reservoir adjacency conditions. The source-reservoir space configuration score... It can be represented as: ; In the formula, This represents the contribution ratio of the k-th candidate source rock set to the i-th target sample. This represents the corresponding source-reservoir adjacency index. The closer the score is to 1, the more favorable the spatial configuration of the source rocks for hydrocarbon supply.
[0080] Then, the fracture-conduction consistency score was calculated. Different types of hydrocarbon accumulation correspond to different fracture-conduction conditions. Near-source stagnant hydrocarbon accumulation has a weak dependence on fracture-conduction, while exogenous migration hydrocarbon accumulation has a strong dependence on fracture-conduction. The fracture-conduction consistency score is... It can be represented as: ; In the formula, This represents the fracture conduction index of the i-th target sample. This represents the range of fracture conduction indices allowed for candidate genetic type c. The distance function from a point to an interval is denoted by . This represents the fracture conduction tolerance parameter. The closer the score is to 1, the higher the degree of matching between the fracture conduction conditions and the genetic type.
[0081] The sedimentary facies consistency score and the preservation condition consistency score can be calculated based on the sedimentary matching matrix and the normalized value of the preservation condition parameter, respectively. The normalized output falls within the range of [0,1], and the closer the score is to 1, the higher the degree of matching.
[0082] After obtaining the consistency scores for each geological rule, a comprehensive geological consistency coefficient is constructed by weighted summation to fully quantify the degree of comprehensive matching between the initial identification results and geological laws. The specific expression is as follows: ; In the formula, Let represent the comprehensive geological consistency coefficient of the i-th target sample with respect to candidate genetic type c. .
[0083] Based on the comprehensive geological consistency coefficient, the initial probabilities of the three genetic types obtained in step A4 are fused with the geological consistency coefficient using Bayesian fusion to obtain the corrected final probabilities, the specific expression of which is: ; In the formula, This represents the final probability of the i-th target sample after geological correction under candidate causal type c. This represents the initial probability of the i-th target sample under candidate cause type c. A variable for traversing all candidate cause types.
[0084] After probability correction, the final causal type and identification confidence level of the target sample are determined. The causal type corresponding to the highest probability after correction is taken as the final causal type of the target sample, and this highest probability is taken as the identification confidence level. The specific expression is as follows: ; ; In the formula, This represents the final cause type of the i-th target sample. This represents the confidence level of identifying the i-th target sample. When... When the i-th target sample is marked as a low-confidence sample, the formula is as follows: This indicates the low-confidence identification threshold.
[0085] In well-section-level applications, to obtain a continuous longitudinal distribution of hydrocarbon accumulation and genesis along the wellbore, the identification results of adjacent target samples are smoothly combined along the well depth direction to form a continuous longitudinal genetic profile. Let the local sample window corresponding to well w at depth z be... Then the well section-level probability profile can be expressed as: ; In the formula, This represents the well segment-level probability that well w belongs to candidate causal type c at depth z. This represents the distance weight of the i-th target sample to depth z. The closer the distance, the greater the weight. This formula can be used to obtain the probability profile of the formation type that changes continuously along the well depth, which can intuitively show the formation and genetic distribution characteristics of the entire well section.
[0086] After the above geological correction and result optimization, the final genetic identification results of the target well area, target well, target layer or target well section are finally output. The output content includes at least: the genetic type corresponding to the target sample, the identification confidence level, the contribution ratio of each set of candidate source rocks, and the vertical genetic distribution results at the well section level. This completes the entire process of unconventional oil and gas genetic identification in the continental basin.
[0087] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any indirect modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology, characterized in that, Includes the following steps: Step A1: Obtain multi-source basic data within the target well area, divide the object to be identified into well sections to form target samples, and establish a reference feature set of candidate source rocks; Step A2: Perform deep correction, interval aggregation, standardization, and missing value completion on the multi-source basic data to construct a basic sample set; Step A3: Based on the basic sample set, construct a genetic identification feature system from three dimensions: source difference, source-reservoir adjacency, and migration and retention. Among them, the source difference feature represents the geochemical affinity between the target sample and each candidate source rock, the source-reservoir adjacency feature represents the spatial configuration relationship between the source rock and the reservoir, and the migration and retention feature represents the migration and retention capacity of oil and gas. Step A4: Based on the genetic identification feature system, perform hierarchical progressive genetic identification. First-level identification distinguishes between single-source hydrocarbon supply and multi-source hydrocarbon supply candidate samples. Second-level identification further distinguishes between mixed-source hydrocarbon supply type and near-source stagnation type hydrocarbon accumulation for multi-source hydrocarbon supply candidate samples. Contribution decomposition processing is performed on samples identified as mixed-source hydrocarbon supply type. Step A5: Combine the preset geological constraint rules to perform consistency correction on the identification results, optimize the probability of each cause type through Bayesian fusion, and output the final cause identification result.
2. The method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology according to claim 1, characterized in that, In step A2, the standardization adopts a robust standardization method, and the standardized parameters are... It can be represented as: ; In the formula, This represents the median of the j-th parameter in the entire sample. This represents the interquartile range of the j-th parameter across all samples. This indicates the smallest positive number whose denominator is zero. The j-th term parameter represents the interval value of the i-th target sample; Missing values are filled using weighted imputation of similar samples from neighboring well sections, adjacent well locations, or the same sedimentary unit. It can be expressed as follows: ; In the formula, This represents the planar distance between the i-th target sample and the r-th candidate sample. This indicates the difference in strata elevation between the two. This represents the difference in sedimentary facies, taking the value 0 when the sedimentary facies are the same and 1 when they are different. , , These represent the attenuation control parameters for planar distance, stratigraphic elevation difference, and sedimentary facies difference, respectively.
3. The method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology according to claim 1, characterized in that, In step A3, the process of constructing the hydrocarbon source difference characteristics includes: The difference distance between the i-th target sample and the k-th candidate source rock is calculated using Mahalanobis distance. ; In the formula, This represents the geochemical feature vector of the i-th target sample. This represents the reference feature center of the k-th candidate source rock. The inverse matrix of the covariance matrix of the k-th candidate source rock is represented. Convert the difference distance into normalized response similarity. : ; In the formula, Indicates the temperature regulation coefficient. This indicates the total number of candidate source rocks. For summation index variables; Constructing multi-source response entropy based on similarity Separation between primary and secondary sources : ; In the formula, Represents the natural logarithm function. Represents a very small positive number; ; In the formula, and Let represent the first and second largest similarity values of the i-th target sample after sorting.
4. The method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology according to claim 3, characterized in that, In step A3, the process of constructing the source-storage adjacency feature includes: Calculate the normalized source-storage distance : ; In the formula, This represents the planar distance between the i-th target sample and the k-th candidate source rock. Indicates vertical distance. and These represent the scale factors for planar distance and vertical distance, respectively; The source-storage distance and source-storage interlayer frequency normalization indices are: The contact relationship index is Thickness matching index is The index of broken connectivity is Integrate into source-storage adjacency index : ; In the formula, Indicates the distance attenuation coefficient. .
5. The method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology according to claim 1, characterized in that, In step A3, the process of constructing transport and retention features includes: Constructing the transport capability index and retention index : ; In the formula, This represents the standardized porosity of the i-th target sample. Indicates standardized penetration rate, Indicates a standardized crack development index. Indicates the standardized transport connectivity index. ; ; In the formula, This represents the standardized oil saturation of the i-th target sample. Indicates standardized capping capability indicators. Indicates standardized storage condition indicators. This indicates the standardization pressure to maintain the indicator. ; A comparative index of transport and retention was constructed based on the transport capacity index and the retention index. : 。 6. The method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology according to claim 1, characterized in that, In step A4, the first-level identification process includes: Constructing a multi-source hydrocarbon supply rating system : ; In the formula, Represents the multi-source response entropy. Indicates the separation degree between primary and secondary sources. This represents the maximum similarity corresponding to the i-th target sample. This represents the normalized transport capacity index. ; The multi-source hydrocarbon supply score is converted into an initial probability for first-level identification using the Sigmoid function: ; ; In the formula, This represents the initial probability that the i-th target sample belongs to the multi-source hydrocarbon supply candidate sample. This represents the initial probability that the i-th target sample belongs to the single-source hydrocarbon supply candidate sample. This represents the first-level kurtosis parameter. This represents the first-level identification threshold.
7. The method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology according to claim 6, characterized in that, In step A4, the secondary identification process includes: Constructing a mixed-source hydrocarbon supply scoring system and near-source retention score : ; ; In the formula, , , , Weighting of mixed-source hydrocarbon supply for scoring. , , , As the weight for the near-source retention score, and These are the source-storage adjacency indices, representing the first and second largest similarities after sorting the similarities of the i-th target sample, respectively. Based on the difference between the two scores, the conditional probability of the second-level identification is obtained through the Sigmoid function: ; ; In the formula, Let represent the conditional probability that the i-th target sample belongs to the mixed-source hydrocarbon supply type, given that it has already been identified as a multi-source hydrocarbon supply candidate. Let represent the conditional probability that the i-th target sample, given that it has been identified as a multi-source hydrocarbon supply candidate, belongs to the near-source retention type hydrocarbon accumulation. This represents the second-order kurtosis parameter. This indicates the threshold for secondary identification.
8. The method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology according to claim 1, characterized in that, In step A4, the contribution decomposition process includes: Construct an optimization objective function with dual geological constraints: ; Simultaneously satisfy ; In the formula, Let represent the objective function for the contribution decomposition of the i-th target sample. This represents the geochemical feature vector of the i-th target sample. This represents a matrix composed of reference feature centers from various candidate source rocks. This represents a vector composed of the weights of each prior contribution. Represents the balance coefficient. This represents the final contribution ratio of the k-th candidate source rock to the i-th target sample.
9. The method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology according to claim 1, characterized in that, In step A5, the preset geological constraint rules include at least maturity window matching rules, source-reservoir space configuration rules, and fault conduction condition rules. A comprehensive geological consistency coefficient is constructed by weighted summation. The corrected final probability is obtained by performing Bayesian fusion using the following formula: ; In the formula, This represents the final probability of the i-th target sample after geological correction under candidate causal type c. This represents the initial probability of the i-th target sample under candidate cause type c. A variable for traversing all candidate cause types.
10. The method for identifying the genesis of unconventional hydrocarbons in continental basins based on AI technology according to claim 9, characterized in that, Step A5 also includes smoothly combining the identification results of adjacent target samples along the well depth direction to form a longitudinal continuous causal profile. The well section-level probability profile is represented as follows: ; In the formula, This represents the well segment-level probability that well w belongs to candidate causal type c at depth z. This represents the distance weight of the i-th target sample with respect to depth z. The local sample window corresponding to well w at depth z.