Characteristic space guided modeling tight sandstone reservoir gas-bearing prediction method and device
By using a feature space-guided modeling method and combining kNN and CNN, the problem of small sample size in the prediction of gas content in tight sandstone reservoirs was solved, achieving high-precision prediction of reservoir gas content, alleviating the shortcomings of deep learning methods, and improving the prediction effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA PETROLEUM & CHEMICAL CORP
- Filing Date
- 2024-11-12
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies for predicting gas content in tight sandstone reservoirs are limited by the small sample size problem of deep learning methods, resulting in low prediction accuracy and difficulty in effectively characterizing complex nonlinear relationships.
A feature space-guided modeling method is adopted. The kNN method is used to initially predict gas content and select high-probability results as pseudo-training data. The CNN model is pre-trained using pseudo-training data and then transferred to actual training samples to establish a stable CNN model, thereby realizing the prediction of gas content in tight sandstone reservoirs.
It alleviates the small sample size problem of deep learning, improves the accuracy of gas content prediction in tight sandstone reservoirs, effectively characterizes the mapping relationship between reservoir seismic response and gas content, and improves the accuracy of prediction.
Smart Images

Figure CN122017982A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of exploration geophysics and artificial intelligence deep learning technology, specifically to a method and apparatus for predicting gas content in tight sandstone reservoirs using feature space-guided modeling. Background Technology
[0002] Among existing seismic reservoir technologies, traditional seismic reservoir prediction technologies include seismic attribute analysis technology, seismic inversion technology, and AVO (Amplitude Variation with Offset) technology.
[0003] Seismic attribute analysis is one of the most commonly used reservoir prediction techniques. Seismic attributes refer to characteristic parameters extracted from seismic data using mathematical methods, comprehensively reflecting changes in seismic data in geometry, statistics, dynamics, and kinematics. These changes are closely related to the spatial variations of reservoir properties and fluid properties. Predicting reservoir properties based on seismic attribute analysis requires calibrating seismic attributes using reservoir properties obtained from well logging interpretation. The methodology can be summarized as follows: first, extract various seismic attributes from well-side seismic data; then, perform correlation analysis with reservoir properties obtained from well logging interpretation to select sensitive attributes; next, calibrate the sensitive attributes using the reservoir properties obtained from well logging, and extrapolate the calibration results from the well point location to the entire work area. Therefore, the geological significance of seismic attributes comes from calibration. Without the calibration process with well logging reservoir properties, seismic attributes are merely geophysical parameters obtained through purely mathematical methods. However, the seismic response of subsurface reservoirs is affected by various factors such as reservoir thickness, subsurface structure, and fluid properties, leading to multiple interpretations in calibration.
[0004] Seismic inversion techniques, used to obtain subsurface elastic parameters or further fluid indicator factors, are widely applied reservoir prediction technologies in industry. Seismic inversion is the process of inferring subsurface characteristics based on seismic data. In a broad sense, seismic inversion technology can quantitatively calculate various geophysical parameters of formations based on seismic data. Because impedance retrieval has a clear physical meaning and has achieved good results in actual reservoir prediction, in a narrow sense, seismic inversion technology specifically refers to seismic impedance retrieval. Impedance retrieval results typically have higher resolution than the original seismic data. Seismic inversion techniques can be categorized based on the method of utilizing well-seismic data, such as narrowband inversion, well-logging-constrained inversion, and seismically constrained well-logging interpolation and extrapolation. Based on the type of seismic data, they can also be divided into pre-stack inversion and post-stack inversion. However, seismic inversion technology itself faces problems such as inaccurate wavelet acquisition and reliance on initial models. In practical applications, the band-limited characteristics of seismic data and field noise interference further increase the difficulty of seismic inversion.
[0005] AVO (Advanced Volatility Occurrence) technology refers to a class of techniques that analyze subsurface reservoirs by utilizing the variation of seismic reflection amplitude with offset. It can be divided into AVO forward modeling and AVO inversion analysis. The theoretical basis of AVO technology is the Zoeppritz equation. The seismic reflection coefficient is determined by the properties of the medium on both sides of the interface. By determining the AVO characteristics of reservoirs under different lithologies and physical properties, hydrocarbon detection or lithology differentiation can be qualitatively performed directly from the seismic record. Based on this understanding, AVO forward modeling technology was developed. However, complex subsurface conditions mean that there is no good one-to-one correspondence between AVO anomalies and reservoir lithology and physical properties, resulting in low accuracy in reservoir prediction using AVO forward modeling. AVO inversion technology establishes a direct relationship between certain formation parameters and reflection coefficients by approximating and simplifying the Zoeppritz equation. However, the assumptions made to achieve the approximation of the Zoeppritz equation limit the application scope of gas-bearing prediction, and the approximation itself lowers the upper limit of the method's accuracy.
[0006] In the prediction of gas content in tight sandstone reservoirs, the application of deep learning is still in its early stages. However, machine learning methods have long been established in geophysics, such as pattern recognition technology which emerged in the 1980s. Some machine learning methods (principal component analysis, support vector machines, clustering, etc.) are also widely used in geophysical processes. Theoretically, deep learning methods have a stronger ability to characterize nonlinear relationships than machine learning methods. However, deep learning requires a large amount of training data to train the prediction model. The small sample size problem may cause the deep learning modeling process to fail to converge completely, resulting in underfitting and reducing the effectiveness of deep learning methods. Researchers have recognized the key issues. Some studies have attempted to reduce the impact of the small sample size problem by designing convolutional neural network structures. However, the nonlinear characterization ability of deep learning depends on the massive number of training parameters in the model; modifying the structure of deep networks is essentially a "trade-off" between the number of training parameters and the amount of training data, and cannot fundamentally solve the small sample size problem. Other studies have used output lithofacies analysis methods and geostatistical interpolation techniques to create new samples. However, the newly generated samples lack physical meaning. Other studies utilize two different deep learning techniques to provide each other with pseudo-training samples for iterative training, and then select the model with higher accuracy to complete the acoustic impedance inversion task. Neither of these deep learning techniques can avoid the small sample size problem, and both may lead to error accumulation. Summary of the Invention
[0007] This invention provides a method and apparatus for predicting the gas content of tight sandstone reservoirs using feature space-guided modeling, in order to solve the aforementioned technical problem in the prior art where deep learning is limited by small sample sizes.
[0008] According to a first aspect of the present invention, a method for predicting the gas content of tight sandstone reservoirs using feature space-guided modeling is provided, comprising:
[0009] Establish a general paradigm for predicting the gas content of tight sandstone reservoirs using supervised learning methods based on pre-stack seismic gathers and well logging gas content curves;
[0010] A feature space-guided modeling method is established, wherein the feature space-guided modeling method includes:
[0011] The kNN method was used to make a preliminary gas content prediction for the target work area, and the high-probability kNN prediction results were selected as pseudo-training data.
[0012] The CNN model is pre-trained using the pseudo-training data, wherein the information of the feature space of the data samples carried in the pseudo-training data is used to guide the establishment of the initial model of the CNN model.
[0013] By using actual training samples to optimize a stable CNN model through transfer learning, gas-bearing prediction of tight sandstone reservoirs can be achieved through feature space-guided modeling.
[0014] Preferably, in establishing a general paradigm for predicting the gas content of tight sandstone reservoirs based on pre-stack seismic gathers and well logging gas content curves using supervised learning methods, the general paradigm is applicable to both machine learning and deep learning methods; wherein,
[0015] Using the gas-bearing curves obtained from well logging interpretation as labels and local waveform responses as features, seismic sampling points are classified into gas-bearing and non-gas-bearing sampling points.
[0016] Preferably, the process of the kNN method performing the classification task is a process of spatially partitioning the feature space based on labeled samples; wherein, the kNN method selects different k nearest neighbor training samples for each test sample to make a decision, and the number of nearest neighbors in the decision process is the value of k. When a small value of k is selected, the kNN method builds a more targeted prediction model for a single test sample, and when a large value of k is selected, more training samples are introduced for reference.
[0017] Preferably, in the initial model for establishing the CNN model, feature space information is transmitted to the CNN modeling process in the form of pseudo-training samples to supplement and constrain the information in the CNN modeling process.
[0018] Preferably, the stable CNN model is obtained through repeated pseudo-training data filtering and CNN model training processes.
[0019] Preferably, the CNN model includes an input layer, a hidden layer, and an output layer; the hidden layer includes a convolutional layer, an activation function, a pooling layer, and a fully connected layer, and the convolutional layer contains the neuron connection method unique to convolutional neural networks; the backpropagation of the convolutional neural network adjusts the weights and biases by minimizing the residuals, so that the convolutional neural network is continuously updated according to the training data; wherein, the backpropagation algorithm uses the idea of stochastic gradient descent.
[0020] Preferably, when the kNN method and the CNN model are applied individually to the task of predicting the gas content of tight sandstone reservoirs, they follow the same general paradigm for predicting the gas content of tight sandstone reservoirs based on supervised learning methods.
[0021] According to a second aspect of the present invention, a device for predicting the gas content of tight sandstone reservoirs using feature space-guided modeling is provided, comprising:
[0022] The first module establishes a general paradigm for supervised learning methods to predict the gas content of tight sandstone reservoirs based on pre-stack seismic gathers and well logging gas content curves; and
[0023] The second module is used to establish a feature space-guided modeling method.
[0024] The second establishment module includes:
[0025] The filtering submodule is used to perform preliminary gas content prediction of the target work area using the kNN method, and to filter high-probability kNN prediction results as pseudo-training data.
[0026] A pre-training submodule is used to pre-train the CNN model using the pseudo-training data, wherein the information carried in the pseudo-training data regarding the feature space of the data samples is used to guide the establishment of an initial model of the CNN model; and
[0027] The transfer learning submodule is used to optimize a stable CNN model using actual training samples through transfer learning, thereby enabling gas-bearing prediction of tight sandstone reservoirs through feature space-guided modeling.
[0028] According to a third aspect of the present invention, an electronic device is provided, comprising:
[0029] Memory; and
[0030] processor;
[0031] The memory is used to store one or more computer instructions; the one or more computer instructions are executed by the processor to implement the method described in any of the above.
[0032] According to a fourth aspect of the present invention, a readable storage medium is provided, wherein computer instructions are stored thereon; wherein, when executed by a processor, the computer instructions implement the method described in any of the preceding claims.
[0033] The technical solution of this invention, after alleviating the small sample problem faced by deep learning methods, effectively applies deep learning methods to the task of predicting the gas content of tight sandstone. It is beneficial to give full play to the advantages of deep learning in characterizing complex nonlinear relationships, clarify the mapping relationship between the seismic response of tight sandstone reservoirs and gas content, and ultimately improve the prediction accuracy of tight reservoirs.
[0034] Since the labels of the training samples do not affect the prediction process of the kNN classifier, the pseudo-training samples are generated based on the feature space distribution of the samples, a process that is entirely data-driven. In the feature space, the distribution range of the pseudo-training samples is greater than or equal to the distribution range of the actual training samples. However, in CNN modeling, fitting the training samples to the labels through a non-linear transformation is an irreversible process of information loss, and the transformation process is prone to getting trapped in local optima. Randomly generating pseudo-samples and inputting them into the CNN model for pre-training, and repeating this step, continuously supplements and constrains the CNN modeling process, which helps alleviate the small sample size problem. Attached Figure Description
[0035] Figure 1 This is a technical flowchart of a method for predicting the gas content of tight sandstone reservoirs using feature space-guided modeling in one embodiment.
[0036] Figure 2 This is a post-stack seismic cross-well profile based on actual data from one embodiment (it should be noted that, in order to more clearly demonstrate the in-phase axis characteristics of the target layer, Figure 3 The data shown is a post-stack seismic profile, but the prediction of gas content is still based on pre-stack seismic data.
[0037] Figure 3 This is a gas-bearing distribution profile predicted by the kNN method in one embodiment;
[0038] Figure 4 One embodiment uses only actual training samples and labels to train a CNN model and predict the gas content distribution profile obtained from the profile.
[0039] Figure 5 One embodiment uses only pseudo-training samples and labels to train a CNN model and predict the gas content distribution profile obtained from the profile.
[0040] Figure 6 One embodiment uses two training samples and labels to simultaneously train a CNN model and predict the gas content distribution profile obtained from the profile.
[0041] Figure 7 This is a gas-bearing distribution profile obtained by a reservoir gas-bearing prediction method that utilizes feature space-guided modeling in one embodiment.
[0042] Figure 8 This is a thin reservoir gas-bearing distribution slice located 19ms above L2 in one embodiment (the circular and cross marks in the figure indicate the well location, where the circular mark indicates that the predicted gas-bearing distribution is consistent with the gas-bearing shown on the well at that point, and the cross mark indicates that the gas-bearing display of the two is inconsistent).
[0043] Figure 9 This is a thin reservoir gas content distribution slice located 13 ms above L2 in one embodiment;
[0044] Figure 10 This is a flowchart of a method for predicting the gas content of tight sandstone reservoirs using feature space-guided modeling in one embodiment.
[0045] Figure 11 One embodiment Figure 10 The flowchart of step S2;
[0046] Figure 12 This is a schematic diagram of a device for predicting the gas content of tight sandstone reservoirs using feature space-guided modeling in one embodiment.
[0047] Figure 13 One embodiment Figure 12 The structural diagram of the second module is shown below. Detailed Implementation
[0048] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0049] To enable those skilled in the art to better understand the present invention, the technical solution of one embodiment will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some, not all, embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort should fall within the scope of protection of the present invention.
[0050] It should be noted that the terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be used interchangeably where appropriate for the embodiments of the invention described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0051] It should be understood that when an element (such as a layer, film, region, or substrate) is described as being "on" another element, the element may be directly on the other element, or there may be an intermediate element present. Moreover, in this invention, when an element is described as being "connected" to another element, the element may be "directly connected" to the other element, or "connected" to the other element via a third element.
[0052] Example 1
[0053] Tight reservoirs are the most abundant unconventional reservoirs and one of the main areas for increasing reserves and production in the oil exploration and development field. To develop and utilize tight reservoirs, it is essential to first accurately locate and characterize them. Deep learning methods show promise in the characterization of tight sandstone reservoirs.
[0054] Please refer to Figure 1 , Figure 10 , Figure 11 This embodiment provides a method for predicting the gas content of tight sandstone reservoirs using feature space-guided modeling. The aim is to effectively utilize deep learning to predict the gas content of tight sandstone reservoirs. Specifically, it uses the distribution of data in the feature space to guide deep learning modeling and establishes a mapping relationship between the pre-stack seismic response of tight sandstone reservoirs and their gas content. This includes the following steps:
[0055] S1. Establish a general paradigm for predicting the gas content of tight sandstone reservoirs using supervised learning methods based on pre-stack seismic gathers and well logging gas content curves;
[0056] S2. Establish a feature space-guided modeling method.
[0057] Step S2 includes the following sub-steps:
[0058] S21. Use the kNN method to make a preliminary gas content prediction for the target work area, and select the high-probability kNN prediction results as pseudo-training data.
[0059] S22. The CNN model is pre-trained using the pseudo-training data, wherein the information of the feature space of the data samples carried in the pseudo-training data is used to guide the establishment of the initial model of the CNN model.
[0060] S23. Optimize a stable CNN model using transfer learning with actual training samples to achieve gas content prediction in tight sandstone reservoirs using feature space-guided modeling.
[0061] In one embodiment, in step S1, the general paradigm is applicable to machine learning methods and deep learning methods. Further, in step S1, using the gas-bearing curve obtained from well logging interpretation as a label and the local waveform response as a feature, seismic sampling points are classified into gas-bearing and non-gas-bearing sampling points.
[0062] In one embodiment, in step S21, the kNN method performs a classification task by spatially partitioning the feature space based on labeled samples. Specifically, the kNN method selects k different nearest neighbor training samples for each test sample for decision-making. The number of nearest neighbors, k, is the value of k. When a small k value is selected, the kNN method builds a more targeted prediction model for a single test sample; when a large k value is selected, more training samples are introduced for reference. In step S22, in the initial model for building the CNN model, feature space information is passed to the CNN modeling process in the form of pseudo-training samples to supplement and constrain the information in the CNN modeling process. In step S23, the stable CNN model is obtained by repeatedly filtering pseudo-training data and training the CNN model.
[0063] In one embodiment, in steps S21-S23, the CNN model includes an input layer, a hidden layer, and an output layer; the hidden layer includes a convolutional layer, an activation function, a pooling layer, and a fully connected layer, and the convolutional layer contains the neuron connection method unique to convolutional neural networks; the backpropagation of the convolutional neural network adjusts the weights and biases by minimizing the residuals, so that the convolutional neural network is continuously updated according to the training data; wherein, the backpropagation algorithm uses the idea of stochastic gradient descent.
[0064] In one embodiment, in step S2, when the kNN method and the CNN model are applied individually to the task of predicting the gas content of tight sandstone reservoirs, they follow the same general paradigm for predicting the gas content of tight sandstone reservoirs based on supervised learning methods.
[0065] In this feature-space guided modeling method for predicting the gas content of tight sandstone reservoirs, the approach to addressing the small sample problem is to use the kNN method to initially predict gas content, and then pass the feature space information of the work area samples as pseudo-training samples to the CNN modeling process, thereby supplementing and constraining the CNN modeling process. Specifically, this feature-space guided modeling method predicts the large-scale lateral distribution of gas content in tight sandstone reservoirs based on the gas content curves of the work area and corresponding pre-stack seismic data. The high-probability prediction results of the kNN classifier are used to label the corresponding samples to be predicted, generating pseudo-training samples. Then, a reasonable CNN model is built and pre-trained using the pseudo-training samples. The feature space information carried in the pseudo-training samples guides the CNN modeling. Finally, the CNN gas content predictor is optimized through transfer learning using actual training samples, achieving feature-space guided modeling for predicting the gas content of tight sandstone reservoirs, alleviating the small sample problem of deep learning, and improving prediction accuracy.
[0066] Example 2
[0067] Please refer to Figure 1 This embodiment provides a method for predicting the gas content of tight sandstone reservoirs using feature space-guided modeling. The following description is from the perspective of technology development or implementation to facilitate reading and understanding; the aim is to clearly and completely explain the technical solution and highlight its inventive points and key features.
[0068] A method for predicting gas content in tight sandstone reservoirs using feature space-guided modeling includes the following steps:
[0069] First, a general paradigm for predicting the gas content of tight sandstone reservoirs based on pre-stack seismic gathers and well logging gas content curves is established (applicable to both machine learning and deep learning methods).
[0070] Then, to address the limitation of small sample size in predicting gas content in tight sandstone reservoirs using deep learning methods, a feature space-guided modeling method is established.
[0071] The specific approach to establishing the feature space-guided modeling method involves first using the k-Nearest Neighbor (kNN) method to perform preliminary gas-bearing predictions for the target work area; then, high-probability kNN predictions are selected as pseudo-training data. Next, the pseudo-training data is used to pre-train a Convolutional Neural Network (CNN) model. At this point, the pseudo-training data carries information about the feature space of the data samples, which can guide the CNN model to establish a reasonable initial model. Subsequently, the pseudo-training sample selection and CNN training process are repeated, and after the CNN model stabilizes, transfer learning is performed on the CNN model using actual training samples. Ultimately, this achieves intelligent gas-bearing prediction of tight sandstone reservoirs using feature space-guided modeling. The specific process is as follows: Figure 1 As shown.
[0072] In the feature space-guided modeling method for predicting the gas content of tight sandstone reservoirs of this invention, the prediction paradigm for tight sandstone reservoir gas content based on supervised learning methods includes:
[0073] Based on whether labeled data is used during model training, learning methods can be divided into supervised and unsupervised methods. Depending on the type of value to be predicted, they can also be categorized into classification and regression tasks. Due to the black-box nature of artificial intelligence, intelligent methods exhibit a unified paradigm when performing supervised classification tasks.
[0074] In the supervised learning-based paradigm for predicting gas content in tight sandstone reservoirs, seismic sampling points are classified into gas-bearing and non-gas-bearing categories. Using the gas-bearing curve obtained from well logging interpretation as a label and the local waveform response as a feature, seismic sampling points are further categorized into gas-bearing and non-gas-bearing categories.
[0075] In the process of creating training samples and labels, well logging data is first matched with seismic data using well-seismic calibration and resampling techniques, based on sonic logging curves and density logging curves. The seismic sampling point category is determined by the gas-bearing curve obtained from well logging interpretation. The well logging gas-bearing curve is obtained by secondary interpretation of the well logging gas saturation curve combined with the lithology, physical properties, electrical properties, and oil-bearing properties of the formation.
[0076] First, the gas saturation curve is obtained based on the rock physics model and conventional logging curve interpretation. Then, the gas saturation curve value for areas with mudstone or poor physical properties is set to zero. Therefore, the gas saturation curve is zero in non-reservoir sections and equal to the gas saturation curve in reservoir sections.
[0077] Practice has shown that the gas-bearing curve corresponds well with the location of the reservoir, but the value of the gas-bearing curve cannot indicate the gas storage capacity of the reservoir. Therefore, to simplify the study and better illustrate the problem, when the gas-bearing curve value is zero, the corresponding seismic data sampling points are classified as non-gas-bearing sampling points; when the gas-bearing curve value is not zero, they are classified as gas-bearing sampling points. Local waveform attributes are extracted from the pre-stack seismic trace set using a fixed time window as the features of the samples. The size of the time window is empirically set to be equal to the wavelet length of the seismic record vertically and equal to the number of offsets in a seismic trace set horizontally. During prediction, the same time window is used to extract the samples to be predicted from the seismic data to be predicted. When using the k-nearest neighbor method to complete the reservoir gas-bearing prediction problem, the category of the sample to be predicted is obtained by minimizing the expected misclassification cost function, which serves as the gas-bearing classification of the central sampling point of the sample to be predicted.
[0078]
[0079] in, The predicted classification result is given, where N represents the number of classes. C(y|j) refers to the prior error of classifying a sample as class j, while C(y|j) refers to the cost of classifying a sample as class y (y=1,...,N) when it is actually class j. It is determined based on the training samples and labels during the network training phase.
[0080] The performance of a model is evaluated using assessment criteria. For classification tasks, a confusion matrix is typically used to summarize the classifier's results, and basic metrics such as accuracy, precision, and recall are derived from the confusion matrix as evaluation criteria. The confusion matrix is a matrix used to analyze the classifier's results. For a binary classifier, its confusion matrix is shown in Table 1:
[0081] Table 1. Confusion Matrix of Binary Classifier
[0082]
[0083] Binary classification can be described as classifying samples into either a positive class or a negative class. The prediction results of a binary classifier can be categorized into four cases as shown in Table 1:
[0084]
[0085] Accuracy represents the overall probability of the model classifying correctly, and is a commonly used evaluation criterion. However, when the model is used in different scenarios or when the samples are imbalanced, people may be more concerned with the classification of a particular class of samples. In this case, accuracy as an evaluation criterion may become distorted. P represents the model's precision. R represents the model's recall. When it is necessary to comprehensively examine precision and recall, F is used. β measure:
[0086]
[0087] Among them, F β This represents a weighted average assessment of precision and recall. When β > 1, the focus is on recall; when β < 1, the focus is on precision; and when β = 1, both precision and recall are considered, resulting in the model's F1 score. In gas-bearing prediction problems, the β value should be adjusted appropriately according to different production stages. During the exploration stage, to characterize reservoir distribution and infer potential resources, the β value can be increased, emphasizing recall. During the development stage, considering drilling costs and other factors, the β value can be decreased, emphasizing precision to ensure the accuracy of the model's predictions of gas-bearing areas.
[0088] The preparation process of the dataset is described below. This invention makes predictions based on pre-stack seismic data and well logging gas-bearing curves, so the required raw data are pre-stack seismic data, seismic stratigraphic data, well stratigraphic data, and conventional well logging curves.
[0089] First, the data is preprocessed. Partial offset stacking of pre-stack seismic data helps suppress noise and enhance the effective information in the seismic data. It also reduces sample dimensionality. For well-seismic data consistency matching, preliminary well-seismic matching is performed using well layer data and seismic horizon data; then, based on the phase axis characteristics of the target layer, fine calibration is performed using sonic logging curves and density logging curves.
[0090] In this invention, the label is a well logging gas-bearing curve, which is obtained through further processing of the gas saturation curve. The gas saturation curve is a commonly used curve in well logging analysis. However, the gas saturation curve is generally a continuous value calculated based on rock physical relationships (empirical formulas or rock physical experiment fitting), which is difficult to accurately reflect the location of gas-bearing reservoirs. When the gas content of the target formation is obtained through cross-analysis of core data and corresponding well logging data from the cored sections, the gas-bearing characteristics of the conventional well logging curve are used. Combining the gas-bearing characteristics of the conventional well logging curve, the gas saturation value of the well section corresponding to the gas saturation curve is set to 1, and all other values are set to 0. Practice has shown that the gas-bearing curve obtained from well logging interpretation can indicate the location of gas-bearing reservoirs relatively well, but the curve value has a weak correlation with the actual gas reserves. Seismic data sampling points are labeled based on the gas-bearing curve. When the gas-bearing curve value corresponding to a sampling point is non-zero, the sampling point is classified as gas-bearing; otherwise, the sampling point is marked as non-gas-bearing.
[0091] The process of sample preparation is described below.
[0092] In this invention, samples are generated based on pre-stack seismic gathers. A time window is set, with the height of the time window being the length of a seismic wavelet and the width of the time window equal to the number of angles (or offsets) contained in a single gather. The time window is slid across the single-trace seismic data, and the seismic records within the time window range are used as samples corresponding to the sampling point at the center of the time window.
[0093] In this invention, when the kNN method and CNN model are applied individually to the task of predicting the gas content of tight sandstone reservoirs, they follow the same set of supervised learning-based gas content prediction paradigms for tight sandstone reservoirs.
[0094] In the feature space-guided modeling method for predicting gas content in tight sandstone reservoirs of this invention, regarding the kNN method and CNN model:
[0095] kNN is a supervised machine learning algorithm. Its principle is simple: during modeling, training samples and their corresponding labels are stored in pairs. In the prediction process, kNN selects the k nearest training samples to the test sample in the feature space based on a distance metric. These k closest training samples in the feature space are figuratively called the k nearest neighbors of the test sample.
[0096] Cover and Hart compared the performance of the nearest neighbor criterion on classification problems with the classic Bayesian optimal classifier. They assumed that the samples are independent and conform to the same data distribution. For a test sample x, there exists an arbitrary small positive number δ such that x always has a training sample z within its δ-neighborhood. That is, for any test sample, its nearest neighbor training sample can always be found in the training set. When the nearest neighbor criterion holds, the probability P(err) that the nearest neighbor classifier classifies x and z into different categories is:
[0097]
[0098] in, Y Let c represent all categories in this classification task, and c be the sample category predicted by the classifier. P(c|z) represents the probability of predicting the category of training sample z as c, and P(c|x) represents the probability of predicting the category of test sample x as c. Based on Bayesian theory, the test sample x is most likely classified into class c′, as follows:
[0099] c′=argmax c∈Y P(c∣x) (7)
[0100] The probabilities of errors by the nearest neighbor classifier and the Bayes optimal classifier are related as follows:
[0101]
[0102] As can be seen, the generalization error rate of the nearest neighbor classifier is no more than twice that of the Bayesian optimal classifier. This sufficiently demonstrates that the nearest neighbor criterion has quite good classification performance. However, in practical applications, to increase the stability of the method, the nearest neighbor criterion is extended to the kNN criterion.
[0103] Essentially, the kNN method performs classification tasks by partitioning the feature space based on labeled samples. The kNN method possesses a certain degree of non-linear characterization. This is because the kNN method selects different k nearest neighbor training samples for each test sample, which is equivalent to building a targeted prediction model for different test samples. Therefore, it is suitable for classifying rare events and performs well in multi-class problems. The number of nearest neighbors referenced during the decision-making process (i.e., the value of k) determines the strength of the non-linear characterization. When a small k value is selected, the kNN method builds a more targeted prediction model for a single test sample, thus introducing stronger non-linearity. However, an excessively small k value may lead to model instability and the risk of overfitting. When a large k value is selected, more training samples can be incorporated for reference, but this may lead to underfitting.
[0104] CNN is a deep learning algorithm inspired by animal visual systems. It contains a multi-layer network structure and can significantly reduce the complexity of deep neural network parameters or models during its application. Today, convolutional neural networks have achieved leading results in fields such as face recognition, handwriting recognition, and speech recognition.
[0105] Similar to traditional backpropagation neural networks, CNNs consist of an input layer, hidden layers, and an output layer. The hidden layers can be further refined into convolutional layers, activation functions, pooling layers, and fully connected layers. The convolutional layers incorporate the unique neuron connection methods of CNNs, using weight sharing to improve the performance of CNNs in problems such as classification.
[0106] Typically, the input layer is a two-dimensional matrix, followed by convolutional layers (feature extraction layers) and pooling layers placed alternately. The number of pooling layers and whether to place them are determined by the actual situation. A fully connected layer connects the last convolutional or pooling layer to the output layer. This network automatically extracts features multiple times, has a certain generalization ability, each layer consists of multiple independent neurons, its structure is simple, the network parameters are few, it has good robustness, and the training speed is fast.
[0107] Convolutional neural networks process input samples in two steps: forward propagation and backward propagation. During forward propagation, the activation value of each neuron is derived from the output of the previous layer through weighted multiplication and accumulation, and this process is passed sequentially to the next layer until the output layer is reached. Then, the error is calculated by comparing the network output with the true target value. The backward propagation phase starts from the output layer and gradually backtracks to the input layers, calculating the error gradient of each layer and adjusting the weights and biases accordingly to reduce the overall prediction error of the network.
[0108] Given a training dataset D containing m sample pairs, each sample pair consists of an input value and a corresponding expected value. The neural network consists of l layers. The input value x represents a single training instance, and the expected value y represents the expected output. The network output obtained after forward propagation is denoted as y'. If the mean squared error is chosen as the loss function, the error E of the neural network after inputting D can be expressed by formula (9):
[0109]
[0110] Where w and b represent the weights and biases of neurons in each layer. The error of the intermediate layers can be obtained based on the output layer using the chain rule. Taking the output layer (i.e., the l-th layer) as an example, the parameter update process is illustrated. First, the gradients of w and b are obtained using formulas (10) and (11):
[0111]
[0112]
[0113] Here, the superscript l of w and b represents the weights and biases of the l-th layer. The weights and biases are updated using the gradient descent algorithm according to these gradients:
[0114]
[0115] Here, α represents the learning rate, which takes a value between 0 and 1, adjusting the step size of gradient descent. Through repeated iterations, the value of the error function is gradually reduced until it reaches its minimum, at which point the neural network completes training.
[0116] In this invention, regarding the feature space guided modeling technique:
[0117] Feature space guidance refers to using the kNN method to perform preliminary classification of the data, that is, to initially divide the feature space of the samples.
[0118] This information is then passed to the CNN model through pseudo-training samples, and the CNN modeling process is supplemented and constrained by repeatedly generating pseudo-training samples randomly.
[0119] The specific approach is as follows: First, the gas-bearing potential of the target work area is predicted using the kNN method (a machine learning method). High-probability kNN predictions are then selected as pseudo-training data. This pseudo-training data is then used to pre-train a CNN model. At this stage, the pseudo-training data carries information about the feature space of the data samples, guiding the CNN model to establish a reasonable initial model. The process of selecting pseudo-training samples and training the CNN is repeated, and after the CNN model stabilizes, transfer learning is performed on the CNN model using actual training samples. Finally, a feature space-guided modeling intelligent gas-bearing potential prediction device for tight sandstone reservoirs is obtained.
[0120] Transfer learning transfers knowledge from a source domain to a target domain, enabling the target domain to achieve better learning results. The dataset containing prior knowledge is called the source domain, and the dataset from which the algorithm learns new knowledge is called the target domain. Essentially, in the transfer process, pre-training the network using the source data is equivalent to initially grouping the parameters and searching for local optima within each group. Combining this with the deep feature mining capabilities of CNNs, fine-tuning the deep network using target domain data is a process of comprehensively examining the grouped local models to obtain the global optimum. This allows the model to have more parameters while conserving computational resources.
[0121] In summary, this feature space-guided modeling method for predicting gas-bearing capacity in tight sandstone reservoirs alleviates the small sample size problem faced by deep learning methods in this task by employing feature space-guided modeling technology. It effectively integrates pre-stack seismic and well logging data, demonstrating the ability to accurately predict the morphology and location of tight reservoirs. This invention is of great significance for the scientific formulation of oilfield development plans and for increasing reserves and production.
[0122] Example 3
[0123] Please refer to Figures 1-9 The invention further illustrates a method for predicting gas content in tight sandstone reservoirs using feature space-guided modeling through case studies.
[0124] To verify the application effect of the method of the present invention, the gas-bearing prediction method of tight sandstone reservoir guided by the feature space modeling of the present invention was tested on a real pre-stack four-dimensional data. It shows that the present invention can alleviate the small sample problem, enable the deep learning method to be effectively applied to the gas-bearing prediction task of tight sandstone reservoir, and has the ability to accurately predict the morphology and location of tight reservoirs.
[0125] The test used actual data from a target area located in a work area in northern China. The target interval is a clastic sedimentary system of terrestrial fluvial-alluvial plain, with a thickness of 110-150m. The upper and middle lithology consists of equally thick interbedded light gray fine sandstone and argillaceous siltstone, while the lower part develops light gray and grayish-white medium and fine sandstone and brownish mudstone. The actual data for the target area includes pre-stack seismic data, seismic stratigraphic data, well-stratified data, and conventional logging curves. The single common reflection point gather contains 45 offsets (300-4500m). The seismic stratigraphic levels are L1 and L2 from top to bottom, and the reservoir is located between these two seismic levels. Disordered in-phase axes can be observed between L1 and L2, and the relationship between seismic response and reservoir gas content is complex. Partial offset stacking was performed using seismic data with offsets ranging from 500-4000m to form a pre-stack seismic common reflection point gather with 16 offsets.
[0126] First, a seismic profile is used to demonstrate the vertical resolution of the method. Figure 3 This displays a post-stack seismic profile. It can be observed that this profile exhibits characteristics common to the work area, with only one strongly negative polarity in-phase axis between layers L1 and L2. The gas-bearing distribution of the profile is predicted using the kNN method alone (e.g.,...). Figure 3 (As shown). Gas-bearing reservoirs exhibit a layered, continuous structure, posing a risk of underfitting. The CNN model is trained using only actual training samples and labels to predict well profiles (e.g., ...). Figure 4 As shown). The predicted reservoir distribution is scattered and discontinuous, which shows that the CNN model is affected by the few-sample problem. The CNN model was trained using only pseudo-training samples and labels and predicted well profiles (e.g.) Figure 5 As shown in the figure). It was observed that the predicted gas-bearing reservoirs were concentrated in certain areas. This phenomenon indirectly indicates that the pseudo-training samples can reflect the most significant differences between gas-bearing and non-gas-bearing samples. The sample feature space of the pseudo-training sample distribution provides a reasonable initial model for the CNN. The CNN model was trained simultaneously using two types of training samples and labels to predict well profiles (e.g., ...). Figure 6 (As shown). The vertical resolution of the predicted gas-bearing areas is improved, but the horizontal distribution is contiguous, which does not conform to the actual reservoir morphology. This is because when two types of training data are input into the CNN model simultaneously, the difference information in the pseudo-training data may act as interference. The gas-bearing distribution profile obtained by the intelligent reservoir gas-bearing prediction method using feature space-guided modeling is shown below. Figure 7 As shown. The predicted reservoir area matches the prior geological information of the work area. Based on Figure 7 The two thin reservoir locations shown are plotted to illustrate the predicted lateral distribution of gas-bearing properties in the tight sandstone reservoirs (e.g., ...). Figure 8 and Figure 9(As shown). The two thin reservoirs of the target layer are located 19 ms and 13 ms above L2, respectively. Overall, the predicted gas-bearing area is in good agreement with the geological background of the work area. The agreement rate with the well logging gas-bearing curve at the well location is over 85%.
[0127] Example 4
[0128] Please refer to Figure 12 , Figure 13 This embodiment provides a feature space-guided modeling device for predicting the gas content of tight sandstone reservoirs, which adopts the following structure:
[0129] 1. First module creation
[0130] The first module 10 is used to establish a general paradigm for supervised learning methods to predict the gas content of tight sandstone reservoirs based on pre-stack seismic gathers and well logging gas content curves.
[0131] 2. Second module creation
[0132] The second module 20 is used to establish a feature space-guided modeling method.
[0133] In one embodiment, the second establishment module 20 adopts the following structure:
[0134] 1. Filtering Submodule
[0135] The filtering submodule 201 is used to perform preliminary gas content prediction of the target work area using the kNN method, and to filter high-probability kNN prediction results as pseudo-training data.
[0136] 2. Pre-training submodule
[0137] The pre-training submodule 202 is used to pre-train the CNN model using the pseudo-training data, wherein the information of the feature space of the data samples carried in the pseudo-training data is used to guide the establishment of the initial model of the CNN model.
[0138] 3. Transfer Learning Submodule
[0139] The transfer learning submodule 203 is used to optimize a stable CNN model using actual training samples through transfer learning, thereby enabling gas content prediction of tight sandstone reservoirs by feature space-guided modeling.
[0140] It should be noted that the apparatus of the present invention is used to implement the methods in the above embodiments, and each module in the apparatus corresponds to each step in the method.
[0141] Example 5
[0142] Based on the same inventive concept, one embodiment of the present invention provides an electronic device, including: a memory and a processor; wherein the memory is used to store one or more computer instructions; the one or more computer instructions are executed by the processor using any of the methods described in the above embodiments.
[0143] Example 6
[0144] Based on the same inventive concept, one embodiment of the present invention provides a readable storage medium storing computer instructions; wherein, when the computer instructions are executed by a processor, they implement the method of any one of the above embodiments.
[0145] One or more of the aforementioned computer instructions can form a program.
[0146] The aforementioned program can run on a processor or be stored in memory (or computer-readable medium). Computer-readable medium includes both permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable medium does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0147] These computer programs may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes can be implemented using different modules, and different steps can be implemented using different modules.
[0148] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for predicting gas content in tight sandstone reservoirs using feature space-guided modeling, characterized in that, include: Establish a general paradigm for predicting the gas content of tight sandstone reservoirs using supervised learning methods based on pre-stack seismic gathers and well logging gas content curves; A feature space-guided modeling method is established, wherein the feature space-guided modeling method includes: The kNN method was used to make a preliminary gas content prediction for the target work area, and the high-probability kNN prediction results were selected as pseudo-training data. The CNN model is pre-trained using the pseudo-training data, wherein the information of the feature space of the data samples carried in the pseudo-training data is used to guide the establishment of the initial model of the CNN model. By using actual training samples to optimize a stable CNN model through transfer learning, gas-bearing prediction of tight sandstone reservoirs can be achieved through feature space-guided modeling.
2. The method for predicting gas content in tight sandstone reservoirs using feature space-guided modeling according to claim 1, characterized in that, In establishing a general paradigm for predicting the gas content of tight sandstone reservoirs based on pre-stack seismic gathers and well logging gas content curves, the general paradigm is applicable to both machine learning and deep learning methods; wherein... Using the gas-bearing curves obtained from well logging interpretation as labels and local waveform responses as features, seismic sampling points are classified into gas-bearing and non-gas-bearing sampling points.
3. The method for predicting gas content in tight sandstone reservoirs using feature space-guided modeling according to claim 1, characterized in that, The kNN method performs a classification task by partitioning the feature space based on labeled samples. Specifically, the kNN method selects k different nearest neighbor training samples for each test sample to make a decision. The number of nearest neighbors, k, is the value of k. When a small k value is selected, the kNN method builds a more targeted prediction model for a single test sample. When a large k value is selected, more training samples are introduced for reference.
4. The method for predicting gas content in tight sandstone reservoirs using feature space-guided modeling according to claim 1, characterized in that, In the initial model for establishing the CNN model, feature space information is passed to the CNN modeling process in the form of pseudo-training samples to supplement and constrain the information in the CNN modeling process.
5. The method for predicting gas content in tight sandstone reservoirs using feature space-guided modeling according to claim 1, characterized in that, The stable CNN model is obtained through repeated pseudo-training data filtering and CNN model training process.
6. The method for predicting gas content in tight sandstone reservoirs using feature space-guided modeling according to claim 1, characterized in that, The CNN model includes an input layer, hidden layers, and an output layer. The hidden layers include convolutional layers, activation functions, pooling layers, and fully connected layers. The convolutional layers contain the neuron connection methods unique to convolutional neural networks. The backpropagation of the convolutional neural network adjusts the weights and biases by minimizing the residuals, so that the convolutional neural network is continuously updated according to the training data. The backpropagation algorithm uses the idea of stochastic gradient descent.
7. The method for predicting gas content in tight sandstone reservoirs using feature space-guided modeling according to any one of claims 1-6, characterized in that, When the kNN method and the CNN model are applied individually to the task of predicting the gas content of tight sandstone reservoirs, they follow the same general paradigm for predicting the gas content of tight sandstone reservoirs based on supervised learning methods.
8. A device for predicting the gas content of tight sandstone reservoirs using feature space-guided modeling, characterized in that, include: The first module is used to establish a general paradigm for supervised learning methods to predict the gas content of tight sandstone reservoirs based on pre-stack seismic gathers and well logging gas content curves. and The second module is used to establish a feature space-guided modeling method. The second establishment module includes: The filtering submodule is used to perform preliminary gas content prediction of the target work area using the kNN method, and to filter high-probability kNN prediction results as pseudo-training data. A pre-training submodule is used to pre-train the CNN model using the pseudo-training data, wherein the information carried in the pseudo-training data regarding the feature space of the data samples is used to guide the establishment of an initial model of the CNN model; and The transfer learning submodule is used to optimize a stable CNN model using actual training samples through transfer learning, thereby enabling gas-bearing prediction of tight sandstone reservoirs through feature space-guided modeling.
9. An electronic device, characterized in that, include: Memory; and processor; The memory is used to store one or more computer instructions; the one or more computer instructions are executed by the processor to implement the method according to any one of claims 1 to 7.
10. A readable storage medium, characterized in that, The readable storage medium stores computer instructions; wherein, when executed by a processor, the computer instructions implement the method described in any one of claims 1 to 7.