A method and system for quantitative seismic prediction of free oil porosity
By combining high-resolution reservoir prediction technology and deep learning random forest algorithm with nuclear magnetic resonance logging and seismic inversion data, a seismic prediction model for free oil porosity was established, which solved the problem of accurate prediction of free oil porosity in the region and enabled accurate identification of shale oil sweet spots and optimization of drilling schemes.
Patent Information
- Application Number
- CN202310653144.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-02
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2043-06-02
AI Technical Summary
Existing technologies struggle to accurately predict free oil porosity in shale oil reservoirs within a regional area, and the relationship between commonly used seismic elastic parameters and free oil porosity is difficult to establish, affecting the precise evaluation of shale oil sweet spots and the identification of high-quality sweet spots.
By employing high-resolution reservoir prediction technology and deep learning random forest algorithm, combined with nuclear magnetic resonance logging data and seismic inversion data, a seismic prediction method for free oil porosity is established through K-means clustering and random forest model. Multiple predictions are performed using seismic elastic parameters and sensitive seismic attributes to form an accurate free oil porosity prediction model.
It achieves a quantitative prediction accuracy of 90% for free oil porosity, which is more than 20% higher than conventional methods, and supports the accurate identification and regional prediction of high-quality sweet spots in shale oil.
Smart Images

Figure CN119065008B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of oil and gas exploration technology, and specifically relates to a method and system for quantitative seismic prediction of free oil porosity. Background Technology
[0002] my country possesses abundant shale oil resources, with recoverable reserves exceeding 13 billion tons, representing a significant resource potential. Accelerating the exploration and development of shale oil resources is a practical way to ensure national energy security. Shale oil can be classified into adsorbed oil and free oil based on its occurrence state. Free oil can be effectively developed through volumetric fracturing, while adsorbed oil, due to its high adsorption energy, requires in-situ conversion into free oil for effective extraction. Therefore, quantitative characterization and regional range prediction of free oil porosity are core technologies for the accurate location and evaluation of high-quality sweet spots in shale oil. However, the porosity of adsorbed and free oil is influenced by various factors such as shale reservoir pore size, crude oil composition, mineral composition, and the development of thin interbedded layers. Continuous quantitative characterization and large-scale prediction are challenging, severely impacting the accurate evaluation of shale oil sweet spots and hindering the improvement of the drilling rate of high-quality sweet spots in horizontal wells and the optimization of segmentation and clustering schemes.
[0003] Currently, existing technologies based on experimental and nuclear magnetic resonance logging techniques have solved the problem of continuous characterization of free oil porosity by accurately calculating the porosity of a single well. These technologies have shown good results in the Jimsar Shale Oil National Demonstration Zone and the Mabei Fengcheng Formation Shale Oil. However, relying solely on data from a single well is insufficient to effectively control the overall sweet spot of shale oil in a region. Accurate prediction of free oil porosity within a regional area is still needed to accurately identify and characterize high-quality sweet spots in shale oil production.
[0004] However, there is currently no mature technology capable of quantitatively predicting free oil porosity. Commonly used seismic prediction techniques often establish a relationship between seismic elastic parameters and free oil porosity, indirectly predicting free oil porosity by inverting the derived elastic parameters. However, the effectiveness of such methods relies on a clear correspondence between seismic elastic parameters and free oil porosity, which is often difficult to obtain in practice. Seismic elastic parameters are influenced by various geological properties, reflecting a comprehensive response to subsurface conditions, including the influence of free oil porosity, as well as other factors such as fluids and the rock skeleton. Therefore, extracting the relationship between seismic elastic parameters and free oil porosity is crucial for quantitative prediction. Summary of the Invention
[0005] To achieve accurate prediction of free oil porosity in shale oil sweet spots within a regional area and meet the current needs of fine-target exploration and development of shale oil, this invention, based on in-depth research into existing free oil prediction methods, proposes a seismic prediction method and system for free oil porosity using high-resolution reservoir prediction technology and deep learning random forest algorithm, targeting the characteristics of shale oil reservoirs.
[0006] The technical solution adopted in this invention is as follows: a method for quantitative seismic prediction of free oil porosity, the method comprising the following steps:
[0007] 1) Utilize nuclear magnetic resonance logging to calculate free oil porosity data for shale oil formations in a single well, and obtain high-quality near-, mid-, and far-offset partially superimposed seismic data;
[0008] 2) Use shale oil rock physical interpretation charts to clarify the relationship between seismic elastic parameters and free oil porosity, and perform targeted processing on the high-resolution seismic inversion elastic data obtained from pre-stack statistical inversion to form machine learning prediction feature value data for free oil porosity.
[0009] 3) Perform K-means cluster analysis on sensitive seismic attribute data volumes and targeted seismic inversion elastic data volumes to establish a random forest free oil content prediction model;
[0010] 4) Obtain the initial prediction data volume of free oil porosity based on the random forest free oil content prediction model, and then input it as the feature value into the random forest free oil porosity prediction model again to complete the second free oil porosity prediction and establish the KRR free oil porosity seismic prediction model.
[0011] 5) Input the machine learning prediction feature value data for free oil porosity and the sensitive seismic attribute data body into the KRR free oil porosity seismic prediction model to obtain the final prediction data body for free oil porosity;
[0012] 6) Verify the final predicted data volume of free oil porosity, and quantitatively characterize the high-quality sweet spots of shale oil within the work area by combining the seismic inversion elastic parameter volume.
[0013] Further, step 2) specifically involves: constructing a rock physical interpretation chart of the shale oil reservoir through rock physical analysis, and determining, based on the chart, the longitudinal wave impedance corresponding to a high-quality sweet spot with an oil saturation >15% that is greater than 9200 g / cm. 3 *m / s and less than 13000g / cm 3 *m / s, P-wave velocity ratio less than 1.85; The high-resolution P-wave impedance and high-resolution P-wave velocity ratio data obtained from pre-stack statistical inversion are processed to remove data corresponding to non-sweet spots in mudstone and tight strata outside the range. Then, various mathematical transformations are performed on the processed data to form machine learning prediction feature value data for free oil porosity.
[0014] Furthermore, in step 3), the K-means clustering analysis is performed using the K-means algorithm, which includes the following steps:
[0015] ① Initialization: Randomly select K sample points as K initial cluster centers;
[0016] ② Cluster the samples: Calculate the distance from each sample to each cluster center, and assign each sample to the cluster containing its nearest neighboring cluster center; the distance between cluster centers is calculated using Euclidean distance d, as shown in the following formula:
[0017]
[0018] Where, x i Let y represent the x-coordinate of the i-th point in the Cartesian coordinate system. i This represents the ordinate value of the i-th point in the rectangular coordinate system;
[0019] ③ Calculate new cluster centers: Calculate the mean of all samples in each cluster as the new cluster centers;
[0020] ④ Stop iterating if the cluster centers do not change or the cutoff condition is reached; otherwise, repeat steps ② and ③.
[0021] Furthermore, random forests are classifiers based on decision trees. The ID3 algorithm is used to calculate the entropy of data feature samples to evaluate the branches of the decision trees. The formula for calculating the entropy of data feature samples is as follows:
[0022]
[0023] Where H(X) is the entropy of the data feature samples, and P i (X) is the probability function of the data features, and n is the number of branch layers;
[0024] Each parent node is divided into multiple child nodes. Each node has a data feature sample entropy. The difference between the data feature sample entropy of the parent node and the weighted average of the data feature sample entropies of all its child nodes is calculated as the data gain optimization decision tree.
[0025]
[0026]
[0027] Where k is the number of nodes, N(x j I(x) represents the number of samples in the lower-level nodes, N represents the total number of samples in the corresponding upper-level nodes, and I(x) represents the number of samples in the lower-level nodes. j Data feature sample entropy; and These represent the entropy of the parent node and the child node, respectively, while Data Gain is the average difference between the entropy of the parent node and the child node.
[0028] Further, step 1) includes: using nuclear magnetic resonance logging to calculate the free oil porosity data of a single-well shale oil layer, wherein a portion of the single wells are marked as learning wells and the remaining single wells are marked as verification wells. The selection criteria for learning wells are that they are representative wells in the work area, can reflect the geological characteristics of the target work area and are evenly distributed within the work area; and obtaining high-quality near, medium and far offset partial superimposed seismic data.
[0029] Furthermore, pre-stack statistical inversion is based on the Markov chain Monte Carlo algorithm, which converges to the population within the actual time bound. During the inversion process, the variogram, probability distribution function, well logging, seismic and lithofacies information are integrated. A corresponding probability distribution model is established based on seismic data analysis. The variogram and probability distribution function are obtained through geological understanding and well data analysis. Starting from the well points, the well points are made to conform to the original seismic information. The inversion software uses a stochastic simulation module to generate inter-well impedance. The wavelet extracted by deterministic inversion is convolved with the reflection coefficient converted from the inter-well impedance to generate a composite seismic record. Then, a new wavelet is extracted to generate a new composite record. This process is repeated iteratively until the matching rate between the composite seismic trace and the original seismic data reaches more than 85%, so that the correlation coefficient meets the requirements of pre-stack inversion. Finally, the inversion parameters are adjusted to complete the inversion calculation.
[0030] Furthermore, the variation function is a distance function that describes the change of the spatial distribution characteristics of a certain attribute with distance, and is expressed as the semivariance of the increment of the regionalized variable Z(x) at two points x and x+h.
[0031]
[0032] Where r(x,h) is the experimental variation function, x represents the absolute position of a point in space, and h represents the lag distance of the variation function.
[0033] Furthermore, the three characteristic values of the variogram—the sill value, the range, and the nugget constant—reflect the spatial variation characteristics of the reservoir parameters; the range is the spatial distance at which the variogram reaches a certain stable value, reflecting the average scale of the regionalized variable carrier in a certain direction or the predicted extension scale of the sand body in a certain direction.
[0034] In addition, the present invention also provides a quantitative seismic prediction system for free oil porosity, the system comprising:
[0035] The calculation module is used to calculate the porosity data of free oil in a single well shale oil layer using nuclear magnetic resonance logging, and to obtain high-quality near-, medium- and far-offset partially superimposed seismic data;
[0036] The sample-specific processing module is used to clarify the relationship between seismic elastic parameters and free oil porosity using shale oil rock physical interpretation charts, and to perform sample-specific processing on the high-resolution seismic inversion elastic data obtained from pre-stack statistical inversion to form machine learning prediction feature value data for free oil porosity.
[0037] The K-means clustering analysis module is used to perform K-means clustering analysis on sensitive seismic attribute data volumes and targeted processed seismic inversion elastic data volumes to establish a random forest free oil content prediction model.
[0038] A module is established to obtain the initial prediction data volume of free oil porosity based on the random forest free oil content prediction model, and then input it as the feature value into the random forest free oil porosity prediction model again to complete the second free oil porosity prediction, and establish the KRR free oil porosity seismic prediction model.
[0039] The free oil porosity final prediction data body module is used to input the free oil porosity machine learning prediction feature value data and the sensitive seismic attribute data body into the KRR free oil porosity seismic prediction model to obtain the free oil porosity final prediction data body.
[0040] The verification module is used to verify the final prediction data of free oil porosity and to quantitatively characterize the high-quality sweet spots of shale oil within the work area by combining the seismic inversion elastic parameter volume.
[0041] Furthermore, the sample-specific processing module includes a determining unit and a processing unit;
[0042] The determining unit is used to construct a petrophysical interpretation chart of shale oil reservoirs through rock physical analysis. Based on the chart, the longitudinal wave impedance corresponding to a high-quality sweet spot with an oil saturation >15% is determined to be greater than 9200 g / cm. 3 *m / s and less than 13000g / cm 3 *m / s, P-wave to S-wave velocity ratio less than 1.85;
[0043] The processing unit is used to process the high-resolution P-wave impedance and high-resolution P-wave / S-wave velocity ratio data obtained from pre-stack statistical inversion, remove the data corresponding to non-sweet spots in mudstone and tight strata outside the range, and then perform various mathematical transformations on the processed data to form machine learning prediction feature value data for free oil porosity.
[0044] This invention establishes a method and system for quantitative seismic prediction of free oil porosity, overcoming the challenge of accurately predicting free oil porosity over a large area in shale oil. During the prediction process, high-resolution seismic elastic parameter data volumes derived from pre-stack statistical inversion are used, and a set of machine learning prediction feature values for free oil porosity is constructed based on rock physical analysis, improving the predictive ability of free oil porosity in thin sweet spots. Finally, using the KRR free oil porosity seismic prediction model, the free oil porosity information in seismic elastic parameters and sensitive seismic attributes is deeply mined to achieve quantitative prediction of free oil porosity, with a prediction accuracy of 90%, which is more than 20% higher than the prediction accuracy of conventional methods.
[0045] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention will be realized and obtained from the description and the drawings. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0047] Figure 1 A flowchart of a method for quantitative seismic prediction of free oil porosity;
[0048] Figure 2 This is a schematic diagram of the fitting process in the analysis of variation.
[0049] Figure 3a This is the pre-stack statistical inversion result of the longitudinal wave impedance profile; Figure 3b Prestack statistical inversion results of P-wave and S-wave velocity ratio profiles;
[0050] Figure 4 A petrophysical chart of shale oil reservoirs;
[0051] Figure 5a The eigenvalue data of the processed high-resolution longitudinal wave impedance profile; Figure 5b The characteristic value data of the processed high-resolution P-wave and S-wave velocity ratio profile;
[0052] Figure 6 The results of K-means cluster analysis are shown in the figure.
[0053] Figure 7 Flowchart for constructing a seismic prediction model for free oil porosity;
[0054] Figure 8a A comparison was made between the free oil porosity fitted by multiple regression of elastic parameters and the free oil porosity calculated by nuclear magnetic resonance logging. Figure 8b This is a comparison between the predicted free oil porosity and the free oil porosity calculated by nuclear magnetic resonance logging in this embodiment of the invention;
[0055] Figure 9 A quantitative prediction profile of free oil porosity in shale oil layers in a case study area;
[0056] Figure 10 A comparison chart of free oil porosity prediction results and nuclear magnetic resonance logging interpretation results;
[0057] Figure 11a Planar distribution diagram of free oil porosity in the main layer of the No. 1 sweetener; Figure 11b Planar distribution diagram of free oil porosity in the main layer of the No. 2 sweetener; Figure 11c This is a planar distribution diagram of the free oil porosity in the main layer of the lower sweet spot No. 1. Figure 11d This is a planar distribution diagram of the free oil porosity in the main layer of the lower sweet spot No. 2.
[0058] Figure 12 This is a schematic diagram of a quantitative seismic prediction system for free oil porosity. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0060] A method for quantitative seismic prediction of free oil porosity, such as Figure 1 As shown, the method includes the following steps:
[0061] 1) Utilize nuclear magnetic resonance logging to calculate free oil porosity data for shale oil formations in a single well, and obtain high-quality near-, mid-, and far-offset partially superimposed seismic data;
[0062] 2) Use shale oil rock physical interpretation charts to clarify the relationship between seismic elastic parameters and free oil porosity, and perform targeted processing on the high-resolution seismic inversion elastic data obtained from pre-stack statistical inversion to form machine learning prediction feature value data for free oil porosity.
[0063] 3) Perform K-means cluster analysis on sensitive seismic attribute data volumes and targeted seismic inversion elastic data volumes to establish a random forest free oil content prediction model;
[0064] 4) The initial prediction data volume of free oil porosity is obtained based on the random forest free oil content prediction model. This data is then used as feature values to input into the random forest free oil porosity prediction model again, completing the second free oil porosity prediction and establishing the KRR free oil porosity seismic prediction model. K refers to K-means clustering analysis, and R refers to the random forest machine learning algorithm. This method uses K-means clustering analysis and the free oil porosity prediction model obtained by the two random forest machine learning algorithms, hence the name KRR free oil porosity seismic prediction model.
[0065] 5) Input the machine learning prediction feature value data for free oil porosity and the sensitive seismic attribute data body into the KRR free oil porosity seismic prediction model to obtain the final prediction data body for free oil porosity;
[0066] 6) Verify the final predicted data volume of free oil porosity, and quantitatively characterize the high-quality sweet spots of shale oil within the work area by combining the seismic inversion elastic parameter volume.
[0067] Step 1) Specifically involves: using nuclear magnetic resonance logging to calculate the free oil porosity data of a single-well shale oil formation. Some of these single wells are marked as learning wells, and the remaining single wells are marked as verification wells. The selection criteria for learning wells are that they are representative wells in the work area, can reflect the geological characteristics of the target work area, and are evenly distributed within the work area. High-quality near-, mid-, and far-offset partial superimposed seismic data are obtained. Using high-quality seismic data to predict free oil porosity can fundamentally ensure the accuracy of the prediction results.
[0068] Step 2) specifically involves: based on the relationship between free oil porosity and seismic elastic parameters analyzed from wellbore, selecting P-wave impedance and P-wave / S-wave velocity ratio, which show good correlation with free oil porosity, as sensitive parameters for free oil porosity. Through rock physical analysis, a rock physical interpretation chart of the shale oil reservoir is constructed (e.g., Figure 4 As shown), according to Figure 4 The chart shown indicates that high-quality desserts with an oil saturation of >15% have a longitudinal wave impedance greater than 9200 g / cm. 3 *m / s and less than 13000g / cm 3 *m / s, P-wave velocity ratio less than 1.85. Based on the analysis results, the high-resolution P-wave impedance and high-resolution P-wave velocity ratio data obtained from pre-stack statistical inversion were processed to remove data corresponding to non-sweet spots in mudstone and tight strata outside the specified range. Figure 5a and Figure 5b These are the high-resolution P-wave impedance profile and high-resolution P-wave / S-wave velocity ratio profile processed using the above method, respectively. The corresponding data for mudstone and tight strata are set to 0.0001, such as 5a and... Figure 5bThe white area in the middle is shown. Then, each processed data is subjected to various mathematical transformations (squaring, taking the reciprocal, taking the logarithm, etc.) to form a set of machine learning prediction feature value data for free oil porosity.
[0069] Table 1 shows the correlation between the eigenvalue data and ordinary elastic parameters of this invention and free oil porosity. As can be seen from the data in Table 1, compared with the high-resolution P-wave impedance and high-resolution P-wave / S-wave velocity ratio data obtained by inversion, the set of machine learning prediction eigenvalue data for free oil porosity constructed in this invention has a higher correlation with free oil porosity, which improves the stability of the machine learning algorithm and thus improves the prediction accuracy of free oil porosity.
[0070] Table 1. Correlation statistics between the eigenvalue data and ordinary elastic parameters of this invention and free oil porosity.
[0071]
[0072] Pre-stack statistical inversion, based on the Markov chain Monte Carlo algorithm that converges to the population within the actual time bound, is an inversion method that organically combines stochastic inversion techniques with pre-stack simultaneous inversion techniques. During the inversion process, information such as variograms, probability distribution functions, well logging data, seismic data, and lithofacies data are integrated. A corresponding probability distribution model is established based on the actual seismic data. Variation and probability distribution functions are obtained through geological understanding and well data analysis. Starting from the well points, ensuring they conform to the original seismic information, and using appropriate inversion software, inter-well impedance is generated using a stochastic simulation module. The wavelet extracted by deterministic inversion is convolved with the reflection coefficient converted from the inter-well impedance to generate a composite seismic record. Then, a new wavelet is extracted to generate a new composite record. This process is iterated until the matching rate between the composite seismic trace and the original seismic data reaches over 85%, meaning the correlation coefficient meets the requirements for pre-stack inversion. Finally, the inversion parameters are adjusted to complete the inversion calculation. Figure 3a and Figure 3b These are high-resolution P-wave impedance profiles and high-resolution P-wave / S-wave velocity ratio profiles obtained from pre-stack statistical inversion, respectively. Since the high frequency of the inversion results originates from well logging data and the lateral trend is controlled by seismic activity, thinner shale oil sweet spots can be continuously identified on the inversion profiles.
[0073] To ensure that stochastic simulations meet the conditions for establishing statistical correlation functions between spatial reservoir parameter points, a constructed variogram function can be used in the spatial data field to describe the relationships between different data points. The variogram function r(x,h) is a distance function that describes the spatial distribution characteristics of a certain attribute as a function of distance. It can be expressed as the semivariance of the increment of the regionalized variable Z(x) at points x and x+h:
[0074]
[0075] In equation (1), the vertical axis of r(x,h) is the experimental variation function, and the horizontal axis of h is the lag distance of the experimental variation function. Figure 2 The diagram shown is a schematic of the analysis of variation fitting. Figure 2 The reservoir contains three main characteristic values: nugget constant, sill value, and range. These values can typically be obtained by fitting the experimental variogram with a theoretical model. Geostatistical inversion first establishes a geological model of the reservoir within the seismic time domain based on geological conditions. Seismic horizons determine the bedding planes, and the original acoustic impedance curves at the well locations need to be placed within the stratigraphic grid. Then, inversion parameters are determined using seismic data and well data, leading to geostatistical inversion calculations. The three main characteristic values of the variogram—sill value, range, and nugget constant—can be used to reflect the spatial variation characteristics of reservoir parameters. The range is the spatial distance at which the variogram reaches a certain stable value. It reflects the average scale of the regionalized variable carrier in a certain direction and can also represent the predicted extension scale of the sand body in a certain direction, thus achieving the goal of predicting the sand body size.
[0076] This invention employs stochastic inversion technology, which requires normally distributed input data. Therefore, in addition to determining the parameters required for conventional seismic inversion and the variograms (also called variance functions) of the P-wave and S-wave impedances and inversion parameters, the correlation coefficients between densities, and the parameter distribution histograms determined through statistical analysis, appropriate transformation methods are needed to convert the non-normal actual data into normally distributed data. The variogram is typically characterized by parameters such as minimum range, maximum range azimuth, vertical range, and maximum range. Then, statistical methods are used to fit the variogram to obtain high-resolution P-wave impedance, P-wave / S-wave velocity ratio, and other seismic inversion elastic data. The selection of inversion parameter combinations is based on the actual conditions of the work area (as shown in Table 2), ensuring accurate identification of thin sweet spots with a thickness of approximately 4m in the inversion results.
[0077] Table 2 Pre-stack statistical inversion parameters
[0078] Horizontal (unit: meter) Vertical (unit: milliseconds) Minimum range 600 2 Maximum range 1500 4
[0079] Step 3) specifically involves: performing K-means clustering analysis on the seismic attributes and their mathematical transformations extracted from the earthquake, and on the set of machine learning prediction feature values for free oil porosity constructed in Step 2). The sample values are divided into K major categories. Based on the sample values of the K major categories, K random forest free oil porosity prediction models are established to predict the first multi-category free oil porosity prediction results. Then, using the multi-category free oil porosity results and their mathematical transformations as samples, a second high-precision prediction using the random forest model is achieved to predict the final free oil porosity prediction results, resulting in the KRR free oil porosity earthquake prediction model. Figure 7The flowchart illustrating the construction process of the aforementioned free oil porosity seismic prediction model is shown. First, the four seismic attributes obtained from K-cluster analysis are input as feature values into the random forest free oil porosity prediction model, resulting in four initial free oil porosity prediction data sets. Then, these four initial free oil porosity prediction data sets are input as feature values into the random forest free oil porosity prediction model again, yielding the final free oil porosity prediction data set. Combining these two random forest machine learning methods for free oil porosity allows the prediction model to fully utilize the input feature value information, improving the accuracy of free oil porosity prediction.
[0080] The K-means algorithm iteratively adjusts the positions of cluster centers and the assignments of all sample points in the direction that minimizes the sum of the distances of all samples in the same class from the cluster center, until the optimal K cluster centers are found and all samples have found their optimal assignments.
[0081] The K-means algorithm consists of four steps:
[0082] ① Initialization: Randomly select K sample points as K initial cluster centers;
[0083] ② Clustering of samples: Calculate the distance from each sample to each cluster center, and classify each sample into the class containing the nearest cluster center. The distance calculation method used in this invention is Euclidean distance d, as shown in formula (2):
[0084]
[0085] Where, x i Let y represent the x-coordinate of the i-th point in the Cartesian coordinate system. i This represents the ordinate value of the i-th point in the rectangular coordinate system;
[0086] ③ Calculate new cluster centers: Calculate the mean of all samples in each cluster as the new cluster centers;
[0087] ④ Stop iterating if the cluster centers no longer change or the cutoff condition is reached; otherwise, repeat steps ② and ③.
[0088] K-means cluster analysis was used to classify the seismic attributes and the targeted high-resolution inverted elastic parameter volumes and their mathematical transformations. Four cluster centers were selected, dividing the data into four major categories. The classification results are as follows: Figure 6 As shown, the four selected cluster centers are as follows: Figure 6 As shown in the pentagram, the seismic attributes and targeted high-resolution inverted elastic parameter data are divided into 4 categories according to the K-means clustering algorithm. Data of the same category are represented by the same symbol, are close to the cluster center, and have similar properties.
[0089] In machine learning, random forests are classifiers based on decision trees. The key to their algorithm lies in determining the splitting attributes of decision tree nodes. There are many node splitting algorithms. The ID3 algorithm is suitable for classifying discretized data such as seismic data and well logging data. It evaluates the decision tree branches by calculating the entropy balance of data samples. The convergence of the decision tree is the process of the entropy decreasing to a certain value and then changing smoothly. The formula for calculating the entropy of data feature samples is shown in (3):
[0090]
[0091] Where H(X) is the entropy of the data feature samples, and P i (X) is the probability function of the data features, and n is the number of branch layers;
[0092] Each parent node is divided into multiple child nodes. Each node has a data feature sample entropy. The difference between the data feature sample entropy of the parent node and the weighted average of the data feature sample entropies of all its child nodes is calculated as the data gain optimization decision tree.
[0093]
[0094]
[0095] Where k is the number of nodes, N(x j I(x) represents the number of samples in the lower-level nodes, and N represents the total number of samples in the corresponding upper-level nodes; j Data feature sample entropy; and These represent the entropy of the parent node and the child node, respectively, while Data Gain is the average difference between the entropy of the parent node and the child node.
[0096] Based on the aforementioned principles, this invention establishes a random forest machine learning model. Free oil porosity calculated from nuclear magnetic resonance logging of learning wells is used as the training sample to train the random forest model. The four clustered data categories are then input as feature values into the trained random forest free oil porosity prediction model, yielding four initial free oil porosity prediction datasets. These datasets are then used as input feature values into the random forest free oil porosity prediction model to complete the second free oil porosity prediction, resulting in the KRR free oil porosity seismic prediction model established in this invention.
[0097] Figure 8a The medium-dark gray curve (representing the predicted value) is the free oil porosity curve obtained by multiple regression of parameters such as measured density, P-wave velocity, S-wave velocity, and wave velocity ratio in Well 1. It has a large error compared with the free oil porosity of the shale oil layer calculated by nuclear magnetic logging (light gray curve, representing the measured value), and the two have a consistency of only 64%. Figure 8bThe medium-dark gray curve (representing the predicted value) is the free oil porosity curve predicted using the random forest free oil porosity quantitative prediction method designed in this invention. It shows a high degree of consistency with the free oil porosity of the shale oil layer calculated by nuclear magnetic resonance logging in Well 1 (light gray curve, representing the measured value), with a similarity of up to 89%. This also proves that the free oil porosity prediction accuracy designed in this invention is far higher than that of general methods.
[0098] Based on the actual exploration and development situation, the shale oil strata in the example work area can be roughly divided into two sweet spots: the upper sweet spot and the lower sweet spot. Each sweet spot is further divided into multiple main layers according to the oil production capacity, which are the target layers for fine characterization of the sweet spots.
[0099] Figure 9 To use this invention to quantitatively predict the free oil porosity of shale oil formations in an example work area, four formation categories were identified based on free oil porosity: Class I oil-bearing formations, Class II oil-bearing formations, Class III oil-bearing formations, and non-oil-bearing formations. The profile shows that this invention can accurately predict the free oil porosity of thin sweet spots of approximately 4 meters. For example, the predicted free oil porosity of the two thin sweet spots above the sweet spot in Well 1 and the two thin sweet spots below the sweet spot in Well 2 correspond well to the well sections shown above the wells. Based on free oil porosity, four formation categories were identified: Class I oil-bearing formations, Class II oil-bearing formations, Class III oil-bearing formations, and non-oil-bearing formations, thus achieving the classification and evaluation of shale oil sweet spots.
[0100] To further verify the accuracy of free oil porosity prediction in this invention, the trajectory of a horizontal well (well 3) was projected onto the free oil porosity prediction profile of this invention, and the accuracy was verified by combining nuclear magnetic resonance and conventional logging data of that well (e.g., Figure 10 As shown in the diagram, the NMR logging interpretation of Well 4 shows a high sweet spot encounter rate exceeding 90%, and its trajectory is generally located within the high free oil porosity interval predicted in the free oil porosity profile of this invention. This well encountered a layer near the target point but did not reach the target layer, corresponding to a region with predicted low free oil porosity. The prediction results are consistent with the actual drilling results, proving the high accuracy of the free oil porosity prediction in this invention. Figure 11a Planar distribution diagram of free oil porosity in the main layer of the No. 1 sweetener; Figure 11b Planar distribution diagram of free oil porosity in the main layer of the No. 2 sweetener; Figure 11c This is a planar distribution diagram of the free oil porosity in the main layer of the lower sweet spot No. 1. Figure 11d The diagram shows the planar distribution of free oil porosity in the No. 2 main oil-producing layer. It can be seen that the distribution patterns of free oil porosity in each main oil-producing section are different, which means that the locations of high-quality sweet spots are different, and the drilling locations are also different.
[0101] Furthermore, this invention also provides a quantitative seismic prediction system for free oil porosity, such as... Figure 12 As shown, the system includes:
[0102] The calculation module is used to calculate the porosity data of free oil in a single well shale oil layer using nuclear magnetic resonance logging, and to obtain high-quality near-, medium- and far-offset partially superimposed seismic data;
[0103] The sample-specific processing module is used to clarify the relationship between seismic elastic parameters and free oil porosity using shale oil rock physical interpretation charts, and to perform sample-specific processing on the high-resolution seismic inversion elastic data obtained from pre-stack statistical inversion to form machine learning prediction feature value data for free oil porosity.
[0104] The K-means clustering analysis module is used to perform K-means clustering analysis on sensitive seismic attribute data volumes and targeted processed seismic inversion elastic data volumes to establish a random forest free oil content prediction model.
[0105] A module is established to obtain the initial prediction data volume of free oil porosity based on the random forest free oil content prediction model, and then input it as the feature value into the random forest free oil porosity prediction model again to complete the second free oil porosity prediction, and establish the KRR free oil porosity seismic prediction model.
[0106] The free oil porosity final prediction data body module is used to input the free oil porosity machine learning prediction feature value data and the sensitive seismic attribute data body into the KRR free oil porosity seismic prediction model to obtain the free oil porosity final prediction data body.
[0107] The verification module is used to verify the final prediction data of free oil porosity and to quantitatively characterize the high-quality sweet spots of shale oil within the work area by combining the seismic inversion elastic parameter volume.
[0108] Specifically, the sample-targeted processing module includes a determining unit and a processing unit;
[0109] The determining unit is used to construct a petrophysical interpretation chart of shale oil reservoirs through rock physical analysis. Based on the chart, the longitudinal wave impedance corresponding to a high-quality sweet spot with an oil saturation >15% is determined to be greater than 9200 g / cm. 3 *m / s and less than 13000g / cm 3 *m / s, P-wave to S-wave velocity ratio less than 1.85;
[0110] The processing unit is used to process the high-resolution P-wave impedance and high-resolution P-wave / S-wave velocity ratio data obtained from pre-stack statistical inversion, remove the data corresponding to non-sweet spots in mudstone and tight strata outside the range, and then perform various mathematical transformations on the processed data to form a set of machine learning prediction feature value data for free oil porosity.
[0111] In summary, the results of practical application in the case study show that this invention can accurately predict the porosity of free oil in unconventional shale reservoirs, thereby enabling precise location of high-quality sweet spots in shale oil. This provides important technical support for the optimization of well location, well network deployment, and well trajectory design for exploration wells, appraisal wells, and development horizontal well groups.
[0112] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for quantitative seismic prediction of free oil porosity, the method comprising the following steps: 1) Utilize nuclear magnetic resonance logging to calculate free oil porosity data for shale oil formations in a single well, and obtain high-quality near-, mid-, and far-offset partially superimposed seismic data; 2) Use shale oil rock physical interpretation charts to clarify the relationship between seismic elastic parameters and free oil porosity, and perform targeted processing on the high-resolution seismic inversion elastic data obtained from pre-stack statistical inversion to form machine learning prediction feature value data for free oil porosity. 3) Perform K-means cluster analysis on sensitive seismic attribute data volumes and targeted processed seismic inversion elastic data volumes to establish a random forest free oil content prediction model; 4) Obtain the initial prediction data volume of free oil porosity based on the random forest free oil content prediction model, and then input it as the feature value into the random forest free oil porosity prediction model again to complete the second free oil porosity prediction and establish the KRR free oil porosity seismic prediction model. 5) Input the machine learning prediction feature value data for free oil porosity and the sensitive seismic attribute data body into the KRR free oil porosity seismic prediction model to obtain the final prediction data body for free oil porosity; 6) Verify the final prediction data of free oil porosity, and quantitatively characterize the high-quality sweet spots of shale oil within the work area by combining the seismic inversion elastic data.
2. The method according to claim 1, wherein, Step 2) specifically refers to: A petrophysical interpretation chart of shale oil reservoirs was constructed through rock physical analysis. Based on the chart, a P-wave impedance greater than 9200 g / cm² was identified for high-quality sweet spots with an oil saturation >15%. 3 m / s and less than 13000 g / cm 3 m / s, P-wave velocity ratio less than 1.85; the high-resolution P-wave impedance and high-resolution P-wave velocity ratio data obtained by pre-stack statistical inversion are processed to remove data corresponding to non-sweet spots in mudstone and tight strata outside the range, and then various mathematical transformations are performed on the processed data to form machine learning prediction feature value data for free oil porosity.
3. The method according to claim 1, wherein, In step 3), the K-means clustering analysis is performed using the K-means algorithm, which includes the following steps: ① Initialization: Randomly select K sample points as K initial cluster centers; ② Cluster the samples: Calculate the distance from each sample to each cluster center, and assign each sample to the cluster containing its nearest neighboring cluster center; the distance between cluster centers is calculated using Euclidean distance. The formula is as follows: ; in, x i This represents the x-coordinate of the i-th point in the Cartesian coordinate system. y i This represents the ordinate value of the i-th point in the rectangular coordinate system; ③ Calculate new cluster centers: Calculate the mean of all samples in each cluster as the new cluster centers; ④ Stop iterating if the cluster centers no longer change or the cutoff condition is reached; otherwise, repeat steps ② and ③.
4. The method according to claim 3, wherein, Random forests are classifiers based on decision trees. They use the ID3 algorithm to calculate the entropy of data feature samples to evaluate the branches of the decision trees. The formula for calculating the entropy of data feature samples is as follows: ; in, Entropy of data feature samples, Let n be the probability function of the data feature samples, and n be the number of branch layers. Each parent node is divided into multiple child nodes. Each node has a data feature sample entropy. The difference between the data feature sample entropy of the parent node and the weighted average of the data feature sample entropies of all its child nodes is calculated as the data gain optimization decision tree. ; ; in, For the number of nodes, The number of samples for a node. This represents the total number of samples in the corresponding parent node. I (x j ) Data feature sample entropy; and Let these represent the entropy of the parent node and the child node, respectively. Data Gain It is the difference between the weighted average of the entropy of the data feature samples of the upper-level node and the lower-level node.
5. The method according to claim 1, wherein, Step 1) includes: using nuclear magnetic resonance logging to calculate the free oil porosity data of a single-well shale oil layer, with a portion of the single wells marked as learning wells and the remaining single wells marked as verification wells. The selection criteria for learning wells are that they are representative wells in the work area, can reflect the geological characteristics of the target work area, and are evenly distributed within the work area; and obtaining high-quality near, medium, and far offset partial superimposed seismic data.
6. The method according to claim 2, wherein, Pre-stack statistical inversion is based on the Markov chain Monte Carlo algorithm, which converges to the population within the actual time bound. During the inversion process, the variogram, probability distribution function, well logging, seismic and lithofacies information are integrated. A corresponding probability distribution model is established based on seismic data analysis. The variogram and probability distribution function are obtained through geological understanding and well data analysis. Starting from the well points, the well points are made to conform to the original seismic information. The inversion software uses a stochastic simulation module to generate inter-well impedance. The wavelet extracted by deterministic inversion is convolved with the reflection coefficient converted from the inter-well impedance to generate a composite seismic record. Then, a new wavelet is extracted to generate a new composite record. This process is repeated iteratively until the matching rate between the composite seismic trace and the original seismic data reaches more than 85%, so that the correlation coefficient meets the requirements of pre-stack inversion. Finally, the inversion parameters are adjusted to complete the inversion calculation.
7. The method according to claim 6, wherein, The variation function is a distance function that describes the spatial distribution characteristics of a certain attribute as a function of distance, and is represented as a regionalized variable. exist and The semivariance of the increment at two points; ; in, Let be the experimental variation function. x Represents the absolute position of a point in space. h This represents the lag moment of the variation function.
8. The method according to claim 7, wherein, The three characteristic values of the variogram function—the sill value, the range, and the nugget constant—reflect the spatial variation characteristics of the reservoir parameters. Range is the spatial distance at which the variogram reaches a stable value, reflecting the average scale of the carrier of regionalized variables in the direction or the predicted extension scale of the sand body in the direction.
9. A quantitative seismic prediction system for free oil porosity, the system comprising: The calculation module is used to calculate the porosity data of free oil in a single well shale oil layer using nuclear magnetic resonance logging, and to obtain high-quality near-, medium- and far-offset partially superimposed seismic data; The sample-specific processing module is used to clarify the relationship between seismic elastic parameters and free oil porosity using shale oil rock physical interpretation charts, and to perform sample-specific processing on the high-resolution seismic inversion elastic data obtained from pre-stack statistical inversion to form machine learning prediction feature value data for free oil porosity. The K-means clustering analysis module is used to perform K-means clustering analysis on sensitive seismic attribute data volumes and targeted processed seismic inversion elastic data volumes to establish a random forest free oil content prediction model. A module is established to obtain the initial prediction data volume of free oil porosity based on the random forest free oil content prediction model, and then input it as the feature value into the random forest free oil porosity prediction model again to complete the second free oil porosity prediction, and establish the KRR free oil porosity seismic prediction model. The free oil porosity final prediction data body module is used to input the free oil porosity machine learning prediction feature value data and the sensitive seismic attribute data body into the KRR free oil porosity seismic prediction model to obtain the free oil porosity final prediction data body. The verification module is used to verify the final prediction data of free oil porosity and to quantitatively characterize the high-quality sweet spots of shale oil within the work area by combining the seismic inversion elastic data.
10. The system according to claim 9, wherein, The sample-specific processing module includes a determination unit and a processing unit; The determining unit is used to construct a petrophysical interpretation chart of shale oil reservoirs through rock physical analysis. Based on the chart, the longitudinal wave impedance corresponding to a high-quality sweet spot with an oil saturation >15% is determined to be greater than 9200 g / cm. 3 m / s and less than 13000 g / cm 3 m / s, P-wave velocity ratio less than 1.85; The processing unit is used to process the high-resolution P-wave impedance and high-resolution P-wave / S-wave velocity ratio data obtained from pre-stack statistical inversion, remove the data corresponding to non-sweet spots in mudstone and tight strata outside the range, and then perform various mathematical transformations on the processed data to form machine learning prediction feature value data for free oil porosity.
Citation Information
Patent Citations
Method for quantificationally predicting sandstone reservoir fluid saturation by combining well and seism
CN101887132A
Compact glutenite gas reservoir quantificational prediction method
CN105044770A