Method for predicting formation pore pressure and related equipment

By correcting outliers and calculating differential pressure on the original well logging data, and combining the random forest model optimized by the Grey Wolf optimization algorithm, the problem of low formation pore pressure prediction accuracy in shale gas exploration has been solved, achieving more accurate pore pressure prediction and improving exploration efficiency.

CN122190742APending Publication Date: 2026-06-12CHINA UNIV OF GEOSCIENCES (BEIJING)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA UNIV OF GEOSCIENCES (BEIJING)
Filing Date
2026-01-30
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of formation pore pressure prediction in shale gas exploration is low, leading to inaccurate drilling location selection and affecting the exploration process and results.

Method used

By acquiring raw logging data, correcting outliers, calculating the pressure difference between overlying strata pressure, pore pressure, and drilling fluid column pressure, and using the Grey Wolf optimization algorithm to determine the hyperparameters of the random forest model, formation pore pressure is predicted.

Benefits of technology

It improves the accuracy and stability of formation pore pressure prediction, and enhances the reliability of drilling design and the efficiency of shale gas exploration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122190742A_ABST
    Figure CN122190742A_ABST
Patent Text Reader

Abstract

The application provides a formation pore pressure prediction method and related equipment, and relates to the technical field of shale gas exploration. The method comprises the following steps: obtaining original logging data, and correcting the original logging data to obtain corrected logging data. The overburden pressure is calculated according to the corrected logging data. Based on the overburden pressure and the corrected logging data, the pressure difference between the pore pressure and the drilling fluid column pressure is calculated. The pressure difference can provide input features with clear physical direction for the random forest model, effectively constrain the original logging data from the rock physics, and help improve the generalization ability and interpretability of the random forest model. Based on the corrected logging data and the pressure difference, the formation pore pressure is predicted by the random forest model. The hyperparameters in the random forest model are determined according to the grey wolf optimization algorithm. The random forest model optimized by the grey wolf optimization algorithm improves the accuracy of the formation pore pressure prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of shale gas exploration technology, and in particular to a method and related equipment for predicting formation pore pressure. Background Technology

[0002] Predicting formation pore pressure is a crucial part of shale gas exploration. Higher accuracy in predicting abnormal formation pore pressure allows for drilling wells closer to gas-rich locations. However, shale pore structures are complex and highly heterogeneous. Traditional formation pressure prediction methods rely on regional empirical parameters; incorrect parameter selection leads to low prediction accuracy, significantly impacting the progress and results of shale gas exploration. Summary of the Invention

[0003] In view of this, the purpose of this application is to propose a method and related equipment for predicting formation pore pressure, so as to solve the problem of low accuracy in predicting formation pore pressure.

[0004] To achieve the above objectives, the first aspect of this application provides a method for predicting formation pore pressure, comprising:

[0005] Obtain raw logging data and perform outlier correction on the raw logging data to obtain corrected logging data; Calculate the pressure of the overlying strata based on the corrected logging data; Based on the overlying strata pressure and the corrected logging data, the pressure difference between the pore pressure and the drilling fluid column pressure is calculated. Based on the corrected logging data and the pressure difference, the formation pore pressure is predicted using a random forest model; wherein the hyperparameters in the random forest model are determined according to the Grey Wolf optimization algorithm.

[0006] Optionally, the outlier correction of the original well logging data includes: In response to determining that the data to be detected in the original well logging data is non-boundary data and that the data to be detected is abnormal data, the data to be detected is replaced by the local mean replacement method; In response to determining that the data to be detected in the original well logging data is boundary data and that the data to be detected is abnormal data, the data to be detected is replaced by the boundary point neighborhood mean replacement method.

[0007] Optionally, in response to determining that the data to be detected in the original well logging data is non-boundary data and that the data to be detected is abnormal data, replacing the data to be detected with a local mean replacement method includes: A sliding window is constructed based on the data to be detected, and the local mean and local standard deviation are calculated based on the data within the sliding window. The local mean and the local standard deviation are used to determine whether the data to be detected is abnormal. If the data to be detected is determined to be abnormal, the local mean is used to replace the data to be detected.

[0008] Optionally, in response to determining that the data to be detected in the original well logging data is boundary data and that the data to be detected is abnormal data, the method of replacing the data to be detected with the boundary point neighborhood mean replacement method includes: Based on the original well logging data, calculate the global mean and global standard deviation; Based on the global mean and the global standard deviation, determine whether the data to be detected is abnormal. If the data to be detected is determined to be abnormal and is located at the beginning boundary of the sequence, construct a window using the neighboring data after the data to be detected. If the data to be detected is determined to be abnormal and is located at the end boundary of the sequence, construct a window using the neighboring data in front of the data to be detected. Calculate the mean of the data within the window and replace the data to be detected with the mean of the data.

[0009] Optionally, after correcting outliers in the original logging data to obtain corrected logging data, the method further includes: Calculate the new global mean and the new global standard deviation based on the revised logging data; Determine whether the corrected logging data is anomalous based on the new global mean and the new global standard deviation. If it is determined to be anomalous, replace the corrected logging data with the new global mean.

[0010] Optionally, the step of calculating the pressure difference between pore pressure and drilling fluid column pressure based on the overlying strata pressure and the corrected logging data includes: Calculate the P-wave velocity and S-wave velocity based on the corrected logging data; The bulk modulus of saturated rock is calculated based on the longitudinal wave velocity, the transverse wave velocity, and the bulk density. Based on the saturated rock bulk modulus and the Geismann fluid substitution equation, the dry rock bulk modulus is determined; The hydrostatic pressure difference is determined based on the pressure of the overlying strata and the hydrostatic pressure. The pressure difference is determined based on the hydrostatic pressure difference and the dry rock bulk modulus.

[0011] Optionally, calculating the bulk modulus of saturated rock based on the P-wave velocity, the S-wave velocity, and the bulk density includes: The shear modulus is calculated based on the bulk density and the transverse wave velocity. The bulk modulus of the saturated rock is calculated based on the bulk density, the longitudinal wave velocity, and the shear module.

[0012] Optionally, determining the pressure difference based on the hydrostatic pressure difference and the dry rock bulk modulus includes: The normal compaction modulus is determined based on the hydrostatic pressure difference. The pressure difference is calculated based on the normal compaction modulus and the dry rock bulk modulus.

[0013] Based on the same inventive concept, a second aspect of this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.

[0014] Based on the same inventive concept, a third aspect of this application also provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method described above.

[0015] As described above, the formation pore pressure prediction method and related equipment provided in this application include the following steps: acquiring raw logging data and correcting outliers in the raw logging data to obtain corrected logging data. By correcting outliers, the raw test data is smoothed. The overlying strata pressure is calculated based on the corrected logging data. The pressure difference between pore pressure and drilling fluid column pressure is calculated based on the overlying strata pressure and the corrected logging data. The pressure difference is essentially a composite index coupling rock mechanical properties, burial conditions, and fluid pressure. It provides input features with clear physical orientation for the random forest model, effectively constraining the raw logging data and improving the generalization ability and interpretability of the random forest model. Based on the corrected logging data and the pressure difference, the formation pore pressure is predicted using the random forest model; the hyperparameters in the random forest model are determined using the Grey Wolf optimization algorithm. The random forest model optimized by the Grey Wolf optimization algorithm improves the accuracy of formation pore pressure prediction. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a schematic flowchart of the formation pore pressure prediction method according to an embodiment of this application; Figure 2 This is a comprehensive logging diagram of an embodiment of this application; Figure 3 This is a schematic diagram comparing the predicted pore pressure curves of embodiments of this application; Figure 4 These are cross-plots and error histograms of different measured and predicted data from embodiments of this application; Figure 5 This is a schematic diagram comparing the predicted pore pressure curves of different data processing methods in embodiments of this application. Figure 6 This is a cross-plot of predicted pore pressure and measured pore pressure for different data processing methods in embodiments of this application. Figure 7 This is a schematic diagram comparing the predicted data results of an embodiment of this application; Figure 8 This is a schematic diagram of the formation pore pressure prediction device according to an embodiment of this application; Figure 9 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.

[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0020] Shale gas exploration and development is a global hot topic. Various exploration results indicate that pore pressure within shale formations is crucial to shale gas production capacity, and the accuracy of pore pressure prediction significantly impacts the final total shale gas yield, serving as a key indicator of shale gas preservation. Accurate pore pressure prediction is also essential for ensuring drilling safety, optimizing drilling design, and guiding oil and gas exploration and development. However, shale often exhibits strong heterogeneity; its rock structure, mineral composition, organic matter content, occurrence mode, and pore structure vary significantly across different regions. This leads to substantial differences in the elastic properties of shale, making simple generalizations difficult.

[0021] At present, pore pressure prediction methods based on empirical formulas have been developed to a relatively mature stage. Commonly used methods include the equivalent depth method, effective stress method, Eaton index method, Bowers method, and Fillippone formula method. However, these traditional methods have many limitations in the calculation process, such as strong dependence on P-wave and S-wave velocities, a large number of required empirical parameters, significant influence from human factors, and the need to calculate different empirical parameters for different environments. As a result, the accuracy and reliability of the prediction results are difficult to guarantee.

[0022] In view of this, this application proposes a method for predicting formation pore pressure. First, the original well logging data is preprocessed to remove outliers. Then, the original well logging data is processed using a velocity-pressure rock physics model to obtain new data—pressure difference (Pd), which is used as a feature parameter. Finally, the hyperparameters of the random forest are optimized using the Grey Wolf algorithm, and the formation pore pressure is predicted by the random forest model.

[0023] The embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0024] This application proposes a method for predicting formation pore pressure, referring to... Figure 1 This includes the following steps: Step 102: Obtain the original logging data and correct outliers in the original logging data to obtain corrected logging data.

[0025] Specifically, the raw logging data are test curves plotted based on the vertical depth of the drilled well. The raw logging data includes bulk density logging data, natural gamma ray logging data, porosity logging data, and sonic transit time logging data. Bulk density refers to the total density of the rock formation. Natural gamma ray refers to the intensity of gamma rays from naturally occurring radioactive elements in the formation. Porosity refers to the proportion of pore volume to the total volume of the rock. Sonic transit time is the reciprocal of the time required for a P-wave to travel through a unit thickness of formation, measured in microseconds per meter. To ensure the accuracy of subsequent formation pore pressure prediction, outlier correction is performed on the raw logging data to obtain more accurate corrected logging data.

[0026] Furthermore, the outlier correction of the original well logging data includes: In response to determining that the data to be detected in the original well logging data is non-boundary data and that the data to be detected is abnormal data, the data to be detected is replaced by the local mean replacement method; In response to determining that the data to be detected in the original well logging data is boundary data and that the data to be detected is abnormal data, the data to be detected is replaced by the boundary point neighborhood mean replacement method.

[0027] Specifically, the raw logging data includes boundary data and non-boundary data. Since non-boundary data cannot be used to construct a complete sliding window, different methods are needed to correct outliers in boundary data and non-boundary data.

[0028] Furthermore, in response to determining that the data to be detected in the original well logging data is non-boundary data, and confirming that the data to be detected is abnormal data, the method of replacing the data to be detected with a local mean replacement method includes: A sliding window is constructed based on the data to be detected, and a local mean and a local standard deviation are calculated based on the data within the sliding window. The local mean and the local standard deviation are used to determine whether the data to be detected is abnormal. If the data to be detected is determined to be abnormal, the local mean is used to replace the data to be detected.

[0029] Specifically, when the data to be detected is non-boundary data, a sliding window with a size of 3 is constructed, centered on the data to be detected. For example, the data to be detected is denoted as data point i. Based on data point i, the preceding data point i-1 and the following data point i+1 are selected to construct the sliding window. The local mean and local standard deviation are calculated based on the data within the sliding window. The local mean is calculated by taking the average of all data within the sliding window. The local standard deviation is calculated by taking the standard deviation of all data within the sliding window.

[0030] Next, the local mean and local standard deviation are used to determine whether the data to be detected is outlier. Specifically, the Z-score of the data to be detected is calculated. The Z-score measures the deviation of a data point from the mean of its dataset and is standardized in units of the dataset's standard deviation. If the Z-score of the data to be detected is greater than a preset threshold, the data is determined to be outlier. For example, the preset threshold is 2.5 times the local standard deviation. The local mean is used to replace the data to be detected to correct outlier non-boundary data. Through the outlier correction of non-boundary data in this embodiment, the random noise in the original logging data can be smoothed, restoring the true trend of the original logging data curve. By combining the local mean and local standard deviation to correct the data to be detected, the true vertical variation characteristics of the formation can be preserved to the maximum extent, providing clean data for subsequent data analysis and effectively improving the accuracy and stability of subsequent formation pore pressure prediction.

[0031] Furthermore, in response to determining that the data to be detected in the original well logging data is boundary data, and confirming that the data to be detected is abnormal data, the data to be detected is replaced using the boundary point neighborhood mean replacement method, including: Based on the original logging data, calculate the global mean and global standard deviation; based on the global mean and global standard deviation, determine whether the data to be tested is abnormal data; if the data to be tested is determined to be abnormal data and is located at the beginning boundary of the sequence, construct a window using the neighboring data after the data to be tested; if the data to be tested is determined to be abnormal data and is located at the end boundary of the sequence, construct a window using the neighboring data before the data to be tested; calculate the mean of the data within the window, and replace the data to be tested with the mean of the data.

[0032] Specifically, when the data to be detected is boundary data, the global mean and global standard deviation are calculated based on the original well logging data. The global mean is calculated by taking the average of all data in the original well logging data. The global standard deviation is calculated by taking the standard deviation of all data in the original well logging data. Then, the global mean and global standard deviation are used to determine whether the data to be detected is abnormal. If the value of the data to be detected exceeds a preset threshold, the data to be detected is determined to be abnormal. For example, the preset threshold can be 2.5 times the global standard deviation. When correcting the data to be detected, an asymmetric neighborhood strategy is used to construct a window based on its boundary position in the data sequence: if the data to be detected is located at the beginning boundary of the sequence, a forward neighborhood window is constructed based on the consecutive data points adjacent to it; if the data to be detected is located at the end boundary of the sequence, a backward neighborhood window is constructed based on the consecutive data points adjacent to it; this window includes all data points adjacent to the data point of the data to be detected. The mean of all data in the one-way neighborhood window is calculated, and the data to be detected is replaced with this mean to correct abnormal boundary data. This embodiment provides a targeted correction scheme for data that the sliding window cannot cover, achieving seamless processing of all original well logging data. Subsequent predictions can avoid introducing boundary errors into the random forest model, providing a reliable data foundation for prediction.

[0033] To further identify outlier data, after correcting for outliers in the original logging data to obtain corrected logging data, the method further includes: Based on the corrected logging data, calculate the new global mean and the new global standard deviation; determine whether the corrected logging data is outlier based on the new global mean and the new global standard deviation; if it is determined to be outlier, replace the corrected logging data with the new global mean.

[0034] Specifically, after correcting the anomalies in the data to be detected in the aforementioned embodiments, a global check is performed on all corrected logging data to prevent errors and omissions. Based on the corrected logging data, a new global mean and a new global standard deviation are calculated. The new global mean is calculated by averaging all data points in the corrected logging data. The new global standard deviation is calculated by calculating the standard deviation of all data points in the corrected logging data. If any value in the corrected logging data exceeds a preset threshold, the corrected logging data is determined to be abnormal data. The preset threshold can be three times the new global standard deviation. The abnormal corrected logging data is replaced with the new global mean. Replacing with the new global mean is necessary because local data may experience correction errors due to the consecutive occurrence of multiple abnormal data points; this serves as a global verification and supplement to the local correction results. By replacing abnormal data in the corrected logging data, outliers that may have been missed after replacing boundary and non-boundary data can be captured. When there are multiple consecutive data points with anomalies in the original logging data, determining whether the corrected logging data is abnormal can identify these overall abnormal logging data, ensuring the overall rationality of the logging data.

[0035] Step 104: Calculate the overlying strata pressure based on the corrected logging data.

[0036] Specifically, the overlying strata pressure is calculated based on the density logging data from the corrected logging data. Overlying strata pressure is a crucial component of the effective stress law. It represents the vertical stress generated by the total weight of all rock matrix and pore fluids above a given point underground. Its calculation formula is achieved by vertically integrating the bulk density logging curve, as shown in the following equation: (1), in, Indicates the pressure of the overlying strata. Indicates the maximum drilling depth. Represents gravitational acceleration. Bulk density, unit: kg / m³ -3 , Indicates the initial depth.

[0037] Step 106: Based on the overlying strata pressure and the corrected logging data, calculate the pressure difference between the pore pressure and the drilling fluid column pressure.

[0038] Specifically, based on the determined overlying strata pressure, a velocity-pressure rock physics model is introduced to calculate the key related characteristic of formation pore pressure, namely, the pressure difference. In practice, the calculation of the pressure difference between pore pressure and drilling fluid column pressure based on the overlying strata pressure and the corrected logging data includes: Calculating the P-wave velocity and S-wave velocity based on the corrected logging data includes: calculating the sonic transit time logging data from the corrected logging data. Calculate P-wave velocity , The unit is m / s. Units are The specific calculation formula is as follows: (2), Determine the transverse wave velocity based on the longitudinal wave velocity. The specific calculation formula is as follows: (3).

[0039] The saturated rock bulk modulus is calculated based on the longitudinal wave velocity, the transverse wave velocity, and the bulk density, including: The shear modulus is calculated based on the bulk density and the transverse wave velocity. The calculation formula is as follows: (4); The bulk modulus of the saturated rock is calculated based on the bulk density, the longitudinal wave velocity, and the shear module. The calculation formula is as follows: (5).

[0040] Based on the saturated rock bulk modulus and the Geismann fluid substitution equation, the dry rock bulk modulus is determined. The specific calculation formula is as follows: (6), (7), Among them, by using equations (6) and (7), the gassmann fluid substitution equation is adopted to inversely calculate the bulk modulus of the dry rock skeleton under the premise of given matrix modulus and different fluid modulus. Indicates the modulus of the mineral matrix. Indicates the original fluid modulus. Indicates the fluid modulus after replacement. Indicates porosity. Indicates intermediate quantity.

[0041] Based on the overlying strata pressure and hydrostatic pressure, the hydrostatic pressure differential is determined. The core of the velocity-pressure model lies in establishing an empirical-theoretical relationship between the dry rock bulk modulus and the formation pressure state. This is achieved by calculating the overlying strata pressure. The difference between the static pressure and the hydrostatic pressure is called the hydrostatic pressure differential. Hydrostatic pressure, also known as water column pressure or hydrostatic pressure, refers to the pressure generated by the weight of a continuous fluid column above a certain depth. Under normal circumstances, formation pore pressure equals hydrostatic pressure, and formation pore pressure is obtained from formation pore pressure logging data.

[0042] Determining the pressure difference based on the hydrostatic pressure difference and the dry rock bulk modulus includes: The normal compaction modulus is determined based on the hydrostatic pressure difference. The calculation formula is as follows: (8), The pressure difference is calculated based on the normal compaction modulus and the dry rock bulk modulus. The calculation formula is as follows: Where n represents the lithology index. Pressure differential It is a composite index that couples rock mechanical properties, burial conditions and fluid pressure. It can directly reflect the effect of formation pore pressure anomalies on the elastic properties of rocks, thus providing a physically oriented input feature for machine learning models (subsequent random forest models).

[0043] Step 108: Based on the corrected logging data and the pressure difference, the formation pore pressure is predicted using a random forest model; wherein the hyperparameters in the random forest model are determined according to the Grey Wolf optimization algorithm.

[0044] Specifically, after obtaining the corrected logging data and differential pressure through the aforementioned steps, these data are input into the random forest model. The data input into the random forest model includes corrected logging data such as bulk density, natural gamma, porosity, sonic transit time, and depth data. The hyperparameters in the random forest model are determined based on the Grey Wolf optimization algorithm. The maximum number of iterations of the Grey Wolf optimization algorithm is set to 30, and the number of individuals in the population is set to 5. The two hyperparameters of the random forest model to be optimized are the number of decision trees and the maximum number of branch splits. The minimum search boundary is set to [80, 3], and the maximum search boundary is set to [600, 35]. After iteration through the Grey Wolf optimization algorithm, the number of optimized decision trees is 80, and the maximum number of splits is 32. These two hyperparameters are then input into the random forest model. The optimized random forest model predicts the formation pore pressure. Formation pore pressure refers to the pressure exerted by fluids (such as gas or water) in the formation pores.

[0045] When geological conditions are complex, increasing the number of decision trees can improve the stability of the random forest model, reduce variance, and better capture complex nonlinear relationships. When geological conditions are simple, reducing the number of decision trees can increase the speed of the random forest model, but too many trees may lead to overfitting. The maximum number of splits limits the overall complexity of the decision trees; when geological conditions are complex, the number of decision nodes should be increased, and when geological conditions are simple, the number of decision nodes should be decreased.

[0046] Based on steps 102 to 108 above, the formation pore pressure prediction method provided in this embodiment includes: acquiring original logging data and correcting outliers in the original logging data to obtain corrected logging data. By correcting outliers, smoothing of the original test data is achieved. Overlying strata pressure is calculated based on the corrected logging data. The pressure difference between pore pressure and drilling fluid column pressure is calculated based on the overlying strata pressure and the corrected logging data. The pressure difference is actually a composite index coupling rock mechanical properties, burial conditions, and fluid pressure. It can provide input features with clear physical orientation for the random forest model, effectively constraining the original logging data and improving the generalization ability and interpretability of the random forest model. Based on the corrected logging data and the pressure difference, formation pore pressure is predicted using the random forest model; wherein the hyperparameters in the random forest model are determined according to the Grey Wolf optimization algorithm. The random forest model optimized by the Grey Wolf optimization algorithm improves the accuracy of formation pore pressure prediction.

[0047] In some embodiments, the training method for the random forest model includes: Obtain the original logging samples and perform outlier correction to obtain corrected logging samples. The outlier correction method is the same as in the previous embodiment and will not be repeated here. Calculate the overlying strata pressure based on the corrected logging samples. Based on the overlying strata pressure and the corrected logging samples, calculate the pressure difference between the pore pressure and the drilling fluid column pressure. Obtain the actual formation pore pressure and construct training samples based on the corrected logging samples, the pressure difference, and the actual formation pore pressure. Iteratively train the initial random forest model using the training samples until the training cutoff condition is met to obtain the trained random forest model.

[0048] To verify the effectiveness of the formation pore pressure prediction method provided in this application, a series of experiments were conducted. The specific experimental methods are as follows: Figure 2A comprehensive logging chart is shown. The study area selected for the experiment is located in Pingqiao District, Fuling, Chongqing, China. For the shale layer in the well logging, curves showing the variation of logging parameters with formation depth (Depth) between 3800m and 5000m that may affect formation pore pressure are included. These curves include volumetric density logging curves (RHOB), natural gamma logging curves (GR), sonic transit time logging curves (AC), porosity logging curves (POR), and measured pore pressure curves. These are then plotted into a comprehensive logging chart. Figure 2 This indicates that there is no abnormal pressure in this well section.

[0049] First, the prediction results of the algorithms were compared. The three algorithms selected were: Random Forest algorithm, Bayes-RF optimized random forest algorithm, and Grey Wolf-RF optimized random forest algorithm (GWO-RF), with GWO-RF being the method used in this application. The three algorithms are based on... Figure 2 Pore ​​pressure was predicted using known well logging data and compared with measured pore pressure. The comparison results are as follows: Figure 3 A schematic diagram comparing the predicted pore pressure curves is shown. The red curve represents the measured pore pressure curve, and the blue curve represents the predicted pore pressure curve. Figure 3 It is evident that the Bayes-RF predicted pore pressure curve deviates most from the measured pore pressure curve, followed by the RF predicted pore pressure curve, while the GWO-RF predicted pore pressure curve is closest to the measured pore pressure curve. Particularly in the 4650-4800m range, Bayes-RF exhibits significant deviation, with its predicted pore pressure curve fluctuating continuously. At 4800m, it shows a prediction result opposite to the measured pore pressure curve. While the Random Forest prediction result is closer to the measured data than Bayes-RF, it also shows a large deviation at 4800m. Meanwhile, the GWO-RF prediction result at 4800m is very close, almost overlapping with the measured data, yielding a better prediction result.

[0050] To more clearly compare the errors between the three algorithms, Figure 4 Cross-plots and error histograms of different measured and predicted data are shown. Figure 4 In diagram 'a', the x-axis represents the measured pore pressure, and the y-axis represents the predicted pore pressure, both in MPa. The red scatter points represent the predictions from a random forest, the green scatter points from a Bayes-RF prediction, and the blue scatter points from a GWO-RF prediction. The closer the scatter points are to the central black diagonal line, the closer they are to the measured data, and the better the prediction performance. Figure 4As can be clearly seen in diagram a, while the random forest's scatter points are relatively concentrated, most deviate from the measured data, overlapping only in a few locations. Bayes-RF's scatter points, although overlapping more with the measured data, are generally more scattered and deviate significantly. GWO-RF's scatter points are closest to the measured data, exhibiting not only high overlap but also a compact structure with minimal dispersion. This comparison demonstrates that GWO-RF has the best prediction performance, while Bayes-RF has the worst.

[0051] exist Figure 4 In b, the x-axis represents the prediction error and the y-axis represents the frequency. This reflects that the closer the error between the predicted pore pressure and the measured pore pressure is to 0 and the higher the frequency, the smaller the error and the higher the prediction accuracy. Figure 4 It is evident from b that GWO-RF has the smallest error between its predicted and measured data because its error frequency is highest near 0, and most error results are concentrated in the (-0.5, 0.5) interval, indicating a relatively small overall error. While the distribution of Random Forest's prediction results is also relatively compact and close to 0, its frequency at 0 is not as high as GWO-RF, and it is slightly more dispersed in the (-1, 1) interval. Bayes-RF's error is more dispersed in the (-1, 1) interval, and is widely distributed and has a high frequency on both sides of 0. Its error is also distributed in the (-3, 3) interval, resulting in the worst prediction performance.

[0052] Table 1 presents a comparison of the errors of the three algorithms. The errors between the predicted pore pressure and the measured pore pressure for each of the three algorithms were calculated. The closer the values ​​of MAE (Mean Absolute Error) and RMSE (Root Mean Square Error) are to 0, the smaller the error. 2 The closer the coefficient of determination (COD) is to 1, the closer the predicted value is to the observed value. The MAE value of GWO-RF is 0.1091, and the RMSE value is 0.1918, both the lowest among the three methods. R0 2 It is 0.9956, the closest value to 1 among the three. The MAE error of GWO-RF is reduced by 49.885% and 65.188% compared to RF and Bayes-RF, respectively; the RMSE error is reduced by 40.379% and 60.108% compared to RF and Bayes-RF, respectively; R... 2 Compared to RF and Bayes-RF, the performance was improved by 0.861% and 2.806% respectively. It can be clearly seen that GWO-RF has the smallest error and is closest to the measured curve.

[0053] Table 1. Error Comparison of Three Algorithms

[0054] After determining GWO-RF as the optimal algorithm, this application also compared the data processing procedures before inputting the raw well logging data into the random forest model optimized by the Grey Wolf optimization algorithm. Three data processing methods are provided: the first is to directly input the raw well logging data into the GWO-RF algorithm model; the second is to perform outlier correction on the raw well logging data before inputting it into the GWO-RF algorithm model; and the third is to perform outlier correction on the raw well logging data and calculate the pressure difference, then input the corrected well logging data and pressure difference into the GWO-RF algorithm model.

[0055] Figure 5 A comparative diagram of predicted pore pressure curves using different data processing methods is provided. Curve A corresponds to the first data processing method, curve B to the second, and curve C to the third. Figure 5 As shown, in the curves corresponding to the first data processing method, the blue curve represents the predicted pore pressure, and the red curve represents the measured pore pressure. Overall, the predicted pore pressures of these three data processing methods are quite close to the measured pore pressures. However, some differences can be observed at 4540m. The predicted pore pressure curve of the first data processing method shows significant fluctuations, while the predicted pore pressure curves of the second and third data processing methods do not exhibit large fluctuations. Similarly, there are differences at approximately 4520m. The predicted pore pressure curves using the first and second data processing methods show large fluctuations at this point, while the predicted pore pressure curve using the third data processing method shows smaller fluctuations than the former two.

[0056] To better compare predicted pore pressure with measured pore pressure Figure 6 The diagram shows a cross-plot of predicted and measured pore pressures for different data processing methods. The x-axis represents the measured pore pressure, and the y-axis represents the predicted pore pressure, both in MPa. The red dots represent the predictions for the first data processing method, the green dots for the second, and the blue dots for the third. The closer the dots are to the central black line, the closer they are to the measured data, and the better the prediction. Figure 6 The scatter plots within the pore pressure range of (14.5, 15) are magnified, revealing noticeable differences within this range. The red scatter plots are more dispersed and deviate from the measured data. The green scatter plots are more concentrated than the red ones, especially around 15.3 MPa. The blue scatter plots are more concentrated and deviate less.

[0057] Table 2 presents a comparison of the errors of the three data processing methods. The errors between the predicted pore pressure and the measured pore pressure for each of the three methods were calculated. For the third method, the MAE value is 0.1045 and the RMSE value is 0.1689, both the smallest among the three methods. This indicates that the error of the third method is the smallest. Meanwhile, R... 2 The value of 0.9966 is the closest to 1 among the three data processing methods, indicating that the curve fit is closest to the measured value. Compared with the first and second data processing methods, the third data processing method reduced the MAE error by 4.216% and 3.330%, respectively, and the RMSE error by 11.940% and 7.096%, respectively. 2 They increased by 0.1% and 0.060% respectively.

[0058] Table 2. Error Comparison of Three Data Processing Methods

[0059] To further validate the formation pore pressure prediction method of this application, this application also tested the method based on new well logging data. The new well logging data comes from "ZHAO Jun, et al. Prediction of shale formation pore pressure based on Zebra Optimization Algorithm-optimized support vector regression". The new well logging data includes data on density, natural gamma, sonic transit time, porosity, compensated neutrons, resistivity, etc. The prediction results of the ZOA-SVR model in the paper are used as the measured pore pressure. Two validation methods are used in the specific validation process: The first validation method uses the density, natural gamma, sonic transit time, and porosity from the new well logging data as features input to the GWO-RF algorithm model. The second validation method corrects outliers for the density, natural gamma, sonic transit time, and porosity from the new well logging data, calculates the pressure difference, and then uses the corrected well logging data and the pressure difference together as features input to the GWO-RF algorithm model. Figure 7 A diagram showing the comparison of predicted data results is provided.

[0060] Figure 7The blue curve represents the prediction results of the ZOA-SVR model. The red curve represents the prediction results of the GWO-RF algorithm model after training the model with the ZOA-SVR model prediction results as actual pore pressure, and then using the trained GWO-RF algorithm model with the first verification method. The yellow curve represents the prediction results of the GWO-RF algorithm model after correcting for outliers and calculating differential pressure from new logging data, using the ZOA-SVR model prediction results as actual pore pressure, and then using the trained GWO-RF algorithm model with the second verification method. It can be clearly seen that the trend of the yellow curve is basically consistent with that of the blue curve, especially at 1517-1518m, where the red curve shows an opposite trend to the blue curve, while the yellow curve remains consistent. This indicates that the second method (i.e., the formation pore pressure prediction method of this application) is effective.

[0061] Table 3 provides the following information: Figure 7 The error comparison table for the two verification methods clearly shows that the second verification method has a smaller error and is closer to the original curve. Compared with the first verification method, the second verification method reduces the MAE error by 36.385%, the RMSE error by 32.857%, and the R... 2 It increased by 12.483%.

[0062] Table 3. Error Comparison of Different Validation Methods

[0063] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.

[0064] It should be noted that some embodiments of this application have been described above. In some cases, the actions or steps described in the above embodiments can be performed in a different order than that shown in the above embodiments and the desired result can still be achieved. In addition, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0065] Based on the same inventive concept, and corresponding to any of the above embodiments, this application also provides a formation pore pressure prediction device.

[0066] refer to Figure 8The formation pore pressure prediction device includes: The correction module 202 is configured to acquire the original logging data and correct outliers in the original logging data to obtain corrected logging data. The first calculation module 204 is configured to calculate the overlying strata pressure based on the corrected logging data; The second calculation module 206 is configured to calculate the pressure difference between pore pressure and drilling fluid column pressure based on the overlying rock pressure and the corrected logging data. The prediction module 208 is configured to predict the formation pore pressure based on the corrected logging data and the pressure difference using a random forest model; wherein the hyperparameters in the random forest model are determined according to the Grey Wolf optimization algorithm.

[0067] In some embodiments, the correction module 202 is configured to replace the data to be detected with a local mean replacement method in response to determining that the data to be detected in the original logging data is non-boundary data and that the data to be detected is abnormal data; and to replace the data to be detected with a boundary point neighborhood mean replacement method in response to determining that the data to be detected in the original logging data is boundary data and that the data to be detected is abnormal data.

[0068] In some embodiments, the correction module 202 is configured to construct a sliding window based on the data to be detected, calculate a local mean and a local standard deviation based on the data within the sliding window, determine whether the data to be detected is abnormal based on the local mean and the local standard deviation, and if the data to be detected is determined to be abnormal, replace the data to be detected with the local mean.

[0069] In some embodiments, the correction module 202 is configured to calculate a global mean and a global standard deviation based on the original logging data; determine whether the data to be detected is abnormal based on the global mean and the global standard deviation; if the data to be detected is determined to be abnormal and is located at the beginning boundary of the sequence, construct a window using the neighboring data after the data to be detected; if the data to be detected is determined to be abnormal and is located at the end boundary of the sequence, construct a window using the neighboring data before the data to be detected, calculate the mean of the data within the window, and replace the data to be detected with the mean of the data.

[0070] In some embodiments, the correction module 202 is configured to calculate a new global mean and a new global standard deviation based on the corrected logging data; determine whether the corrected logging data is abnormal based on the new global mean and the new global standard deviation; and if it is determined to be abnormal, replace the corrected logging data with the new global mean.

[0071] In some embodiments, the second calculation module 206 is configured to calculate the P-wave velocity and S-wave velocity based on the corrected logging data; calculate the saturated rock bulk modulus based on the P-wave velocity, the S-wave velocity, and the bulk density; determine the dry rock bulk modulus based on the saturated rock bulk modulus and the Geismann fluid substitution equation; determine the hydrostatic pressure difference based on the overlying strata pressure and the hydrostatic pressure; and determine the pressure difference based on the hydrostatic pressure difference and the dry rock bulk modulus.

[0072] In some embodiments, the second calculation module 206 is configured to calculate the shear modulus based on the bulk density and the transverse wave velocity; and to calculate the saturated rock bulk modulus based on the bulk density, the longitudinal wave velocity, and the shear module.

[0073] In some embodiments, the second calculation module 206 is configured to calculate and determine the normal compaction modulus based on the hydrostatic pressure difference; and to calculate and determine the pressure difference based on the normal compaction modulus and the dry rock bulk modulus.

[0074] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.

[0075] The apparatus described above is used to implement the corresponding formation pore pressure prediction method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0076] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the formation pore pressure prediction method described in any of the above embodiments.

[0077] Figure 9 This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.

[0078] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0079] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0080] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0081] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0082] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0083] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0084] The electronic devices described above are used to implement the corresponding formation pore pressure prediction methods in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0085] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to execute the formation pore pressure prediction method as described in any of the above embodiments.

[0086] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0087] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the formation pore pressure prediction method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0088] Based on the same concept, corresponding to any of the above embodiments, this application also provides a computer program product, including computer program instructions, which, when run on a computer, cause the computer to perform the method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0089] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application is limited to these examples; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in detail for the sake of brevity.

[0090] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0091] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0092] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.

Claims

1. A method for predicting formation pore pressure, characterized in that, include: Obtain raw logging data and perform outlier correction on the raw logging data to obtain corrected logging data; Calculate the pressure of the overlying strata based on the corrected logging data; Based on the overlying strata pressure and the corrected logging data, the pressure difference between the pore pressure and the drilling fluid column pressure is calculated. Based on the corrected logging data and the pressure difference, the formation pore pressure is predicted using a random forest model; wherein the hyperparameters in the random forest model are determined according to the Grey Wolf optimization algorithm.

2. The method according to claim 1, characterized in that, The outlier correction of the original well logging data includes: In response to determining that the data to be detected in the original well logging data is non-boundary data and that the data to be detected is abnormal data, the data to be detected is replaced by the local mean replacement method; In response to determining that the data to be detected in the original well logging data is boundary data and that the data to be detected is abnormal data, the data to be detected is replaced by the boundary point neighborhood mean replacement method.

3. The method according to claim 2, characterized in that, The step of replacing the data to be detected with a local mean replacement method in response to determining that the data to be detected in the original well logging data is non-boundary data and that the data to be detected is abnormal data includes: A sliding window is constructed based on the data to be detected, and the local mean and local standard deviation are calculated based on the data within the sliding window. The local mean and the local standard deviation are used to determine whether the data to be detected is abnormal. If the data to be detected is determined to be abnormal, the local mean is used to replace the data to be detected.

4. The method according to claim 2, characterized in that, The step of replacing the data to be detected with the boundary point neighborhood mean replacement method in response to determining that the data to be detected in the original well logging data is boundary data and that the data to be detected is abnormal data includes: Based on the original well logging data, calculate the global mean and global standard deviation; Based on the global mean and the global standard deviation, determine whether the data to be detected is abnormal data. If the data to be detected is determined to be abnormal data and the data to be detected is located at the beginning boundary of the sequence, construct a window using the neighboring data after the data to be detected. If the data to be detected is determined to be abnormal data and the data to be detected is located at the end boundary of the sequence, construct a window using the neighboring data in front of the data to be detected. Calculate the mean value of the data within the window, and replace the data to be detected with the mean value.

5. The method according to claim 2, characterized in that, After correcting outliers in the original logging data to obtain corrected logging data, the method further includes: Calculate the new global mean and the new global standard deviation based on the revised logging data; Determine whether the corrected logging data is anomalous based on the new global mean and the new global standard deviation. If it is determined to be anomalous, replace the corrected logging data with the new global mean.

6. The method according to claim 1, characterized in that, The calculation of the pressure difference between pore pressure and drilling fluid column pressure based on the overlying strata pressure and the corrected logging data includes: Calculate the P-wave velocity and S-wave velocity based on the corrected logging data; The bulk modulus of saturated rock is calculated based on the longitudinal wave velocity, the transverse wave velocity, and the bulk density. Based on the saturated rock bulk modulus and the Geismann fluid substitution equation, the dry rock bulk modulus is determined; The hydrostatic pressure difference is determined based on the pressure of the overlying strata and the hydrostatic pressure. The pressure difference is determined based on the hydrostatic pressure difference and the dry rock bulk modulus.

7. The method according to claim 6, characterized in that, The calculation of the bulk modulus of saturated rock based on the longitudinal wave velocity, the transverse wave velocity, and the bulk density includes: The shear modulus is calculated based on the bulk density and the transverse wave velocity. The bulk modulus of the saturated rock is calculated based on the bulk density, the longitudinal wave velocity, and the shear module.

8. The method according to claim 6, characterized in that, Determining the pressure difference based on the hydrostatic pressure difference and the dry rock bulk modulus includes: The normal compaction modulus is determined based on the hydrostatic pressure difference. The pressure difference is calculated based on the normal compaction modulus and the dry rock bulk modulus.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 8.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 8.