A method for predicting pore pressure by optimizing the Eaton method using GMM
By applying Gaussian hybrid clustering algorithm and particle swarm optimization algorithm in the Eaton method, the Eaton index is optimized, and the problem of difficulty in accurately predicting the pore pressure of the whole well section of the complex formation in the existing technology is solved, and efficient and accurate pore pressure prediction of complex formations is achieved.
Patent Information
- Application Number
- CN202410955688.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-17
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-07-17
AI Technical Summary
The prior art is difficult to accurately predict the pore pressure of the whole well section of complex formations during drilling, especially when drilling in non-sedimental rock formations, and the single Eaton method is difficult to effectively predict.
The Gaussian hybrid clustering algorithm (GMM) is used to segment the formation, and the parameters in the Eaton method are optimized to adapt to different formation conditions. Combined with the particle swarm optimization (PSO) algorithm and the least squares method, the Eaton index is optimized, and the pore pressure in the whole well section is predicted.
The accuracy and practicality of predicting the pore pressure of the whole well section of the complex formation is achieved, and the calculation formula and Eaton index can be selected according to different formation conditions, which is suitable for drilling process of drilling into complex formations.
Smart Images

Figure CN118965962B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for optimizing the prediction of pore pressure by the Eaton method using GMM, belonging to the technical field of oil and gas drilling. Background Art
[0002] Pore pressure is the basis for designing the safety density window of drilling fluid and selecting drilling tools, and is the foundation of drilling design. Predicting pore pressure is a key link in oil and gas drilling. The Eaton method comprehensively considers formation compaction and combines logging information to predict pore pressure, which is an efficient and practical method at present. However, as the depth increases, the drilled formations become increasingly complex, and a single Eaton method is difficult to predict the pore pressure of the entire well section. Therefore, the Gaussian mixture clustering algorithm can be used to segment the formation, and the parameters in the Eaton method can be optimized according to different formation segments to adapt to different formations, so as to accurately predict the pore pressure of the entire well section.
[0003] At present, there are various methods for predicting pore pressure in different formations. Among them, the invention patent with the publication number of CN118030038A by CNOOC (China) Limited and CNOOC Research Institute Beijing discloses a method and system for predicting formation pressure in a salt-bearing formation. This method uses the Eaton method to calculate the formation pressure P 碎屑岩 of clastic rock formations and the formation pressure P 脏盐 of dirty salt formations; sets the formation pressure P 纯盐岩 of pure salt rock formations as the hydrostatic pressure P 静水 , and combines and adds P 碎屑岩 , P 纯盐岩 and P 脏盐 to calculate and obtain the formation pressure result P 含盐地层 of the salt-bearing formation. The invention patent with the publication number of CN109736784A by Yangtze University discloses a method for predicting and calculating pore pressure in sedimentary rock formations. This method uses the logging data of sedimentary rock formations to draw a scatter plot to determine the normal compaction section, so as to determine the characteristic parameters of the normal distribution; analyzes the probability of abnormal formation pressure by determining the trend line equation of the well depth and the logarithm of acoustic time difference in the normal compaction mudstone section; calculates the equivalent density of formation pressure in the undercompacted section by the ratio method or the Eaton method.
[0004] Although the above two patents and other existing technical solutions have achieved certain technical effects, the technical means are single, and they can only predict formations with simple lithology, and have little guiding significance for complex formations; when drilling through non-sedimentary rock formations, the technical means such as those in CN109736784A will not be able to effectively predict the formation pore pressure. Summary of the Invention
[0005] The present invention mainly overcomes the deficiencies in the prior art and proposes a method for predicting pore pressure by optimizing the Eaton method using GMM.
[0006] The technical solution provided by the present invention to solve the above technical problems is: a method for predicting pore pressure by optimizing the Eaton method using GMM, comprising the following steps:
[0007] Step 1: Preprocess the acoustic travel time in the logging data and exclude the null values and outliers therein;
[0008] Step 2: Plot the scatter diagram of the preprocessed acoustic travel time and estimate the number of clusters to be clustered based on the scatter diagram;
[0009] Step 3: Based on the preprocessed acoustic travel time, use the EM (expectation maximization) method to continuously iterate the E-step and M-step to make the GMM parameters tend to be stable, so as to obtain the optimal GMM model;
[0010] Step 4: Input the preprocessed acoustic travel time and the estimated number of clusters in Step 2 into the optimized GMM model in Step 3 to cluster the elements of each cluster in the acoustic travel time;
[0011] Step 5: Use the PSO (particle swarm) algorithm to obtain the parameters of the compaction theory of the acoustic travel time of each cluster, and based on each cluster element clustered in Step 4, optimize and select the theoretical acoustic travel time;
[0012] Step 6: Substitute the measured acoustic travel time and the theoretical acoustic travel time in Step 5 into the Eaton method, and optimize the Eaton index based on the measured pore pressure by the least squares method;
[0013] Step 7: Output the pore pressure of the entire well section based on the calculation model optimized in Steps 5 and 6.
[0014] The present invention has the following beneficial effects: A method for predicting pore pressure by optimizing the Eaton method using GMM proposed by the present invention is simple and easy to implement, conforms to engineering practice, can optimize the calculation formula and Eaton index according to different formation conditions, and has strong practical value for drilling in complex formations. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a flow block diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0017] As Figure 1 shown, the present invention provides a method for predicting pore pressure by optimizing the Eaton method using GMM, including the following steps:
[0018] Step 1: Preprocess the acoustic travel time in the logging data and exclude null values and outliers. The preprocessing includes smoothing and normalizing the data;
[0019] The method of smoothing is moving average, and the window width is set to 5:
[0020] yy(1) = y(1)
[0021] yy(2) = (y(1) + y(2) + y(3)) / 3
[0022] yy(3) = (y(1) + y(2) + y(3) + y(4) + y(5)) / 5
[0023] yy(4) = (y(2) + y(3) + y(4) + y(5) + y(6)) / 5
[0024] ……
[0025] yy(n) = (y(n - 2) + y(n - 1) + y(n) + y(n + 1) + y(n + 2)) / 5
[0026] In the formula, y(n) represents the nth data; yy(n) represents the nth data after smoothing;
[0027] The normalization method adopts forward normalization:
[0028]
[0029] In the formula, A represents the parameter before normalization; M represents the parameter after normalization; A min ,A max respectively represent the minimum and maximum values of the parameter.
[0030] Step 2: Plot the scatter plot of the acoustic travel time after preprocessing, estimate the number of clusters to be clustered based on the line chart, and input the Gaussian mixture clustering algorithm.
[0031] Step 3: Based on the preprocessed acoustic travel time, use the EM (Expectation-Maximization) method to continuously iterate the E-step and M-step to make the GMM parameters tend to be stable, so as to obtain the optimal GMM model. Among them, the EM (Expectation-Maximization) method is to solve the parameters of the maximum likelihood function estimation from the complete data containing hidden variables through iteration. The iterative steps mainly include the E-step and M-step. The E-step guesses the hidden data of the model to obtain the expected value, and the M-step performs the maximum likelihood estimation on the complete data to solve the parameters of the Gaussian mixture model.
[0032] Among them, the likelihood function can be expressed as:
[0033]
[0034] In the formula, x1...x m represents the observed data of the mixture model; p(x i , θ) represents the model probability function; θ represents a parameter vector composed of one or more unknown parameters.
[0035] There are hidden variables in the complete data. Directly solving the maximum value of the likelihood estimation function to estimate the parameters of the mixture model will make the process very complicated. First, the hidden variables can be found, then take the logarithm of the above formula, establish the lower bound of the logarithm of the likelihood function, and continuously iterate and optimize it:
[0036]
[0037] In the formula, p(x i , z i , θ) represents the model probability function after determining the hidden variables;
[0038] The maximum likelihood function does not consider the probability generated by the model itself, that is, the prior probability:
[0039]
[0040] In the formula, p(z N ) represents the likelihood probability of the hidden variable z N itself; α m represents the weight of the m-th Gaussian distribution in the mixture model, and satisfies
[0041] Based on the prior probability, Bayes' theorem combined with the initial values of the Gaussian mixture model parameters can determine the influence degree of the m-th classification model in the Gaussian mixture model on the observed data (x1, x2,...x j ,..., x N ), that is, the maximum posterior probability:
[0042]
[0043] wherein, φ(x j |θ m ) represents the probability density function of the m-th Gaussian distribution model when the observed value is x j .
[0044] Step 4: Input the preprocessed acoustic travel time and the number of clusters estimated in Step 2 into the optimized GMM model in Step 3 to cluster the elements of each cluster in the acoustic travel time; wherein, Step 3 mainly optimizes μ m , and α m , μ m where represents the mean of the m-th Gaussian distribution, represents the variance matrix of the m-th Gaussian distribution; the core of GMM is the Gaussian mixture clustering model, and the elements of each cluster clustered from the preprocessed acoustic travel time conform to different Gaussian distribution models. By weighted training of multiple Gaussian distribution models in each cluster, the probability that the elements of each cluster belong to the cluster can be obtained;
[0045] Let the random variable be x, and the Gaussian mixture model is composed of M Gaussian distributions. This model can be expressed as:
[0046]
[0047] wherein, θ m represents φ(x|θ m ) can be expressed as:
[0048]
[0049] Step 5: Use PSO to obtain the parameters of the acoustic time compaction theory for each cluster, and based on the elements of each cluster clustered in Step 4, optimize the theoretical acoustic travel time; the compaction theory formula is as follows:
[0050]
[0051] wherein, DNTH represents the acoustic travel time of sound waves propagating in the formation theoretically, μs / ft; T PMUD represents the acoustic travel time of the rock skeleton, μs / ft; VD represents the vertical depth of the wellbore, m; C1 and C2 represent adjustable coefficients, which are optimized by PSO and are dimensionless;
[0052] Then, use PSO to optimize the parameters of the formula to obtain the calculated value. PSO starts from a random solution and searches for the optimal solution through iteration. Finally, the fitness function determines the quality of the solution; the particle velocity and position update formulas at the next iteration t + 1 of PSO are:
[0053] v i,d (t + 1) = ωv i,d (t) + c1r1(pbesti,d (t) - x i,d (t)) + c a r2(gbest d (t) - x i,d (t))
[0054] x i,d (t + 1) = x i,d (t) + v i,d (t + 1)
[0055] Among them, x i,d (t) represents the position of particle i in the d - dimensional space at the iteration number t; vi ,d (t) represents the velocity of particle i in the d - dimensional space at the iteration number t; pbest i (t) = [p i,1 (t), p i,2 (t), …, p i,D (t)] represents the historical best position of particle i in the d - dimensional space at the iteration number t; gbest(t) = [g1(t), g2(t), …, g D (t)] represents the historical best position of all particles in the entire particle swarm in the d - dimensional space at the iteration number t; r1 and r2 represent random numbers taking values in the interval [0, 1]; c1 and c2 represent learning factors, usually taking values in the interval [0, 2] to control the learning step size; w represents the inertia weight, dimensionless;
[0056] Finally, the correlation coefficient is used to judge the quality of the optimization. The correlation coefficient can be expressed as:
[0057]
[0058] In the formula, r(x, y) represents the correlation between data x and y, dimensionless; Cov(x, y) represents the covariance between data x and y, dimensionless; Var(x) and Var(y) respectively represent the variances of data x and y, dimensionless.
[0059] Step 6: Substitute the measured acoustic travel - time and the theoretical acoustic travel - time in Step 5 into the Eaton method, and optimize the Eaton exponent based on the measured pore pressure by the least - squares method. The Eaton formula can be expressed as:
[0060]
[0061] In the formula, p p represents the predicted pore pressure, g / cm 3 ; σ v represents the overburden pressure of the wellbore, g / cm 3 ; p wRepresents the hydrostatic pressure of the wellbore, g / cm 3 ; AC true Represents the measured acoustic travel time, μs / ft; n represents the Eaton index, dimensionless.
[0062] Step 7: Output the pore pressure of the entire well section based on the calculation model optimized in Steps 5 and 6.
[0063] As mentioned above, it is not any form of limitation to the present invention. Although the present invention has been disclosed through the above embodiments, it is not intended to limit the present invention. Any person skilled in the art, within the scope of the technical solution of the present invention, can make changes or modifications to equivalent embodiments by using the technical content disclosed above. However, as long as it does not depart from the technical solution of the present invention, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A method for predicting pore pressure using GMM to optimize Eaton's method, characterized in that: The following steps are involved: Step 1: pre-process the acoustic time difference in the logging data and exclude the null values and abnormal values; Step 2: Draw a scatter plot of the acoustic time difference after preprocessing, and estimate the number of clusters that need to be clustered based on the scatter plot; Step 3: Based on the preprocessed acoustic wave time difference, the EM (expectation maximization) method is used to continuously iterate the E-step and M-step to stabilize the GMM parameters, so as to obtain the optimal GMM model; Step 4: input the preprocessed acoustic time difference and the number of clusters estimated in step 2 into the GMM model optimized in step 3 to cluster the elements of each cluster in the acoustic time difference; Step 5: Use PSO to obtain the parameters of the acoustic wave time difference compaction theory for each cluster, and optimize the theoretical acoustic wave time difference based on each cluster element clustered in step 4; Step 6, substituting the measured acoustic time difference and the theoretical acoustic time difference in step 5 into the Eaton method, and optimizing the Eaton index based on the measured pore pressure by the least square method; Step 7: Output the pore pressure of the entire well section based on the calculation model optimized in steps 5 and 6.
2. The method according to claim 1, characterized in that In the step 1, the preprocessing includes smoothing and normalization; The smoothing method is moving average, and the window width is set to 5: yy(1)=y(1) yy(2)=(y(1)+y(2)+y(3)) / 3 yy(3)=(y(1)+y(2)+y(3)+y(4)+y(5)) / 5 yy(4)=(y(2)+y(3)+y(4)+y(5)+y(6)) / 5 …… yy(n)=(y(n-2)+y(n-1)+y(n)+y(n+1)+y(n+2)) / 5 In the formula, y(n) represents the nth data; yy(n) represents the nth data after smoothing; The normalization method uses forward normalization: In the formula, A represents the parameters before normalization; M represents the parameters after normalization; A min , A max Respectively represent the minimum and maximum values of the parameter.
3. The method according to claim 1, characterized in that: In step 2, the pre-processed acoustic time difference is also called complete data, which includes the observed observation data X = {x1, x2, ..., x j ,…,x N } and unobserved implicit variables Z = {z1,z2,…,z N }.
4. The method according to claim 1, characterized in that: In step 3, the EM (expectation maximization) method is to solve the parameters of the maximum likelihood function estimation from the complete data containing implicit variables through iteration; the iterative steps mainly include E-step and M-step, the E-step guesses the implicit data of the model to obtain the expected value, and the M-step performs maximum likelihood estimation on the complete data to solve the parameters of the Gaussian mixture model; Among them, the likelihood function can be expressed as: In the formula, x1...x m represents the observed data of the mixed model; p(x i ,θ) represents the model probability function; θ represents a parameter vector composed of one or more unknown parameters; The complete data contains implicit variables. Directly solving the maximum value of the likelihood estimation function to estimate the parameters of the mixed model will make the process very complicated. We can first find the implicit variables, then take the logarithm of the above formula, establish the lower bound of the logarithm of the likelihood function, and continuously iterate and optimize it: In the formula, p(x i ,z i ,θ) represents the model probability function after determining the latent variables; The maximum likelihood function does not take into account the probability generated by the model itself, that is, the prior probability: In the formula, p(z N ) represents the implicit variable z N The likelihood probability of itself; α m Represents the weight of the mth Gaussian distribution in the mixture model and satisfies Based on the prior probability, Bayesian theorem combined with the initial values of the Gaussian mixture model parameters can determine the mth classification model in the Gaussian mixture model for the observed data (x1, x2, ...x j ,...,x N ), that is, the maximum posterior probability: In the formula, φ(x j |θ m ) indicates that the observed value is x j is the probability density function of the mth Gaussian distribution model.
5. The method according to claim 1, characterized in that In step 4, step 3 mainly optimizes μ m , and α m , μ m where represents the mean of the mth Gaussian distribution, Represents the variance matrix of the mth Gaussian distribution; the core of the Gaussian mixture clustering algorithm is the Gaussian mixture clustering model. The cluster elements clustered by the pre-processed acoustic wave time difference conform to different Gaussian distribution models. By weighted training of multiple Gaussian distribution models in each cluster, the probability of each cluster element belonging to the cluster can be obtained; Assume that the random variable is x, and the Gaussian mixture model is composed of M Gaussian distributions. The model can be expressed as: In the formula, θ m express φ(x|θ m ) can be expressed as:
6. The method according to claim 1, characterized in that In step 4, the compaction theory calculation formula is: Where, DNTH is the theoretical acoustic time difference of sound waves propagating in the formation, us / ft; T PMUD represents the acoustic time difference of rock skeleton, us / ft; VD represents the vertical depth of wellbore, m; C1 and C2 represent adjustable coefficients, which are optimized by PSO and dimensionless; PSO starts from a random solution and searches for the optimal solution through iteration. Finally, the fitness function determines the quality of the solution. The particle speed and position update formula at the next iteration t+1 of PSO is: v i,d (t+1)=ωv i,d (t)+c1r1(pbest i,d (t)-x i,d (t))+c2r2(gbest d (t)-x i,d (t)) x i,d (t+1)=x i,d (t)+v i,d (t+1) Among them, x i,d (t) represents the position of particle i in d-dimensional space when the number of iterations is t; v i,d (t) represents the velocity of particle i in d-dimensional space when the number of iterations is t; pbest i (t) = [p i,1 (t), p i,2 (t),…,p i,D (t)] represents the historical optimal position of particle i in d-dimensional space when the number of iterations is t; gbest(t) = [g1(t), g2(t), …, g D (t)] represents the historical optimal position of all particles in the entire particle swarm in the d-dimensional space when the number of iterations is t; r1 and r2 represent random numbers between the interval [0,1]; c1 and c2 represent learning factors, which are usually between the interval [0,2] and are used to control the learning step size; ω represents the inertia weight, which is dimensionless.
7. The method according to claim 1, characterized in that The correlation coefficient in step 4 is usually used to measure the linear relationship between two sets of data and can be expressed as: In the formula, r(x, y) represents the correlation between data x and g, dimensionless; Cou(x, g) represents the covariance of data x and g, dimensionless; Var(x) and Var(y) represent the variance of data x and y, dimensionless.
8. The method according to claim 1, characterized in that The Eaton formula in step 5 can be expressed as: In the formula, p p is the predicted pore pressure, g / cm 3 ; σ v Indicates the overburden pressure of the wellbore, g / cm 3 ;p w Indicates the hydrostatic pressure of the wellbore, g / cm 3 ; AC true represents the measured sound wave time difference, us / ft; n represents the Eaton index, dimensionless; The pore pressure of the entire well section can be obtained by optimizing the DNTH input in step 4 using the least squares method.
Citation Information
Patent Citations
Sedimentary rock stratum pore pressure prediction calculation method
CN109736784A
Formation pressure prediction method and system for salt-containing formation
CN118030038A
High-pressure container detection method based on infrared image splicing fusion algorithm
CN110706191A
Apparatus and methods for improved subsurface data processing systems
CN111542819A