Intelligent Prediction Method for Total Organic Carbon Content in Shales Based on Empirical Constraints

By combining empirical constraints and GWO-SVR algorithm, the SVR model is optimized by using DEN-GR to correct the △LogR model, the problem of insufficient accuracy of the prediction of shale total organic carbon content is solved, and higher prediction accuracy and reliability are achieved, supporting the exploration and development of shale reservoirs.

CN119418822BActive Publication Date: 2025-07-11CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411457233.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-18
Publication Date
2025-07-11
Estimated Expiration
2044-10-18

AI Technical Summary

Technical Problem

The prior art lacks accuracy in the prediction of total organic carbon content of shale, especially in the case of complex geological background, the mapping relationship between conventional logging curves and total organic carbon content is weak, resulting in great limitations in the evaluation of shale reservoirs.

Method used

Combining empirical constraints and GWO-SVR algorithm, the △LogR model is corrected through DEN-GR, the SVR model parameters are optimized using the GWO algorithm, and combining experimental data and conventional logging curves to achieve accurate prediction of the total organic carbon content of shale.

Benefits of technology

It improves the accuracy and reliability of the prediction of total organic carbon content of shale, overcomes the limitations of geological background environment, and provides a basis for shale reservoir exploration and development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119418822B_ABST
    Figure CN119418822B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent prediction method for total organic carbon content in shale based on empirical constraints, which relates to the technical field of geophysical logging. According to the logging data in the research area, the present invention reconstructs the acoustic impedance curve, obtains a plurality of TOC sample data, constructs a sample data set including a training set and a test set, determines an empirical model for the total organic carbon content based on the DEN-GR corrected △LogR model, and introduces it into the fitness function of the GWO algorithm. The GWO algorithm and the TOC sample data in the training set are used to optimize the SVR model, determine the optimal solutions of the penalty coefficient and kernel function parameters in the SVR model, obtain the trained SVR model, and use the TOC sample data in the test set to verify the accuracy of the SVR model for predicting the total organic carbon content in shale, and apply the verified SVR model to the intelligent prediction of the total organic carbon content in shale. The present invention combines empirical constraints with the GWO-SVR algorithm to achieve accurate prediction of the total organic carbon content in shale, providing a basis for guiding the exploration and development of shale reservoirs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of geophysical logging, and particularly relates to an intelligent prediction method for total organic carbon content in shale based on empirical constraints. Background Art

[0002] As an important unconventional oil and gas resource, shale oil and gas have become important strategic replacement resources. The total organic carbon content (TOC), as an important parameter in shale reservoirs, can directly characterize the generation, storage, and production capacity of shale gas / shale oil. Usually, after collecting source rock samples, experts obtain high-precision TOC content through laboratory analysis and measurement. This method usually requires crushing the samples and using instruments such as Rock-Eval pyrolyzers for pyrolysis analysis. Although the obtained TOC content results are direct and accurate, it requires destructive testing on core or cuttings samples; moreover, it is to calculate TOC by measuring the carbon content in the samples using an elemental analyzer. These two methods are costly and time-consuming, and are not conducive to the promotion of the entire well section.

[0003] In order to further obtain the total organic carbon content distribution of continuous longitudinal formations based on experimental data, experts in this field have proposed various empirical formulas, such as using conventional logging curve fitting, the △logR method, etc. However, with the increasing pressure of increasing oil reserves and production in oilfields, the standard for TOC accuracy is also increasing day by day. There are significant limitations in relying solely on empirical evaluation methods. Moreover, when faced with complex geological backgrounds and low-background contrast conditions, the mapping relationship between conventional logging curves and total organic carbon content is weak during shale reservoir evaluation.

[0004] The existence of the above problems has restricted the application of conventional linear processing and traditional empirical methods in predicting the total organic carbon content in shale. Therefore, how to make full use of existing experimental data and conventional logging curves, combine empirical constraints and the non-linear processing ability of intelligent machine learning, and propose an intelligent prediction method for total organic carbon content in shale based on empirical constraints is of great significance for the research of total organic carbon content in shale. Summary of the Invention

[0005] The present invention aims to solve the above problems and proposes an intelligent prediction method for total organic carbon content in shale based on empirical constraints. By combining empirical constraints with the GWO-SVR algorithm, based on the DEN-GR corrected △LogR model and introducing it into the GWO-SVR algorithm, through introducing the total organic carbon content empirical model into the model parameter optimization process of the SVR model, accurate prediction of the total organic carbon content in shale is realized, providing a basis for the exploration and development of shale reservoirs.

[0006] The present invention adopts the following technical solutions:

[0007] Intelligent prediction method for total organic carbon content of shale based on empirical constraints, comprising the following steps:

[0008] Step 1, obtain the logging data of the study area, reconstruct the acoustic impedance curve, select the logging curves related to the total organic carbon content, and obtain multiple sample data of the total organic carbon content of shale;

[0009] Step 2, construct a sample data set, and randomly allocate the sample data of the total organic carbon content of shale to the training set and the test set of the sample data set;

[0010] Step 3, based on the DEN-GR corrected △LogR model, determine the empirical model of the total organic carbon content in the study area;

[0011] Step 4, optimize the SVR model based on the GWO algorithm, introduce the empirical model of the total organic carbon content into the gray wolf optimization process, combine the sample data of the total organic carbon content of shale in the training set, and use the GWO algorithm to determine the optimal model parameters of the SVR model to obtain the trained SVR model;

[0012] Step 5, input the sample data of the total organic carbon content of shale in the test set into the trained SVR model, use the trained SVR model to predict the total organic carbon content of the sample data of the total organic carbon content of shale, and verify the accuracy of the trained SVR model.

[0013] Preferably, in the said Step 1, according to the acoustic travel time curve and the density curve, use the ratio between the acoustic travel time and the formation density at different depths to reconstruct the acoustic impedance curve;

[0014] The logging data of the study area are the conventional logging curves that have been logged in the study area, including the acoustic travel time curve, the density curve, the neutron porosity curve, the natural gamma curve and the resistivity curve. According to the acoustic travel time curve and the density curve, use the ratio between the acoustic travel time and the formation density at different depths to reconstruct the acoustic impedance curve;

[0015] Based on the total organic carbon content curve that has been logged, select the logging curves that can change with the total organic carbon content from the conventional logging curves, combine with the acoustic impedance curve to obtain the logging curves related to the total organic carbon content, and based on the logging curves related to the total organic carbon content and the total organic carbon content curve at different depths, obtain multiple sample data of the total organic carbon content of shale.

[0016] Preferably, the sample data of the total organic carbon content of shale consists of multiple characteristic parameter values and the measured value of the total organic carbon content of shale at the same depth. Among them, the characteristic parameter values are determined according to the logging curves related to the total organic carbon content, and the measured value of the total organic carbon content of shale is determined according to the total organic carbon content curve.

[0017] Preferably, in step 3, based on the well logging curves in the study area, the △LogR model is corrected based on DEN-GR, and the model parameters are determined by fitting the well logging curves in the study area. The empirical model for total organic carbon content is determined as follows:

[0018] TOC = [a × lg(GR) + b × DEN + c] × ΔlogR

[0019] Where:

[0020] The △LogR model corrected based on DEN-GR is:

[0021]

[0022] In the formula, TOC is the total organic carbon content; a, b, and c are all model parameters; GR is the natural gamma value; DEN is the formation density value; R t is the resistivity, with the unit of Ω·m; R baseline is the resistivity corresponding to the baseline in the resistivity curve, with the unit of Ω·m; Δt is the acoustic time difference value, with the unit of μs / m; Δt baseline is the acoustic time difference value corresponding to the baseline in the acoustic time difference curve, with the unit of μs / m.

[0023] Preferably, in step 4, optimizing the SVR model based on the GWO algorithm includes the following sub-steps:

[0024] Step 4.1, initialize the GWO algorithm, set the type of gray wolf individuals, the positions of the wolf packs, and the fitness function in the GWO algorithm, and determine the number of search agents, the maximum number of optimizations, and the initial value of the search space of the GWO algorithm;

[0025] Step 4.2, set the initial values of the penalty coefficient C and the kernel function parameter gamma in the SVR model, and input the sample data of the shale total organic carbon content in the training set into the SVR model and the empirical model of total organic carbon content;

[0026] Step 4.3, based on the GWO algorithm, perform an optimization search to obtain the solutions searched by each individual in the wolf pack, calculate the mean square error using the SVR model, and cooperate with the empirical model of total organic carbon content to calculate the penalty term to determine the fitness function corresponding to each solution;

[0027] Step 4.4, compare the fitness function values of each solution, determine the best solution, update the positions of the wolf pack according to the positions of the Alpha wolf, Beta wolf, and Delta wolf, and re-perform an optimization search based on the GWO algorithm to obtain the solutions searched by each individual in the wolf pack, and update the coefficient vector and the solution when the wolf pack surrounds the prey;

[0028] Step 4.5: Repeat Step 4.3 and Step 4.4 until the preset maximum number of optimization times is reached. Stop the optimization search using the GWO algorithm, output the optimal solutions of the penalty coefficient C and the kernel function parameter gamma in the SVR model, and obtain the trained SVR model.

[0029] Preferably, the fitness function is:

[0030]

[0031] In the formula, Fitness is the fitness; w is the input sample; MSE is the mean square error; Penalty is the penalty term based on the total organic carbon content empirical model, which is used to ensure the data scale balance between the mean square error and the penalty term; λ is the penalty factor, which is used to weigh the importance of the mean square error MSE and the penalty term; N is the total number of shale total organic carbon content sample data in the training set; i is the sample serial number; y i is the measured TOC value of the i-th shale total organic carbon content sample data; is the predicted TOC value of the i-th shale total organic carbon content sample data; j is the characteristic parameter serial number; M is the total number of characteristic parameters; x j is the value of the j-th characteristic parameter; is the partial derivative of the predicted TOC value with respect to the characteristic parameter value x j , which is used to represent the change rate of the predicted TOC value with respect to the j-th characteristic parameter; is the partial derivative of the TOC empirical value calculated by the total organic carbon content empirical model with respect to the j-th characteristic parameter, which is used to represent the change rate of the TOC empirical value with respect to the j-th characteristic parameter.

[0032] Preferably, in the GWO algorithm, the individuals within the wolf pack are divided into four categories according to their social ranks, which are Alpha wolf, Beta wolf, Delta wolf, and Omega wolf from the highest to the lowest social rank. Among them, the Alpha wolf is the leader of the wolf pack, ranking first in the wolf pack, and is responsible for making decisions and guiding the actions of the wolf pack; the Beta wolf ranks second in the wolf pack and is used to assist the Alpha wolf in making decisions and guiding the actions of other lower-rank wolves; the Delta wolf ranks third in the wolf pack and is responsible for scouting and guarding; the Omega wolf ranks fourth in the wolf pack and is used to execute the orders of the Alpha wolf, Beta wolf, and Delta wolf;

[0033] When performing the optimization search based on the GWO algorithm, the solution searched by the Alpha wolf is the best solution, the solution searched by the Beta wolf is the second-best solution, the solution searched by the Delta wolf is the third-best solution, and the solution searched by the Omega wolf is the remaining alternative solutions;

[0034] The collective hunting process of the wolf pack alternates between global search and local search, including four stages: surrounding the prey, hunting, attacking the prey, and searching for the prey;

[0035] The expression for surrounding the prey is:

[0036]

[0037] In the formula, t is the current optimization times; is the distance vector between the gray wolf and the prey; is the distance vector between the gray wolf and the prey at the next optimization; are all coefficient vectors; is the position vector of the prey; is the position vector of the gray wolf; represents a linear decrease from 2 to 0 during the optimization process; are all random vectors in [0, 1];

[0038] The expression for hunting is:

[0039]

[0040]

[0041] In the formula are the distance vectors of Alpha wolf, Beta wolf, and Delta wolf respectively; is the coefficient vector of Alpha wolf, is the coefficient vector of Beta wolf, is the coefficient vector of Delta wolf, are all used to adjust the distance of individuals in the wolf pack; are the current positions of Alpha wolf, Beta wolf, and Delta wolf respectively; is the new potential position updated by Omega wolf according to the position and distance of Alpha wolf, is the new potential position updated by Omega wolf according to the position and distance of Beta wolf, is the new potential position updated by Omega wolf according to the position and distance of Delta wolf; is the coefficient vector determined by Omega wolf according to the position change of Alpha wolf, is the coefficient vector determined by Omega wolf according to the position change of Beta wolf, is the coefficient vector determined by Omega wolf according to the position change of Delta wolf; is the final position of Omega wolf. When Omega wolf tends to deviate from the prey; when When it is, Omega wolves tend to approach the prey.

[0042] Preferably, the SVR model selects a radial basis function as the kernel function to handle complex non-linear problems; the model parameters of the SVR model include the penalty coefficient C and the kernel function parameter gamma. Among them, the penalty coefficient C determines the penalty strength of the SVR model for misclassification and is used to control the tolerance of the SVR model to errors; the kernel function parameter gamma determines the radiation range of the SVR model and is used to control the influence range of a single sample data.

[0043] Preferably, in step 5, the sample data of the total organic carbon content of shale in the test set is input into the trained SVR model. The trained SVR model predicts the predicted value of the total organic carbon content of shale based on the sample data of the total organic carbon content of shale, and compares it with the measured value of the total organic carbon content of shale in the sample data of the total organic carbon content of shale. Calculate the accuracy rate predicted by the trained SVR model and compare it with the preset precision value. If the accuracy rate predicted by the trained SVR model is lower than the preset precision value, return to step 4 to continue training the SVR model. Otherwise, obtain the verified SVR model and directly use the verified SVR model to predict the total organic carbon content of shale.

[0044] The present invention has the following beneficial effects:

[0045] The present invention proposes an intelligent prediction method for the total organic carbon content of shale based on empirical constraints. This method uses TOC experimental data and conventional logging curves as the basic data support, and based on the DEN-GR corrected △LogR model, determines the empirical model of the total organic carbon content in the study area and introduces it into the GWO algorithm. By introducing an empirical penalty term in the fitness function of the GWO algorithm to constrain the consistency between the predicted value and the overall trend of the measured total organic carbon content of shale, and using the GWO algorithm to realize the parameter optimization of the SVR model. The method of the present invention combines empirical constraints with machine learning, overcomes the limitation of the current total organic carbon content prediction method of shale in shale reservoir evaluation by the geological background environment, and effectively improves the accuracy and reliability of the total organic carbon content prediction of shale. Description of the Drawings

[0046] Figure 1 It is a flow chart of an intelligent prediction method for the total organic carbon content of shale based on empirical constraints of the present invention.

[0047] Figure 2 It is a flow chart of the GWO algorithm of the present invention.

[0048] Figure 3 It is a prediction diagram of the total organic carbon content of shale in Well Y in Example 2. Detailed Embodiments

[0049] Taking a certain research area as an example, the specific implementation manners of the present invention will be further described in conjunction with the accompanying drawings:

[0050] Example 1

[0051] The present invention provides an intelligent prediction method for total organic carbon content of shale based on empirical constraints. As Figure 1 shown, the method specifically includes the following steps:

[0052] Step 1: Select the area where a certain shale reservoir is located as the research area, and obtain the logging data of the research area, that is, the conventional logging curves of the wells already logged in the research area, including the acoustic travel-time curve, density curve, neutron porosity curve, natural gamma curve, and resistivity curve. According to the acoustic travel-time curve and density curve, the impedance curve is reconstructed using the ratio between the acoustic travel-time and formation density at different depths to reflect the physical properties of the rock.

[0053] Then, based on the total organic carbon (TOC) content curve of the wells already logged, select the logging curves that can vary with the total organic carbon content from the conventional logging curves, and combine them with the impedance curve to obtain the logging curves related to the total organic carbon content. Based on the logging curves related to the total organic carbon content and the total organic carbon content curve at different depths, a plurality of sample data of the total organic carbon content of shale are obtained.

[0054] The sample data of the total organic carbon content of shale consists of multiple characteristic parameter values and the measured value of the total organic carbon content of shale at the same depth. Among them, the characteristic parameter values are determined according to the logging curves related to the total organic carbon content, and the measured value of the total organic carbon content of shale is determined according to the total organic carbon content curve.

[0055] Step 2: Construct a sample data set, and randomly allocate the sample data of the total organic carbon content of shale to the training set and test set of the sample data set. Among them, the training set is used to train the SVR model, and the test set is used to test the accuracy of the prediction of the total organic carbon content of shale by the trained SVR model.

[0056] Step 3: Based on the DEN-GR corrected △LogR model, determine the empirical model of the total organic carbon content of the research area.

[0057] The traditional △LogR model is a TOC calculation method based on acoustic travel-time, resistivity, and maturity parameters proposed by Passey et al. in 1990. When the traditional △LogR model is applied to calculate the TOC content of deep hydrocarbon source rocks, the error is relatively large. Therefore, in this example, the traditional △LogR model is corrected based on DEN-GR and fitted with the logging data of the research area to determine the empirical model of the total organic carbon content of the research area as:

[0058] TOC = [a×lg(GR) + b×DEN + c]×ΔlogR

[0059] Wherein,

[0060] The ΔLogR model based on DEN-GR correction is as follows:

[0061]

[0062] In the formula, TOC is the total organic carbon content; a, b, and c are all model parameters; GR is the natural gamma value; DEN is the formation density value; R t is the resistivity, with the unit of Ω·m; R baseline is the resistivity corresponding to the baseline in the resistivity curve, with the unit of Ω·m; Δt is the acoustic time difference, with the unit of μs / m; Δt baseline is the acoustic time difference corresponding to the baseline in the acoustic time difference curve, with the unit of μs / m.

[0063] In this embodiment, taking well X in the study area as an example, according to the well logging data in the study area, the resistivity R corresponding to the baseline in the resistivity curve is determined baseline to be 1 Ω·m, and the acoustic time difference Δt corresponding to the baseline in the acoustic time difference curve baseline is 80 μs / m. According to the well logging data in the study area, the empirical model of the total organic carbon content of well X is determined as:

[0064] TOC = [9.67×lg(GR) - 10.62×DEN + 9.19]×ΔlogR.

[0065] Step 4: Optimize the SVR model based on the GWO algorithm, introduce the empirical model of the total organic carbon content into the gray wolf optimization process, combine with the shale total organic carbon content sample data in the training set, and use the GWO algorithm to determine the optimal model parameters of the SVR model to obtain the trained SVR model.

[0066] Optimizing the SVR model based on the GWO algorithm includes the following sub-steps:

[0067] Step 4.1: Initialize the GWO algorithm (i.e., the gray wolf optimization algorithm), set the type of gray wolf individuals, the positions of the wolf packs, and the fitness function in the GWO algorithm, and determine the number of search agents, the maximum number of optimizations, and the initial value of the search space of the GWO algorithm.

[0068] In this embodiment, the GWO algorithm can automatically adjust the model parameters in the SVR (Support Vector Regression) model according to the characteristics and internal structure of the dataset. Specifically, it adjusts the penalty coefficient C and the kernel function parameter gamma of the SVR model. The goal of the SVR model is to determine the optimal hyperplane in the given dataset, such that the distances from all data points in the dataset to the optimal hyperplane are as small as possible. Among them, the kernel function, the penalty coefficient C, and the kernel function parameter gamma are the key parameters of the SVR model. In this embodiment, the radial basis function (RBF) is selected as the kernel function to handle complex non-linear problems; the penalty coefficient C determines the penalty strength of the SVR model for misclassification and is used to control the tolerance of the SVR model to errors. The larger the value of the penalty coefficient C, the heavier the penalty of the SVR model for errors; the kernel function parameter gamma determines the radiation range of the SVR model and is used to control the influence range of individual sample data. If the kernel function parameter gamma is too large, it may lead to overfitting, and if it is too small, the model may be too simple; the GWO algorithm is used to optimize and determine the penalty coefficient C and the kernel function parameter gamma of the SVR model, and by adjusting the penalty coefficient C and the kernel function parameter gamma, the complexity of the SVR model and the accuracy of the prediction results are balanced.

[0069] Furthermore, using the total organic carbon content empirical model, the penalty term based on the △LogR model with empirical constraints is used to determine the consistency of the overall trend between the predicted value of the penalty term constrained SVR model and the calculated value of the TOC empirical model. In this embodiment, the fitness function is set as follows:

[0070]

[0071] In the formula, Fitness is the fitness; w is the input sample; MSE is the mean square error; Penalty is the penalty term based on the total organic carbon content empirical model, which is used to ensure the data scale balance between the mean square error and the penalty term; λ is the penalty factor, which is used to weigh the importance of the mean square error MSE and the penalty term and ensure the data scale balance between the mean square error and the penalty term. Generally, it is determined by cross-validation. In this embodiment, five-fold cross-validation is adopted to determine that the value of the penalty factor λ is 0.1; N is the total number of shale total organic carbon content sample data in the training set; i is the sample serial number; y i is the measured value of TOC of the i-th shale total organic carbon content sample data; is the predicted value of TOC of the i-th shale total organic carbon content sample data; j is the characteristic parameter serial number; m is the total number of characteristic parameters; x j is the value of the j-th characteristic parameter; is the partial derivative of the TOC predicted value with respect to the characteristic parameter value x j and is used to represent the change rate of the TOC predicted value with respect to the j-th characteristic parameter; The partial derivative of the TOC empirical value calculated by the total organic carbon content empirical model with respect to the j-th characteristic parameter, which is used to represent the change rate of the TOC empirical value with respect to the j-th characteristic parameter.

[0072] Step 4.2: Set the initial values of the penalty coefficient C and the kernel function parameter gamma in the SVR model, and input the shale total organic carbon content sample data in the training set into the SVR model and the total organic carbon content empirical model.

[0073] Step 4.3: Based on the GWO algorithm for optimization search, obtain the solutions searched by each individual in the wolf pack, calculate the mean square error using the SVR model, and cooperate with the total organic carbon content empirical model to calculate the penalty term to determine the fitness function corresponding to each solution.

[0074] In this embodiment, in the GWO algorithm, the individuals in the wolf pack are divided into four categories according to the social rank, which are Alpha wolf, Beta wolf, Delta wolf, and Omega wolf from high to low social rank. Among them, the Alpha wolf is the leader of the wolf pack, ranking first in the wolf pack, responsible for making decisions and guiding the actions of the wolf pack; the Beta wolf ranks second in the wolf pack, assisting the Alpha wolf in making decisions and guiding the actions of other lower-rank wolves; the Delta wolf ranks third in the wolf pack, responsible for scouting and warning; the Omega wolf ranks fourth in the wolf pack, executing the orders of the Alpha wolf, Beta wolf, and Delta wolf.

[0075] When performing optimization search based on the GWO algorithm, the solution searched by the Alpha wolf is the best solution, the solution searched by the Beta wolf is the second-best solution, the solution searched by the Delta wolf is the third-best solution, and the solution searched by the Omega wolf is the remaining alternative solution.

[0076] The collective hunting process of the wolf pack alternates between global search and local search, including four stages: surrounding the prey, hunting, attacking the prey, and searching for the prey. Among them, the expression for surrounding the prey is:

[0077]

[0078] In the formula, t is the current optimization number; is the distance vector between the gray wolf and the prey; is the distance vector between the gray wolf and the prey in the next optimization; are all coefficient vectors; is the position vector of the prey; is the position vector of the gray wolf; indicates a linear decrease from 2 to 0 during the optimization process; All are random vectors in [0, 1].

[0079] The expression of the hunting is as follows:

[0080]

[0081]

[0082] In the formula, are the distance vectors of Alpha wolf, Beta wolf, and Delta wolf respectively; is the coefficient vector of Alpha wolf, is the coefficient vector of Beta wolf, is the coefficient vector of Delta wolf, All are used to adjust the distances of individuals in the wolf pack; are the current positions of Alpha wolf, Beta wolf, and Delta wolf respectively; is the new potential position updated by Omega wolf according to the position and distance of Alpha wolf, is the new potential position updated by Omega wolf according to the position and distance of Beta wolf, is the new potential position updated by Omega wolf according to the position and distance of Delta wolf; is the coefficient vector determined by Omega wolf according to the position change of Alpha wolf, is the coefficient vector determined by Omega wolf according to the position change of Beta wolf, is the coefficient vector determined by Omega wolf according to the position change of Delta wolf; is the final position of Omega wolf. When Omega wolf tends to deviate from the prey; when Omega wolf tends to approach the prey.

[0083] Step 4.4: Compare the fitness function values of each solution, determine the best solution, update the positions of the wolf pack according to the positions of Alpha wolf, Beta wolf, and Delta wolf, re-optimize and search based on the GWO algorithm to obtain the solutions searched by each individual in the wolf pack, and update the coefficient vectors and solutions when the wolf pack surrounds the prey.

[0084] Step 4.5: Repeat Step 4.3 and Step 4.4 until the preset maximum number of optimization times is reached, stop the optimization search using the GWO algorithm, output the optimal solutions of the penalty coefficient C and the kernel function parameter gamma in the SVR model, and obtain the trained SVR model.

[0085] In this embodiment, each time the optimization search is performed based on the GWO algorithm, ifFigure 2 As shown in the figure, first, use the wolf pack search to obtain multiple solutions. By evaluating the fitness functions of each solution and sorting them, after obtaining the best solution, reset the positions of the wolf pack according to the positions of the Alpha wolf, Beta wolf, and Delta wolf, update the coefficient vector and the solution and evaluate them until the preset maximum number of optimization times is reached, that is, determine the penalty coefficient C and the kernel function parameter gamma in the best SVR model, and set the SVR model according to the optimal solutions of the penalty coefficient C and the kernel function parameter gamma.

[0086] Step 5: Input the sample data of the total organic carbon content of shale in the test set into the trained SVR model, and use the trained SVR model to predict the total organic carbon content of the sample data of the total organic carbon content of shale to verify the accuracy of the trained SVR model.

[0087] In this embodiment, input the sample data of the total organic carbon content of shale in the test set into the trained SVR model, use the trained SVR model to predict the predicted value of the total organic carbon content of shale according to the sample data of the total organic carbon content of shale, and compare it with the measured value of the total organic carbon content of shale in the sample data of the total organic carbon content of shale, calculate the accuracy rate predicted by the trained SVR model and compare it with the preset precision value. If the accuracy rate predicted by the trained SVR model is lower than the preset precision value, return to Step 4 to continue training the SVR model. Otherwise, obtain the verified SVR model and directly use the verified SVR model to predict the total organic carbon content of shale.

[0088] Embodiment 2

[0089] In this embodiment, the traditional empirical method based on the △LogR model, the grey wolf optimization algorithm - support vector machine algorithm GWO - SVR, and the intelligent prediction method for the total organic carbon content of shale based on empirical constraints proposed in Embodiment 1 are respectively used to predict the total organic carbon content of shale in Well Y, and the prediction results of the total organic carbon content of shale in Well Y are obtained respectively, as Figure 3 shown.

[0090] Figure 3 In the figure, the TOClogR curve is the curve of the total organic carbon content of shale predicted by the traditional empirical method, the GWO - SVR curve is the curve of the total organic carbon content of shale predicted by the grey wolf optimization algorithm - support vector machine algorithm, and the EC - GWO - SVR curve is the curve of the total organic carbon content of shale predicted by the intelligent prediction method for the total organic carbon content of shale based on empirical constraints proposed in the present invention.

[0091] Through comparative analysis, it can be seen that the method of the present invention optimizes the support vector regression model based on the experience-constrained gray wolf algorithm. By introducing the △LogR model based on experience constraints into the gray wolf algorithm to optimize the support vector regression model, combining experience constraints with machine learning based on the existing experimental data and conventional logging curves in the study area, accurate prediction of the total organic carbon content of shale in the whole well section to be logged is achieved, which is conducive to guiding the exploration and development of shale reservoirs.

[0092] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by those skilled in the art within the substantial scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. An intelligent prediction method for the total organic carbon content of shale based on empirical constraints, characterized in that, It includes the following steps: Step 1: Obtain the logging data of the study area, reconstruct the wave impedance curve, select the logging curves related to the total organic carbon content, and obtain multiple shale total organic carbon content sample data; Step 2: Construct a sample data set and randomly assign the shale total organic carbon content sample data to the training set and test set of the sample data set; Step 3: Based on the DEN-GR corrected △LogR model, determine the total organic carbon content empirical model of the study area; Step 4: Optimize the SVR model based on the GWO algorithm, introduce the total organic carbon content empirical model into the gray wolf optimization process, combine with the shale total organic carbon content sample data in the training set, and use the GWO algorithm to determine the optimal model parameters of the SVR model to obtain the trained SVR model; Step 5: Input the shale total organic carbon content sample data in the test set into the trained SVR model, use the trained SVR model to predict the total organic carbon content of the shale total organic carbon content sample data, and verify the accuracy of the trained SVR model; In step 4, optimizing the SVR model based on the GWO algorithm includes the following sub-steps: Step 4.1: Initialize the GWO algorithm, set the type of gray wolf individuals, the position of the wolf pack, and the fitness function in the GWO algorithm, and determine the number of search agents, the maximum number of optimizations, and the initial value of the search space of the GWO algorithm; Step 4.2: Set the initial values of the penalty coefficient C and the kernel function parameter gamma in the SVR model, and input the shale total organic carbon content sample data in the training set into the SVR model and the total organic carbon content empirical model; Step 4.3: Based on the GWO algorithm, search for the solutions found by each individual in the wolf pack, calculate the mean square error using the SVR model, and cooperate with the total organic carbon content empirical model to calculate the penalty term to determine the fitness function corresponding to each solution; Step 4.4: Compare the fitness function values of each solution, determine the best solution, update the position of the wolf pack according to the positions of the Alpha wolf, Beta wolf, and Delta wolf, re-search for the solutions found by each individual in the wolf pack based on the GWO algorithm, and update the coefficient vector and solution when the wolf pack surrounds the prey; Step 4.5: Repeat steps 4.3 and 4.4 until the preset maximum number of optimizations is reached, stop searching for optimization using the GWO algorithm, output the optimal solutions of the penalty coefficient C and the kernel function parameter gamma in the SVR model, and obtain the trained SVR model.

2. The intelligent prediction method for the total organic carbon content of shale based on empirical constraints according to claim 1, wherein, In step 1, according to the acoustic travel time curve and the density curve, use the ratio between the acoustic travel time at different depths and the formation density to reconstruct the wave impedance curve; The logging data of the study area are the conventional logging curves that have been logged in the study area, including the acoustic travel time curve, density curve, neutron porosity curve, natural gamma curve, and resistivity curve. According to the acoustic travel time curve and the density curve, use the ratio between the acoustic travel time at different depths and the formation density to reconstruct the wave impedance curve; Based on the total organic carbon content curve obtained from well logging, select the well logging curves that can vary with the total organic carbon content from the conventional well logging curves, and combine with the acoustic impedance curve to obtain the well logging curves related to the total organic carbon content. Based on the well logging curves related to the total organic carbon content and the total organic carbon content curve at different depths, obtain multiple shale total organic carbon content sample data.

3. The intelligent prediction method for the total organic carbon content of shale based on empirical constraints according to claim 2, wherein The shale total organic carbon content sample data consists of multiple characteristic parameter values and the measured value of the shale total organic carbon content at the same depth. Among them, the characteristic parameter values are determined according to the well logging curves related to the total organic carbon content, and the measured value of the shale total organic carbon content is determined according to the total organic carbon content curve.

4. The intelligent prediction method for the total organic carbon content of shale based on empirical constraints according to claim 1, wherein In step 3, according to the well logging curves in the study area, based on the DEN-GR corrected △LogR model, combine with the fitting of the well logging curves in the study area to determine the model parameters, and determine the total organic carbon content empirical model as: TOC = [a×lg(GR)+b×DEN+c]×ΔlogR Where The △LogR model based on DEN-GR correction is: Where TOC is the total organic carbon content; a, b, and c are all model parameters; GR is the natural gamma value; DEN is the formation density value; R t is the resistivity, with the unit of Ω·m; R baseline is the resistivity corresponding to the baseline in the resistivity curve, with the unit of Ω·m; Δt is the acoustic time difference, with the unit of μs / m; Δt baseline is the acoustic time difference corresponding to the baseline in the acoustic time difference curve, with the unit of μs / m.

5. The intelligent prediction method for total organic carbon content of shale based on empirical constraints according to claim 1, characterized in that, The fitness function is: In the formula, Fitness is the fitness; w is the input sample; MSE is the mean square error; Penalty is the penalty term based on the total organic carbon content empirical model, which is used to ensure the data scale balance between the mean square error and the penalty term; λ is the penalty factor, which is used to balance the importance of the mean square error MSE and the penalty term; N is the total number of sample data of the total organic carbon content of shale in the training set; i is the sample serial number; y i is the measured TOC value of the i-th shale total organic carbon content sample data; is the predicted TOC value of the i-th shale total organic carbon content sample data; j is the serial number of the characteristic parameter; M is the total number of characteristic parameters; x j is the value of the j-th characteristic parameter; is the partial derivative of the predicted TOC value with respect to the characteristic parameter value x j and is used to represent the change rate of the predicted TOC value with respect to the j-th characteristic parameter; is the partial derivative of the TOC empirical value calculated by the total organic carbon content empirical model with respect to the j-th characteristic parameter, and is used to represent the change rate of the TOC empirical value with respect to the j-th characteristic parameter.

6. The intelligent prediction method for the total organic carbon content of shale based on empirical constraints according to claim 1, wherein In the GWO algorithm, the individuals in the wolf pack are divided into four categories according to the social rank, which are Alpha wolf, Beta wolf, Delta wolf, and Omega wolf from high to low social rank. Among them, the Alpha wolf is the leader of the wolf pack, ranking first in the wolf pack, and is responsible for making decisions and guiding the actions of the wolf pack; the Beta wolf ranks second in the wolf pack and is used to assist the Alpha wolf in making decisions and guiding the actions of other lower-rank wolves; the Delta wolf ranks third in the wolf pack and is responsible for scouting and alerting; the Omega wolf ranks fourth in the wolf pack and is used to execute the orders of the Alpha wolf, Beta wolf, and Delta wolf; When performing optimization search based on the GWO algorithm, the solution found by the Alpha wolf is the best solution, the solution found by the Beta wolf is the second-best solution, the solution found by the Delta wolf is the third-best solution, and the solution found by the Omega wolf is the remaining alternative solutions; The collective hunting process of the wolf pack adopts alternating global search and local search, including four stages, namely surrounding the prey, hunting, attacking the prey, and searching for the prey; The expression for surrounding the prey is: where \(t\) is the current optimization iteration number; is the distance vector between the grey wolf and the prey; is the distance vector between the grey wolf and the prey in the next optimization; are both coefficient vectors; is the position vector of the prey; is the position vector of the grey wolf; represents a linear decrease from 2 to 0 during the optimization process; are both random vectors in \([0, 1]\); The expression for hunting is: In the formula, are the distance vectors of Alpha wolf, Beta wolf, and Delta wolf respectively; is the coefficient vector of Alpha wolf, is the coefficient vector of Beta wolf, is the coefficient vector of Delta wolf, All are used to adjust the distance of individuals in the wolf pack; are the current positions of Alpha wolf, Beta wolf, and Delta wolf respectively; is the new potential position updated by Omega wolf according to the position and distance of Alpha wolf, is the new potential position updated by Omega wolf according to the position and distance of Beta wolf, is the new potential position updated by Omega wolf according to the position and distance of Delta wolf; is the coefficient vector determined by Omega wolf according to the position change of Alpha wolf, is the coefficient vector determined by Omega wolf according to the position change of Beta wolf, is the coefficient vector determined by Omega wolf according to the position change of Delta wolf; is the final position of Omega wolf. When , Omega wolf tends to deviate from the prey; when , Omega wolf tends to approach the prey.

7. The intelligent prediction method for the total organic carbon content of shale based on empirical constraints according to claim 1, characterized in that The SVR model selects the radial basis function as the kernel function to handle complex nonlinear problems; the model parameters of the SVR model include the penalty coefficient C and the kernel function parameter gamma. Among them, the penalty coefficient C determines the penalty strength of the SVR model for misclassification and is used to control the tolerance of the SVR model to errors; the kernel function parameter gamma determines the radiation range of the SVR model and is used to control the influence range of a single sample data.

8. The intelligent prediction method for total organic carbon content of shale based on empirical constraints according to claim 1, characterized in that In the said step 5, input the sample data of the total organic carbon content of shale in the test set into the trained SVR model. Use the trained SVR model to predict the predicted value of the total organic carbon content of shale based on the sample data of the total organic carbon content of shale, and compare it with the measured value of the total organic carbon content of shale in the sample data of the total organic carbon content of shale. Calculate the prediction accuracy of the trained SVR model and compare it with the preset precision value. If the prediction accuracy of the trained SVR model is lower than the preset precision value, then return to step 4 to continue training the SVR model. Otherwise, obtain the verified SVR model and directly use the verified SVR model to predict the total organic carbon content of shale.