Shale TOC content prediction method, electronic equipment and medium
The construction of a TOC content prediction model for terrestrial shale through the BP network optimized by genetic algorithms solves the problem of insufficient prediction accuracy of TOC content in the existing technology, and achieves higher prediction accuracy and nonlinear relationship modeling of the model.
Patent Information
- Application Number
- CN202311505595.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to accurately predict the TOC content of terrestrial shales, especially because of the strong heterogeneity of terrestrial shales and high clay mineral content, which leads to a low correlation between the TOC content and a single elastic parameter, making it difficult to obtain reliable TOC content through linear fitting.
A BP network optimized by genetic algorithm is used to construct a TOC prediction model through laboratory core test data, and the model is applied to well logging data. Using the nonlinear mapping and adaptive learning ability of the BP network, a model that can effectively predict the TOC content of terrestrial shale is established.
This method can effectively reduce the error of TOC content prediction, improve the prediction accuracy, solve the problem of insufficient linear fit prediction accuracy, and establish a TOC prediction model without the need for parameters such as diagenesis and geological factors.
Smart Images

Figure CN119990380A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of geophysical technology, and more specifically, to a shale TOC content prediction method, electronic equipment and medium. Background Art
[0002] TOC content is one of the important parameters for predicting sweet spots in shale reservoirs. Currently, prediction of TOC content in marine shale reservoirs is mostly achieved by establishing a linear correlation between TOC content and density. Compared with marine shales, the correlation between TOC content and a single elastic parameter (density, P-wave velocity, S-wave velocity, elastic modulus, etc.) is often low due to the strong heterogeneity of continental shale development layers, high clay mineral content, large variation in debris content, long sedimentation time span, and large differences in thermal evolution. Therefore, it is difficult to directly obtain reliable TOC content from elastic parameter linear fitting.
[0003] Therefore, it is necessary to develop a shale TOC content prediction method, electronic equipment and medium.
[0004] The information disclosed in the background technology section of the present invention is only intended to deepen the understanding of the general background technology of the present invention, and should not be regarded as acknowledging or suggesting in any form that the information constitutes the prior art already known to those skilled in the art. Summary of the invention
[0005] The present invention proposes a shale TOC content prediction method, electronic device and medium. Starting from laboratory core test data, a BP network optimized by genetic algorithm is used to predict the TOC content of continental shale. The TOC prediction model is built by using the good global search ability of genetic algorithm and the nonlinear mapping and adaptive learning ability of BP network, and then the model is applied to the logging data. The method of this patent has a solid experimental basis for predicting the TOC content of continental shale, and can make full use of the TOC content information contained in different elastic parameters, solving the problem of insufficient accuracy in predicting TOC content using linear fitting.
[0006] In a first aspect, the present disclosure provides a method for predicting shale TOC content, comprising:
[0007] Obtain laboratory core test data and well logging data for continental shale target layers;
[0008] Performing initialization processing on the test data and the well logging data to obtain initialized test data and initialized well logging data;
[0009] Dividing the initialized test data into a training set and a test set;
[0010] Performing model training using the training set and testing using the test set to obtain a trained model;
[0011] The initialized logging data is brought into the trained model to predict the TOC content of continental shale logging.
[0012] Preferably, the initialization process includes consistency check and normalization process.
[0013] Preferably, performing model training using the training set includes:
[0014] Determine the topological structure of the BP network, use the logging data as the input of the BP network, use the TOC content as the output, and use the BP network with one hidden layer to establish a prediction model;
[0015] Determine the range of the number of neurons in the hidden layer according to the number of neurons in the input layer, select different numbers of neurons in the hidden layer within the range for prediction, compare the prediction results with the measured data, and determine the optimal number of neurons in the hidden layer by calculating the root mean square error between the two;
[0016] Initialize the prediction model, assign a random value within (-1, 1) to each weight threshold, encode the weight and threshold of the prediction model by real number encoding, and initialize the population;
[0017] Initialize the genetic algorithm, and perform chromosome encoding on the weights and thresholds of the prediction model. During encoding, the length of the chromosome gene is the sum of all weights and thresholds in the network;
[0018] The inverse square root of the error obtained by training the BP network according to the training set is used as the fitness value of the genetic algorithm, and the smaller the error of the individual, the greater the fitness;
[0019] According to the rule of survival of the fittest, excellent individuals are selected from the current population as parents to produce the next generation of individuals;
[0020] When two parent individuals undergo a crossover operation, the gene chain codes are exchanged at the crossover operation point, thus forming two new individuals;
[0021] Randomly select an individual from the population and obtain a new individual through mutation;
[0022] It is determined whether the fitness value satisfies the termination condition. If so, the optimal weight and threshold are obtained. If not, new individuals are regenerated to obtain a new population, and the individual fitness values of the new population are calculated until the termination condition is met.
[0023] Preferably, the chromosome code string Y is in the form of:
[0024] Y=(w 11 ,w 12 ,…,w ij ,…,w mn ,w1,w2,…,w j ,…,w n ,b1,b2,…,b j ,…,b n ,b)
[0025] In the formula, w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, w j is the weight between the jth neuron in the hidden layer and the neuron in the output layer, b j is the threshold of the jth neuron in the hidden layer, b is the threshold of the neuron in the output layer, m is the number of neurons in the input layer, n is the number of neurons in the hidden layer, i = 1, 2, ..., m, j = 1, 2, ..., n;
[0026] The chromosome length l is:
[0027] l=m×n+n×1+n+1
[0028] Where m is the number of neurons in the input layer, and n is the number of neurons in the hidden layer.
[0029] Preferably, when two parent individuals perform a crossover operation, the gene chain codes are exchanged at the crossover operation point, thereby forming two new individuals including:
[0030] Suppose the two parent individuals are X=(x1,…,x i ,…,x l ) and Y=(y1,…,y i ,…,y l ), then the two offspring individuals X ′ =(x1 ′ ,…,x i ′ ,..,x l ′ ) and Y ′ =(y1 ′ ,…,y i ′ ,..,y l ′ )for:
[0031]
[0032] Among them, r is a random number.
[0033] Preferably, randomly selecting an individual from the population and obtaining a new individual through mutation includes:
[0034] Let the individual be X=(x1,…,x i ,…,x l ), and x i ∈[a i , b i ], then after mutation, the individual gene x i ′ for:
[0035]
[0036]
[0037] In the formula, a i , b i are the upper and lower limits of each variable, G, G max are the current population size and the maximum population size, r1 and r2 are random numbers between 0 and 1, and b is a parameter related to the number of iterations.
[0038] Preferably, the test set is tested by:
[0039] Assigning the optimal weight and threshold to the BP network, and calculating the error through the test set;
[0040] According to the gradient descent algorithm, the adjustment amount of the weight is proportional to the gradient descent of the error to obtain the adjusted weight;
[0041] Determine whether the error meets the termination condition. If so, output the trained model. If not, repeat the above steps.
[0042] Preferably, the adjustment amount of the weight is:
[0043]
[0044]
[0045] Among them, Δw ij and Δw j are the adjustment values of the neuron weights of the input layer and hidden layer, and hidden layer and output layer, respectively; η is the learning rate; e is the output error of the BP network; and w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, w j is the weight between the jth neuron in the hidden layer and the neuron in the output layer.
[0046] In a second aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:
[0047] A memory storing executable instructions;
[0048] A processor runs the executable instructions in the memory to implement the shale TOC content prediction method.
[0049] In a third aspect, the embodiments of the present disclosure further provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the shale TOC content prediction method is implemented.
[0050] Its beneficial effects are:
[0051] (1) In the process of predicting TOC content in continental shale, the BP network TOC content prediction method based on genetic algorithm optimization can effectively reduce the prediction error. The powerful global optimization ability of the genetic algorithm avoids the local optimal problem and greatly improves the prediction ability of the BP network, thereby making the prediction of TOC content more accurate.
[0052] (2) Since the BP network is used to simulate the complex nonlinear relationship between density, longitudinal wave velocity, and shear wave velocity, a prediction model for TOC content in continental shale can be established without the need for parameters such as diagenesis and geological factors, and the prediction accuracy can meet user needs.
[0053] The methods and apparatus of the present invention have other features and advantages that will be apparent from, or will be described in detail in, the accompanying drawings and subsequent detailed descriptions incorporated herein, which together serve to explain the specific principles of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The above and other objects, features and advantages of the present invention will become more apparent through a more detailed description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present invention.
[0055] Figure 1 A flow chart of a BP network optimized by a genetic algorithm according to an embodiment of the present invention is shown.
[0056] Figure 2 A flow chart showing the steps of a method for predicting shale TOC content according to an embodiment of the present invention.
[0057] Figure 3 A schematic diagram showing consistency check between continental shale core test data and well logging data according to an embodiment of the present invention is shown.
[0058] Figure 4 A schematic diagram showing the comparison between the model predicted TOC and the test set TOC according to an embodiment of the present invention is shown.
[0059] Figure 5 A schematic diagram showing the prediction results of TOC content in continental shale well logging according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0060] The preferred embodiments of the present invention will be described in more detail below. Although the preferred embodiments of the present invention are described below, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein.
[0061] The present invention provides a method for predicting shale TOC content, comprising:
[0062] Obtain laboratory core test data and well logging data for continental shale target layers;
[0063] Performing initialization processing on the test data and the logging data to obtain initialized test data and initialized logging data;
[0064] Divide the initialized test data into training set and test set;
[0065] Train the model through the training set and test it through the test set to obtain the trained model;
[0066] The initialized logging data is brought into the trained model to predict the TOC content of continental shale logging.
[0067] In one example, the initialization process includes a consistency check and a normalization process.
[0068] In one example, training a model using a training set includes:
[0069] Determine the topological structure of the BP network, use the logging data as the input of the BP network, use the TOC content as the output, and use the BP network with one hidden layer to establish a prediction model;
[0070] Determine the range of the number of neurons in the hidden layer according to the number of neurons in the input layer, select different numbers of neurons in the hidden layer within the range for prediction, compare the prediction results with the measured data, and determine the optimal number of neurons in the hidden layer by calculating the root mean square error between the two;
[0071] Initialize the prediction model, assign random values within (-1,1) to each weight threshold, encode the weights and thresholds of the prediction model through real number encoding, and initialize the population;
[0072] Initialize the genetic algorithm and encode the weights and thresholds of the prediction model into chromosomes. During encoding, the length of the chromosome gene is the sum of all weights and thresholds in the network.
[0073] The inverse square root of the error obtained by BP network training according to the training set is used as the fitness value of the genetic algorithm. The smaller the individual error, the greater the fitness.
[0074] According to the rule of survival of the fittest, excellent individuals are selected from the current population as parents to produce the next generation of individuals;
[0075] When two parent individuals undergo a crossover operation, the gene chain codes are exchanged at the crossover operation point, thus forming two new individuals;
[0076] Randomly select an individual from the population and obtain a new individual through mutation;
[0077] Determine whether the fitness value meets the termination condition. If so, obtain the optimal weight and threshold. If not, regenerate new individuals to obtain a new population, and calculate the individual fitness value of the new population until the termination condition is met.
[0078] In one example, the chromosome encoding string Y is in the form:
[0079] Y=(w 11 ,w 12 ,...,w ij ,…,w mn ,w1,w2,…,w j ,…,w n ,b1,b2,…,b j ,...,b n ,b)
[0080] In the formula, w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, w j is the weight between the jth neuron in the hidden layer and the neuron in the output layer, b j is the threshold of the jth neuron in the hidden layer, b is the threshold of the neuron in the output layer, m is the number of neurons in the input layer, n is the number of neurons in the hidden layer, i = 1, 2, ..., m, j = 1, 2, ..., n;
[0081] The chromosome length l is:
[0082] l=m×n+n×1+n+1
[0083] Where m is the number of neurons in the input layer, and n is the number of neurons in the hidden layer.
[0084] In one example, when two parent individuals undergo a crossover operation, the gene chain codes are exchanged at the crossover operation point, thereby forming two new individuals including:
[0085] Suppose the two parent individuals are X=(x1,…,xi ,…,x l ) and Y=(y1,...,y i ,...,y l ), then the two offspring individuals X′=(x′1,...,x′ i ,..,x′ l ) and Y′=(y′1,…,y′ i ,..,y′ l )for:
[0086]
[0087] Among them, r is a random number.
[0088] In one example, an individual is randomly selected from the population, and a new individual is obtained by mutation including:
[0089] Let the individual be X=(x1,…,x i ,…,x l ), and x i ∈[a i , b i ], then after mutation, the individual gene x i ′ for:
[0090]
[0091]
[0092] In the formula, a i , b i are the upper and lower limits of each variable, G, G max are the current population size and the maximum population size, r1 and r2 are random numbers between 0 and 1, and b is a parameter related to the number of iterations.
[0093] In an example, testing with the test set you see:
[0094] Assign the optimal weights and thresholds to the BP network and calculate the error using the test set;
[0095] According to the gradient descent algorithm, the adjustment amount of the weight is proportional to the gradient descent of the error to obtain the adjusted weight;
[0096] Determine whether the error meets the termination condition. If so, output the trained model. If not, repeat the above steps.
[0097] In one example, the weights are adjusted by:
[0098]
[0099]
[0100] Among them, Δw ij and Δw j are the adjustment values of the neuron weights of the input layer and hidden layer, and hidden layer and output layer, respectively; η is the learning rate; e is the output error of the BP network; and w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, w j is the weight between the jth neuron in the hidden layer and the neuron in the output layer.
[0101] Specifically, the density, P-wave velocity and S-wave velocity measured in the core laboratory of the target strata of continental shale were used to establish a BP network continental shale TOC content prediction model optimized by genetic algorithm, and then the logging data of density, P-wave velocity and S-wave velocity were substituted to obtain the predicted logging TOC content.
[0102] Obtain laboratory core test data and logging data of the continental shale target layer. The core data include TOC content, density, P-wave velocity, and S-wave velocity. The logging data include density, P-wave velocity, and S-wave velocity. Check the consistency of core elastic parameters and logging elastic parameters at the same depth to ensure that the model obtained from core data can be reasonably applied to logging data. At the same depth, core data and logging data are very close. Data initialization: Convert core test data and logging data of different units and dimensions into dimensionless data within the same range.
[0103] Divide the training set and the test set: Use the initialized core test data as the sample set, and divide the sample set into the training set and the test set in a ratio of approximately 8:2 according to the number of existing core test data.
[0104] Figure 1 A flow chart of a BP network optimized by a genetic algorithm according to an embodiment of the present invention is shown.
[0105] like Figure 1 As shown, the model is established by BP network optimized by genetic algorithm.
[0106] Model training: Given the initial BP network weight threshold, use the training set to obtain the optimal BP network weight threshold based on genetic algorithm optimization, including the following steps:
[0107] The topological structure of the BP network was determined, density, longitudinal wave velocity, and shear wave velocity were selected as the input of the BP network, TOC content was selected as the output of the BP network, and a prediction model was established using a BP network with one hidden layer. The number of neurons in the hidden layer and the number of neurons in the input layer followed the following relationship:
[0108] n≤2m+1
[0109] Where m is the number of neurons in the input layer, and n is the number of neurons in the hidden layer. According to the number of neurons in the input layer m=3, the upper limit of the number of neurons in the hidden layer n is determined to be 7. Different numbers of neurons in the hidden layer are selected within this range for prediction. The prediction results are compared with the measured data. The root mean square error between the two is calculated to determine the optimal number of neurons in the hidden layer n=5.
[0110] The root mean square error formula is:
[0111]
[0112] In the formula, n is the number of data, T p is the predicted value of TOC content, T e It is the measured value of TOC content.
[0113] Initialize the BP network, assign random values within (-1,1) to each weight threshold, and use real number coding to encode the network weights and thresholds to initialize the population.
[0114] The genetic algorithm is initialized to encode the weights and thresholds of the initial BP network. During encoding, the length of the chromosome gene is the sum of all weights and thresholds in the network.
[0115] The chromosome code string Y is in the form of:
[0116] Y=(w 11 ,w 12 ,...,w ij ,…,w mn ,w1,w2,…,w j ,...,w n ,b1,b2,...,b j ,…,b n ,b)
[0117] Where w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, w j is the weight between the jth neuron in the hidden layer and the neuron in the output layer, b j is the threshold of the jth neuron in the hidden layer, b is the threshold of the neuron in the output layer, m is the number of neurons in the input layer, n is the number of neurons in the hidden layer, i = 1, 2, ..., m, j = 1, 2, ..., n.
[0118] The chromosome length l is:
[0119] l=m×n+n×1+n+1
[0120] Where m is the number of neurons in the input layer, and n is the number of neurons in the hidden layer.
[0121] According to the training set data, the inverse square root of the error obtained by BP network training is used as the fitness value of the genetic algorithm. The smaller the individual error, the greater the fitness. The fitness function is:
[0122]
[0123] Where F is the fitness, T is the predicted TOC content value, is the expected value of TOC content.
[0124] By calculating the fitness of all individuals in the population and selecting the better individuals from the current population as parents to produce the next generation of individuals according to the rule of survival of the fittest, the probability of each individual being selected is determined by the following method:
[0125]
[0126] Where p k is the probability of the kth individual being selected, F k is the fitness of the kth individual, and N is the total number of individuals in the population.
[0127] Perform selection, crossover and mutation operations. When two parent individuals perform a crossover operation, the gene chain codes are exchanged at the crossover operation point to form two new individuals.
[0128] Assume that the two parent individuals are X=(x1,...,x i ,...,x l ) and Y=(y1,...,y i ,…,y l ), then the two offspring individuals X′=(x′1,…,x′ i ,..,x′ l ) and Y′=(y′1,…,y′ i ,..,y′ l ) can be expressed as:
[0129]
[0130] Where r is a random number.
[0131] An individual is randomly selected from the population and mutated to obtain a new individual with a certain probability.
[0132] Assume that an individual is X=(x1,...,x i ,…,x l ), and x i ∈[a i , b i], then the individual gene x′ after mutation i for:
[0133]
[0134]
[0135] Where a i , b i are the upper and lower limits of each variable, G, G max are the current population size and the maximum population size, r1 and r2 are random numbers between 0 and 1, and b is a parameter related to the number of iterations.
[0136] Determine whether the fitness value meets the termination conditions. If so, obtain the optimal weight and threshold. If not, repeatedly perform selection, crossover, and mutation operations to generate a new population, and calculate the individual fitness of the new population to find the optimal individual. The larger the fitness value, the better the individual.
[0137] Model evaluation: Use the test set data to calculate the prediction error until the accuracy meets the end condition, and output the trained model, including:
[0138] The calculated optimal weights and thresholds are assigned to the BP network, and the test error is calculated using the test set data.
[0139] Assume x i is the i-th input signal of the BP network input layer, w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, b j is the threshold of the jth neuron in the hidden layer, then the total input x of the jth neuron in the hidden layer is j for:
[0140]
[0141] Where m is the number of neurons in the input layer.
[0142] According to the sigmoid function tansig, the output y of the jth neuron in the hidden layer can be obtained j for:
[0143]
[0144] Where x j is the total input of the jth neuron in the hidden layer.
[0145] Assume that w j is the weight between the jth neuron in the hidden layer and the neurons in the output layer, b is the threshold of the total output signal, then the total input signal x of the neurons in the output layer of the BP network is:
[0146]
[0147] Where y j is the output of the jth neuron in the hidden layer, and n is the number of neurons in the hidden layer.
[0148] According to the linear function purelin, the total output T of the BP network can be obtained as:
[0149] T=x
[0150] Where x is the total input signal of the neurons in the output layer of the BP network.
[0151] Assume that the expected output of the BP neural network is Then the error e is:
[0152]
[0153] Where T is the total output of BP network.
[0154] Update the weights and thresholds. According to the gradient descent algorithm, the adjustment of the weights is proportional to the gradient descent of the error, that is:
[0155]
[0156]
[0157] Δw ij and Δw j are the adjustment values of the neuron weights of the input layer and hidden layer, and hidden layer and output layer, respectively; η is the learning rate; e is the output error of the BP neural network; and w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, w j is the weight between the jth neuron in the hidden layer and the neuron in the output layer.
[0158] Then the adjusted weight w′ ij and w′ j for:
[0159] w′ ij =w ij +Δw ij
[0160] w′ j =w j +Δw j
[0161] Where w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, w j is the weight between the jth neuron in the hidden layer and the neuron in the output layer, Δwij and Δw j They are the adjustment amounts of the neuron weights between the input layer and the hidden layer, and between the hidden layer and the output layer, respectively.
[0162] Determine whether the test error meets the termination condition. If so, output the trained model. If not, repeat the above steps.
[0163] The initialized logging elastic parameters are used as input and substituted into the model to predict the TOC content of continental shale logging.
[0164] The present invention also provides an electronic device, which includes: a memory storing executable instructions; and a processor, which runs the executable instructions in the memory to implement the above-mentioned shale TOC content prediction method.
[0165] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned shale TOC content prediction method is implemented.
[0166] To facilitate understanding of the solutions and effects of the embodiments of the present invention, four specific application examples are given below. Those skilled in the art should understand that the examples are only for facilitating understanding of the present invention, and any specific details thereof are not intended to limit the present invention in any way.
[0167] Example 1
[0168] Figure 2 A flow chart showing the steps of a method for predicting shale TOC content according to an embodiment of the present invention.
[0169] like Figure 2 As shown, the shale TOC content prediction method includes: step 101, obtaining laboratory core test data and logging data of the continental shale target layer; step 102, initializing the test data and logging data to obtain initialized test data and initialized logging data; step 103, dividing the initialized test data into a training set and a test set; step 104, training the model through the training set and testing through the test set to obtain a trained model; step 105, bringing the initialized logging data into the trained model to predict the continental shale logging TOC content.
[0170] The density, P-wave velocity and S-wave velocity measured in the core laboratory of the target layer of continental shale were used to establish a BP network TOC content prediction model for continental shale optimized by genetic algorithm. Then the logging data of density, P-wave velocity and S-wave velocity were substituted to obtain the predicted logging TOC content.
[0171] Figure 3A schematic diagram showing consistency check between continental shale core test data and well logging data according to an embodiment of the present invention is shown.
[0172] Obtain laboratory core test data and well logging data of the continental shale target layer in a certain area of the Sichuan Basin. The core data includes TOC content, density, P-wave velocity, and S-wave velocity. The well logging data includes density, P-wave velocity, and S-wave velocity. Check the consistency of the core elastic parameters and the well logging elastic parameters at the same depth to ensure that the model obtained from the core data can be reasonably applied to the well logging data, such as Figure 3 As shown, at the same depth, the core data and the logging data are very close; data initialization: the core test data and logging data of different units and dimensions are converted into dimensionless data within the same range, among which the TOC content data unit is %, the density data unit is g / cc, and the P-wave velocity data and S-wave velocity data unit is km / s.
[0173] Divide the training set and the test set: Use the initialized core test data as the sample set, and divide the sample set into the training set and the test set in a ratio of approximately 8:2 according to the number of existing core test data.
[0174] Model training: Given the initial BP network weight threshold, use the training set to obtain the optimal BP network weight threshold based on genetic algorithm optimization, including the following steps:
[0175] The topological structure of the BP network was determined, density, longitudinal wave velocity, and shear wave velocity were selected as the input of the BP network, TOC content was selected as the output of the BP network, and a prediction model was established using a BP network with one hidden layer. The number of neurons in the hidden layer and the number of neurons in the input layer followed the following relationship:
[0176] n≤2m+1
[0177] Where m is the number of neurons in the input layer, and n is the number of neurons in the hidden layer. According to the number of neurons in the input layer m=3, the upper limit of the number of neurons in the hidden layer n is determined to be 7. Different numbers of neurons in the hidden layer are selected within this range for prediction. The prediction results are compared with the measured data. The root mean square error between the two is calculated to determine the optimal number of neurons in the hidden layer n=5.
[0178] The root mean square error formula is:
[0179]
[0180] In the formula, n is the number of data, T p is the predicted value of TOC content, T e It is the measured value of TOC content.
[0181] Initialize the BP network, assign random values within (-1,1) to each weight threshold, and use real number coding to encode the network weights and thresholds to initialize the population.
[0182] The genetic algorithm is initialized to encode the weights and thresholds of the initial BP network. During encoding, the length of the chromosome gene is the sum of all weights and thresholds in the network.
[0183] The chromosome code string Y is in the form of:
[0184] Y=(w 11 ,w 12 ,…,w ij ,…,w mn ,w1,w2,…,w j ,...,w n ,b1,b2,...,b j ,…,b n ,b)
[0185] Where w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, w j is the weight between the jth neuron in the hidden layer and the neuron in the output layer, b j is the threshold of the jth neuron in the hidden layer, b is the threshold of the neuron in the output layer, m is the number of neurons in the input layer, n is the number of neurons in the hidden layer, i = 1, 2, ..., m, j = 1, 2, ..., n.
[0186] The chromosome length l is:
[0187] l=m×n+n×1+n+1
[0188] Where m is the number of neurons in the input layer, and n is the number of neurons in the hidden layer.
[0189] According to the training set data, the inverse square root of the error obtained by BP network training is used as the fitness value of the genetic algorithm. The smaller the individual error, the greater the fitness. The fitness function is:
[0190]
[0191] Where F is the fitness, T is the predicted TOC content value, is the expected value of TOC content.
[0192] By calculating the fitness of all individuals in the population and selecting the better individuals from the current population as parents to produce the next generation of individuals according to the rule of survival of the fittest, the probability of each individual being selected is determined by the following method:
[0193]
[0194] Where p k is the probability of the kth individual being selected, F k is the fitness of the kth individual, and N is the total number of individuals in the population.
[0195] Perform selection, crossover and mutation operations. When two parent individuals perform a crossover operation, the gene chain codes are exchanged at the crossover operation point to form two new individuals.
[0196] Assume that the two parent individuals are X=(x1,…,x i ,…,x l ) and Y=(y1,…,y i ,…,y l ), then the two offspring individuals X ′ =(x1 ′ ,…,x i ′ ,..,x l ′ ) and Y ′ =(y1 ′ ,…,y i ′ ,..,y l ′ ) can be expressed as:
[0197]
[0198] Where r is a random number.
[0199] An individual is randomly selected from the population and mutated to obtain a new individual with a certain probability.
[0200] Assume that an individual is X=(x1,…,x i ,…,x l ), and x i ∈[a i , b i ], then after mutation, the individual gene x i ′ for:
[0201]
[0202]
[0203] Where a i , b i are the upper and lower limits of each variable, G, G max are the current population size and the maximum population size, r1 and r2 are random numbers between 0 and 1, and b is a parameter related to the number of iterations.
[0204] Determine whether the fitness value meets the termination conditions. If so, obtain the optimal weight and threshold. If not, repeatedly perform selection, crossover, and mutation operations to generate a new population, and calculate the individual fitness of the new population to find the optimal individual. The larger the fitness value, the better the individual.
[0205] Figure 4 A schematic diagram showing the comparison between the model predicted TOC and the test set TOC according to an embodiment of the present invention is shown.
[0206] Model evaluation: Use the test set data to calculate the prediction error until the accuracy meets the end condition, output the trained model, and compare the TOC content prediction results of the test set. Figure 4 As shown, including:
[0207] The calculated optimal weights and thresholds are assigned to the BP network, and the test error is calculated using the test set data.
[0208] Assume x i is the i-th input signal of the BP network input layer, w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, b j is the threshold of the jth neuron in the hidden layer, then the total input x of the jth neuron in the hidden layer is j for:
[0209]
[0210] Where m is the number of neurons in the input layer.
[0211] According to the sigmoid function tansig, the output y of the jth neuron in the hidden layer can be obtained j for:
[0212]
[0213] Where x j is the total input of the jth neuron in the hidden layer.
[0214] Assume that w j is the weight between the jth neuron in the hidden layer and the neurons in the output layer, b is the threshold of the total output signal, then the total input signal x of the neurons in the output layer of the BP network is:
[0215]
[0216] Where y j is the output of the jth neuron in the hidden layer, and n is the number of neurons in the hidden layer.
[0217] According to the linear function purelin, the total output T of the BP network can be obtained as:
[0218] T=x
[0219] Where x is the total input signal of the neurons in the output layer of the BP network.
[0220] Assume that the expected output of the BP neural network is Then the error e is:
[0221]
[0222] Where T is the total output of BP network.
[0223] Update the weights and thresholds. According to the gradient descent algorithm, the adjustment of the weights is proportional to the gradient descent of the error, that is:
[0224]
[0225]
[0226] Δw ij and Δw j are the adjustment values of the neuron weights of the input layer and hidden layer, and hidden layer and output layer, respectively; η is the learning rate; e is the output error of the BP neural network; and w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, w j is the weight between the jth neuron in the hidden layer and the neuron in the output layer.
[0227] Then the adjusted weight w′ ij and w′ j for:
[0228] w′ ij =w ij +Δw ij
[0229] w′ j =w j +Δw j
[0230] Where w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, w j is the weight between the jth neuron in the hidden layer and the neuron in the output layer, Δw ij and Δw j They are the adjustment amounts of the neuron weights between the input layer and the hidden layer, and between the hidden layer and the output layer, respectively.
[0231] Determine whether the test error meets the termination condition. If so, output the trained model. If not, repeat the above steps.
[0232] Figure 5 A schematic diagram showing the prediction results of TOC content in continental shale well logging according to an embodiment of the present invention is shown.
[0233] The initialized logging elastic parameters are used as input and substituted into the model to predict the TOC content of continental shale logging, such as Figure 5 shown.
[0234] Example 2
[0235] The present disclosure provides an electronic device, which includes: a memory storing executable instructions; and a processor running the executable instructions in the memory to implement the above-mentioned shale TOC content prediction method.
[0236] An electronic device according to an embodiment of the present disclosure includes a memory and a processor.
[0237] The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.
[0238] The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In one embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory.
[0239] Those skilled in the art should be able to understand that in order to solve the technical problem of how to obtain a good user experience, the present embodiment may also include well-known structures such as a communication bus and an interface, and these well-known structures should also be included in the protection scope of the present disclosure.
[0240] For detailed description of this embodiment, reference may be made to the corresponding descriptions in the aforementioned embodiments, which will not be repeated here.
[0241] Example 3
[0242] An embodiment of the present disclosure provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the shale TOC content prediction method is implemented.
[0243] According to the computer-readable storage medium of the embodiment of the present disclosure, non-transitory computer-readable instructions are stored thereon. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the above-mentioned methods of each embodiment of the present disclosure are executed.
[0244] The above-mentioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or mobile hard disk), media with built-in rewritable non-volatile memory (e.g., memory card) and media with built-in ROM (e.g., ROM box).
[0245] Those skilled in the art should understand that the purpose of the above description of the embodiments of the present invention is only to exemplarily illustrate the beneficial effects of the embodiments of the present invention, and is not intended to limit the embodiments of the present invention to any given examples.
[0246] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.
Claims
1. A method for predicting shale TOC content, characterized in that: include: Obtain laboratory core test data and well logging data for continental shale target layers; Performing initialization processing on the test data and the well logging data to obtain initialized test data and initialized well logging data; Dividing the initialized test data into a training set and a test set; Performing model training using the training set and testing using the test set to obtain a trained model; The initialized logging data is brought into the trained model to predict the TOC content of continental shale logging.
2. The method for predicting shale TOC content according to claim 1, wherein: The initialization process includes consistency checking and normalization processing.
3. The method for predicting shale TOC content according to claim 1, wherein: Model training using the training set includes: Determine the topological structure of the BP network, use the logging data as the input of the BP network, use the TOC content as the output, and use the BP network with one hidden layer to establish a prediction model; Determine the range of the number of neurons in the hidden layer according to the number of neurons in the input layer, select different numbers of neurons in the hidden layer within the range for prediction, compare the prediction results with the measured data, and determine the optimal number of neurons in the hidden layer by calculating the root mean square error between the two; Initialize the prediction model, assign a random value within (-1, 1) to each weight threshold, encode the weight and threshold of the prediction model by real number encoding, and initialize the population; Initialize the genetic algorithm, and perform chromosome encoding on the weights and thresholds of the prediction model. During encoding, the length of the chromosome gene is the sum of all weights and thresholds in the network; The inverse square root of the error obtained by training the BP network according to the training set is used as the fitness value of the genetic algorithm, and the smaller the error of the individual, the greater the fitness; According to the rule of survival of the fittest, excellent individuals are selected from the current population as parents to produce the next generation of individuals; When two parent individuals undergo a crossover operation, the gene chain codes are exchanged at the crossover operation point, thus forming two new individuals; Randomly select an individual from the population and obtain a new individual through mutation; It is determined whether the fitness value satisfies the termination condition. If so, the optimal weight and threshold are obtained. If not, new individuals are regenerated to obtain a new population, and the individual fitness values of the new population are calculated until the termination condition is met.
4. The method for predicting shale TOC content according to claim 3, wherein: The chromosome code string Y is in the form of: Y=(w 11 ,w 12 ,...,w ij ,...w mn ,w1,w2,...,w j ,...,w n ,b1,b2,...,b j ,...,b n ,b) In the formula, w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, w j is the weight between the jth neuron in the hidden layer and the neuron in the output layer, b j is the threshold of the jth neuron in the hidden layer, b is the threshold of the neuron in the output layer, m is the number of neurons in the input layer, n is the number of neurons in the hidden layer, i = 1, 2, ..., m, j = 1, 2, ..., n; The chromosome length l is: l=m×n+n×1+m+1 Where m is the number of neurons in the input layer, and n is the number of neurons in the hidden layer.
5. The method for predicting shale TOC content according to claim 3, wherein: When two parent individuals undergo a crossover operation, the gene chain codes are exchanged at the crossover operation point, thereby forming two new individuals including: Suppose the two parent individuals are X=(x1,…,x i ,…,x l ) and Y=(y1,…,y i ,…,y l ), then the two offspring individuals X′=(x1′,…,x i ′,..,x l ′) and Y′=(y1′,…,y i ′,..,y l ')for: Among them, r is a random number.
6. The method for predicting shale TOC content according to claim 3, wherein: Randomly select an individual from the population, and obtain new individuals through mutation, including: Let the individual be X=(x1,…,x i ,…,x l ), and x i ∈[a i , b i ], then after mutation, the individual gene x i 'for: In the formula, a i 、b i are the upper and lower limits of each variable, G, G max are the current population size and the maximum population size, r1 and r2 are random numbers between 0 and 1, and b is a parameter related to the number of iterations.
7. The method for predicting shale TOC content according to claim 3, wherein: Test it with the test set and you will see: Assigning the optimal weight and threshold to the BP network, and calculating the error through the test set; According to the gradient descent algorithm, the adjustment amount of the weight is proportional to the gradient descent of the error to obtain the adjusted weight; Determine whether the error meets the termination condition. If so, output the trained model. If not, repeat the above steps.
8. The method for predicting shale TOC content according to claim 7, wherein: The weight adjustment is: Among them, Δw ij and Δw j are the adjustment values of the neuron weights of the input layer and hidden layer, and hidden layer and output layer, respectively; η is the learning rate; e is the output error of the BP network; and w ij is the weight between the i-th input signal in the input layer and the j-th neuron in the hidden layer, w j is the weight between the jth neuron in the hidden layer and the neuron in the output layer.
9. An electronic device, characterized in that: The electronic device comprises: A memory storing executable instructions; A processor, wherein the processor runs the executable instructions in the memory to implement the shale TOC content prediction method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the shale TOC content prediction method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Identification method of total content of organic carbon in shale
CN106168677A
Method for predicting total organic carbon (TOC) content of shale
CN106568918A
Shale gas total organic carbon prediction method based on deep learning and interpolation regression
CN114943060A
Method for evaluating organic carbon content of shale oil reservoir based on genetic optimization neural network algorithm
CN115032361A
Fly ash carbon content prediction method based on genetic simulated annealing parameter optimization
CN115809744A