A method, system and medium for predicting TOC of source rocks

By using a prediction model of the convolutional neural network-bidirectional gating cyclic unit-attention mechanism algorithm in the TOC prediction of source rocks, combined with multiple logging data and actual measured data, the problem of insufficient prediction accuracy in the existing technology is solved, and higher prediction accuracy and more effective formation information mining are achieved.

CN119395766BActive Publication Date: 2025-05-16SOUTHWEST PETROLEUM UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411493977.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2025-05-16
Estimated Expiration
2044-10-24

AI Technical Summary

Technical Problem

The existing TOC prediction methods for source rocks are difficult to meet the exploration needs in terms of accuracy, especially in areas with low exploration levels, and prediction accuracy is difficult to ensure.

Method used

A prediction model based on the convolutional neural network-bidirectional gating cyclic unit-attention mechanism algorithm is adopted, and combined with the recombinant multiple logging data and the TOC data obtained by partially measured TOC data, the overall source rock TOC content in the study area is predicted.

Benefits of technology

It significantly improves the accuracy of TOC prediction of source rocks, and can more effectively dig stratigraphic information to meet exploration needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119395766B_ABST
    Figure CN119395766B_ABST
Patent Text Reader

Abstract

The present invention provides a prediction method, system and medium for source rock TOC, which relates to the technical field of TOC prediction, including: obtaining the measured total organic carbon data and multiple logging data obtained from the core sample experiment of the source rock in the sampling well in the study area, and performing normalization and reorganization to obtain the preliminary logging curve value; judging the processed preliminary logging curve value according to the data fuzzy region model to obtain the effective data prediction interval, and removing the abnormal values ​​of the data related to the prediction result; constructing a first prediction model based on the convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm; inputting the measured total organic carbon data obtained in the experiment and the final logging-related data into the first prediction model after training, and obtaining the predicted total organic carbon data of the unexperimented part of the sampling well. The method uses the convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm to construct a prediction model to predict the TOC content of the entire source rock in the study area, and the prediction accuracy is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of TOC prediction, and in particular to a method, system and medium for predicting TOC of a hydrocarbon source rock. Background Art

[0002] The size of total organic carbon (TOC) is often related to whether shale oil and gas is high-yield and enriched. The prediction of TOC content has become one of the important indicators that must be predicted in the field of unconventional oil and gas exploration such as shale gas and shale oil. Due to the strong heterogeneity of shale TOC content and its rapid vertical and horizontal changes, it is difficult to use ordinary prediction algorithms to accurately predict the TOC content of source rocks.

[0003] At present, the commonly used prediction methods for source rock TOC mainly include autonomous machine learning algorithms, virtual well construction prediction, neural network algorithms and seismic information quantitative prediction algorithms. Among them, the autonomous machine learning algorithm (a shale TOC content prediction method and device based on machine learning, patent number: CN202211423221.0) requires a large amount of well-seismic combined data, the sampling data sample requirements are too large, the data precision is too high, and the existing data in some unsampled well areas are difficult to meet the requirements; virtual well construction prediction (a method for predicting source rock TOC in sparse well areas, patent number: CN202211286246.0) has harsh preset geological conditions, large parameter requirements, and cumbersome and complicated calculation process, which is difficult to meet the requirements in some well areas; neural network algorithm (shale total organic carbon prediction based on graph neural network Methods, systems and devices, using neural networks to predict TOC data, patent number: CN202111411640.8) It is difficult to solve the overfitting problem when the data is redundant, and its internal neural network is built based on a random screening mechanism, and it is difficult to ensure the stability of each prediction result; the seismic information quantitative prediction algorithm (a method and device for quantitatively predicting the mass fraction of organic carbon in source rocks by seismic information, combining two-dimensional seismic data and logging data to predict TOC, patent number: CN201610922188.4) The method is too complex, the model accuracy requirements are high, and it combines a lot of seismic information. It is difficult for some study areas to meet the data provision conditions.

[0004] In summary, the current prediction of TOC content in source rocks is mostly based on direct or indirect prediction of TOC content based on autonomous machine learning algorithms, virtual well construction predictions, neural network algorithms, and seismic information quantitative prediction algorithms. There is a lack of in-depth mining of stratigraphic information, and these prediction methods are often unable to meet exploration needs in terms of accuracy, especially in areas with low exploration levels, where prediction accuracy is difficult to guarantee. Summary of the invention

[0005] To solve the above problems, the present invention provides a method for predicting TOC of source rocks. The method is based on reorganized multiple logging data and partially measured TOC data, and adopts a convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm to construct a prediction model to predict the TOC content of the entire source rock in the study area. The method has high prediction accuracy.

[0006] To achieve the above object, the present invention provides the following technical solutions.

[0007] A method for predicting TOC of a source rock comprises the following steps:

[0008] Obtain the measured total organic carbon data and multiple logging data obtained from the core sample experiments of the source rocks in the sampling wells in the study area, and perform normalization and reorganization to obtain the preliminary logging curve values;

[0009] The processed preliminary logging curve values ​​are judged according to the data fuzzy region model to obtain the effective data prediction interval, and the abnormal values ​​of the data related to the prediction results are removed to obtain the final logging related data;

[0010] A first prediction model based on a convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm is constructed; the measured total organic carbon data obtained from the experiment and the final logging-related data are input into the trained first prediction model to obtain the predicted total organic carbon data of the unexperimented part of the sampling well; wherein, the first prediction model is based on a gated recurrent unit network model, an adaptive reverse layer is added to the forward layer of the gated recurrent unit network model, and the front and back hidden vectors are concatenated to form a bidirectional gated recurrent unit network model; when the bidirectional gated recurrent unit network model is trained, each part is nested with an attention mechanism module to participate in the training.

[0011] Preferably, it also includes:

[0012] Perform fuzzy prediction on the logging data of non-sampling wells according to the data fuzzy region model, remove abnormal values ​​of data related to the prediction results, and obtain the logging related data of non-sampling wells;

[0013] A second prediction model based on a convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm is constructed; based on the measured total organic carbon data and predicted total organic carbon data obtained from the experiment of the sampling wells, and the final logging-related data of the sampling wells and non-sampling wells, the predicted total organic carbon data of the non-sampling wells are obtained through the trained second prediction model;

[0014] A third prediction model based on a convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm was constructed; based on the predicted total organic carbon data of sampling wells and non-sampling wells, the total organic carbon and total organic carbon plane distribution data of source rock sections of other non-sampling wells in the study area were obtained through the trained third prediction model;

[0015] Among them, the second prediction model and the third prediction model are both based on the gated recurrent unit network model, an adaptive reverse layer is added to the forward layer of the gated recurrent unit network model, and the front and back hidden vectors are spliced ​​to form a bidirectional gated recurrent unit network model; when the bidirectional gated recurrent unit network model is trained, each part is nested with an attention mechanism module to participate in the training.

[0016] Preferably, the method of obtaining the measured total organic carbon data obtained from the core samples of hydrocarbon source rocks in the sampling wells in the study area includes the following steps:

[0017] Collect multiple core samples of source rocks from source rock development intervals of sampling wells in the study area according to preset sampling intervals and preset sampling quantities;

[0018] The core samples of each source rock are placed in a total organic carbon analyzer for analysis to obtain the measured total organic carbon data of the core samples of each source rock.

[0019] Preferably, the plurality of logging data include a natural gamma ray GR curve, an acoustic time difference AC curve, a resistivity RT curve, a density DEN curve, a compensated thermal neutron CNL curve, a uranium U curve and a potassium K curve of the sampling well;

[0020] The acquisition of the preliminary logging curve value includes:

[0021] The natural gamma ray GR curve, acoustic time difference AC curve, resistivity RT curve, density DEN curve, compensated thermal neutron CNL curve, uranium U curve and potassium K curve of the sampling well are integrated, the window data are interpolated and omitted, and the missing values ​​and predicted values ​​of different window sizes are calculated by the following formula.

[0022] Preferably, judging the processed preliminary logging curve value according to the data fuzzy region model to obtain the effective data prediction interval, removing the abnormal values ​​of the data related to the prediction result, and obtaining the final logging related data includes the following steps:

[0023] Analyze the precision and recall of well logging data;

[0024] The precision is as follows:

[0025]

[0026] In the formula, TP is a true positive example and FP is a false positive example;

[0027] The recall rate Recall is shown as follows:

[0028]

[0029] Where, FN is a false negative example;

[0030] The number of positive samples is flattened to verify the mutual restriction parameters between precision and recall, and to measure the accuracy of the model input data:

[0031]

[0032] When the value of F1 is smaller, the true positive TP increases relatively, and the false positive FP decreases relatively, that is, the precision Precision and recall Recall both increase relatively. F1 weights both the precision Precision and recall Recall to determine the effective data prediction interval;

[0033] According to the effective data prediction interval, the abnormal values ​​of the data related to the prediction results are removed to obtain the final logging related data.

[0034] Preferably, the first prediction model is trained by dividing the training set and the test set by a selection cross-validation method, comprising the following steps:

[0035] The union of k-1 subsets is used as the training set of the first prediction model, and the remaining subsets are used as the test set. K training tests are performed, and the mean of the k test results is taken to adjust the feature dimension parameters.

[0036] Set the initial learning rate to 0.01 and the number of learning times to 350;

[0037] Establish a regression layer sequence, adjust the convolution kernel to [2,1], set the step size to [1,1], and the number of channels to 32;

[0038] Adjust the data tiling method, establish the data tiling layer, global average pooling layer and sequence folding layer; adjust the Adam gradient descent algorithm, verify the descent factor, and adjust the parameter settings;

[0039] Compare the output results of the training set and the test set. If the accuracy of the result exceeds the preset threshold, it passes. Otherwise, continue to adjust the parameters of the first prediction model.

[0040] Preferably, it also includes:

[0041] All the predicted total organic carbon data were analyzed for correlation, and the overall distribution of total organic carbon in the study area was obtained.

[0042] Based on the core data and logging data of the sampling wells in the study area, the distribution sections of high total organic carbon content in the source rock development sections of the sampling wells are determined; according to the changing characteristics of the natural gamma GR curve, the source rock development sections of the sampling wells are determined; according to the vertical distribution of the predicted total organic carbon data, the distribution sections of the total organic carbon content in the unsampled wells are determined.

[0043] A prediction system for source rock TOC, the system comprising:

[0044] processor;

[0045] a memory having stored thereon a computer program executable on the processor;

[0046] Wherein, when the computer program is executed by the processor, the steps of the method for predicting source rock TOC are implemented.

[0047] A computer-readable storage medium stores a data processing program, which implements the steps of the method for predicting source rock TOC when executed by a processor.

[0048] Beneficial effects of the present invention:

[0049] The present invention proposes a prediction method, system and medium for source rock TOC. The method is based on reorganized multiple logging data and partially measured TOC data, and uses a convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm to build a prediction model to predict the TOC content of the entire source rock in the study area. Among them, the method is based on the GRU model, by adding an adaptive reverse layer on the basis of the GRU forward layer, and splicing the front and back hidden vectors to form a bidirectional gated recurrent unit learning network with a bidirectional multi-output learning function. The learning network integrates the ability of past and future information learning, fully considers the forward and reverse time sequence of data, improves the shortcomings of GRU unidirectional prediction, and greatly enhances the ability to mine and predict parameters related to formation information. On this basis, the attention mechanism algorithm is added, so that the model can simulate the weight adjustment strategy designed for the behavior of the human brain to automatically perceive local important information, continuously change the model output size during the calculation process, and embed each part of the BiGRU network learning model to participate in the training, fully consider the correlation between various features, and effectively improve the prediction accuracy of source rock TOC. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is a flow chart for predicting source rock TOC according to an embodiment of the present invention;

[0051] Figure 2 It is the GRU network composition;

[0052] Figure 3 : is a bidirectional gated recurrent unit BiGRU network structure diagram of an embodiment of the present invention;

[0053] Figure 4 is the Attention module structure of an embodiment of the present invention;

[0054] Figure 5Schematic diagram of prediction results of an embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0056] Example 1

[0057] A method for predicting TOC of a source rock of the present invention has a specific process as follows: Figure 1 As shown, the following steps are included:

[0058] S1: Obtain the measured total organic carbon data and multiple logging data obtained from the core sample experiments of the source rocks in the sampling wells in the study area, and perform normalization and reorganization to obtain the preliminary logging curve values.

[0059] S2: According to the data fuzzy region model, the processed preliminary logging curve value is judged to obtain the effective data prediction interval, and the abnormal values ​​of the data related to the prediction result are removed to obtain the final logging related data.

[0060] S3: Construct the first prediction model based on the convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm; input the measured total organic carbon data obtained in the experiment and the final logging-related data into the trained first prediction model to obtain the predicted total organic carbon data of the unexperimented part of the sampling well.

[0061] S4: Perform fuzzy prediction on the logging data of the non-sampling wells according to the data fuzzy region model, remove abnormal values ​​of the data related to the prediction result, and obtain the logging related data of the non-sampling wells.

[0062] S5: Construct a second prediction model based on the convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm; according to the measured total organic carbon data and predicted total organic carbon data obtained from the experiment of the sampling wells, and the final logging-related data of the sampling wells and non-sampling wells, the predicted total organic carbon data of the non-sampling wells are obtained through the trained second prediction model.

[0063] S6: Construct a third prediction model based on the convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm; according to the predicted total organic carbon data of sampling wells and non-sampling wells, obtain the total organic carbon and total organic carbon planar distribution data overview of the source rock sections of other non-sampling wells in the study area through the trained third prediction model.

[0064] In S6, specifically:

[0065] S6.1: Perform correlation analysis on all the predicted total organic carbon data obtained to obtain an overview of the total organic carbon planar distribution data in the study area.

[0066] S6.2: Based on the core data and logging data of the sampling wells in the study area, determine the distribution sections of high total organic carbon content in the source rock development sections of the sampling wells; determine the source rock development sections of the sampling wells based on the changing characteristics of the natural gamma GR curve; determine the distribution sections of the total organic carbon content in the unsampled wells based on the vertical distribution of the predicted total organic carbon data.

[0067] The first prediction model, the second prediction model and the third prediction model are all based on the gated recurrent unit network model. An adaptive reverse layer is added to the forward layer of the gated recurrent unit network model GRU, and the front and back hidden vectors are concatenated to form a bidirectional gated recurrent unit network model. When training the bidirectional gated recurrent unit network model, each part is nested with an attention mechanism module to participate in the training.

[0068] Furthermore, the convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm is improved by the GRU model to increase its calculation of the reverse layer order and concatenate the front and back hidden layer vectors to form a bidirectional learning function, so that the prediction of TOC data related to the physical and electrical properties of the stratum sediments can be more accurate, the learning depth is deeper, and the learning memory is retained longer. Among them, the GRU network structure is mainly composed of the reset gate Rt and the update gate Zt, such as Figure 2 As shown, Figure 2 This is the GRU network structure diagram. The following formula is the calculation formula of the GRU hidden unit:

[0069]

[0070] In the formula, X t is the input at time t; R t , Z t They are reset gate and update gate respectively; W z , W r , W are parameter training matrices; The current node activation status; h t-1 and h t are the node outputs at the previous moment and the current moment respectively; σ is the Sigmoid function.

[0071] A reverse layer is added to the GRU forward layer, and the front and back hidden layer vectors are concatenated to form a bidirectional gated recurrent unit (BiGRU) network. The network structure is as follows: Figure 3 shown.

[0072] Multiple logging data include the natural gamma GR curve, acoustic time difference AC curve, resistivity RT curve, density DEN curve, compensated thermal neutron CNL curve, uranium U curve and potassium K curve of the sampling well. In the process of substituting the normalized calculation results into the BiGRU model, the influence of various logging data will vary. Relatively high weights are given to important features with large influences, and lower weights are used to reduce the influence of secondary features, thereby improving the model accuracy and prediction value accuracy. Therefore, the attention mechanism strategy is introduced. Its characteristics are that the attention mechanism (Attention) is a weight adjustment strategy designed to model the behavior of the human brain automatically perceiving local important information. The model output size is continuously changed during the calculation process, and can be nested in each part of the BiGRU network learning model to participate in training, fully considering the correlation between various features. The structure of the Attention module is as follows: Figure 4 As shown. In the Attention module, X1, X2, X3...X n is n input vectors, q is the query vector, and the query vector is obtained by linear transformation of the input vector, where the weight matrix w is a trainable parameter. The query vector representation formula is as follows:

[0073] q=wx i , i=1, 2, 3..., n;

[0074] S1, S2, S3...Sn are scaled dot product models, which are used to calculate the correlation between the query vector and each input vector. The scaled dot product model formula is as follows:

[0075]

[0076] Among them, d is the input dimension. Applying the scaled dot product model for correlation calculation is more efficient. The calculation result conforms to the normal distribution, avoiding the situation where the update gradient is too small or even 0, and accelerating the convergence speed of the softmax activation function. a1, a2, a3…a n represents the weighted probability obtained by normalizing the correlation result through the softmax activation function, and the obtained weighted probability is respectively combined with the input vectors x1, x2, x3...x n The corresponding multiplication and summation are used to obtain a new output, and the above calculation is repeated until all output results are updated. Finally, the Attention module selects a feature dimension that is most decisive for the target from the input vector and assigns it a higher weight.

[0077] Specifically, obtaining the measured total organic carbon data obtained from the core sample experiment of the source rock of the sampling well in the study area includes the following steps:

[0078] From the source rock development sections of the sampling wells in the study area, multiple source rock core samples are collected according to the preset sampling intervals and the preset sampling quantity.

[0079] The core samples of each source rock are placed in a total organic carbon analyzer for analysis to obtain the measured total organic carbon data of the core samples of each source rock.

[0080] The acquisition of the preliminary logging curve value includes the following steps:

[0081] The natural gamma ray GR curve, acoustic time difference AC curve, resistivity RT curve, density DEN curve, compensated thermal neutron CNL curve, uranium U curve and potassium K curve of the sampling well are integrated, the window data are interpolated and omitted, and the missing values ​​and predicted values ​​of different window sizes are calculated by the following formula.

[0082] Further, the processed preliminary logging curve values ​​are judged according to the data fuzzy region model to obtain the effective data prediction interval, and the abnormal values ​​of the data related to the prediction result are removed to obtain the final logging related data, including the following steps:

[0083] Analyze the precision and recall of well logging data;

[0084] The precision is as follows:

[0085]

[0086] In the formula, TP is a true positive example and FP is a false positive example.

[0087] The recall rate Recall is shown as follows:

[0088]

[0089] Where FN is the false negative example.

[0090] The number of positive samples is flattened to verify the mutual restriction parameters between precision and recall, and to measure the accuracy of the model input data:

[0091]

[0092] When the value of F1 is smaller, the true positive TP increases relatively and the false positive FP decreases relatively, that is, the precision Precision and recall Recall both increase relatively. F1 weights both the precision Precision and the recall Recall to determine the effective data prediction interval.

[0093] According to the effective data prediction interval, the abnormal values ​​of the data related to the prediction results are removed to obtain the final logging related data.

[0094] The first prediction model uses the selective cross-validation method to divide the training set and the test set for training. The union of k-1 subsets is used as the training set of the first prediction model, and the remaining subsets are used as the test set. K training tests are performed, and the average of k test results is taken to adjust the feature dimension parameters; the initial learning rate is set to 0.01, and the number of learning times is 350; the regression layer sequence is established, the convolution kernel is adjusted to [2,1], the step size is set to [1,1], and the number of channels is 32; the data tiling method is adjusted, and the data tiling layer, global average pooling layer and sequence folding layer are established; the Adam gradient descent algorithm is adjusted, the descent factor is verified, and the parameter settings are adjusted; the output results of the training set and the test set are compared. If the accuracy of the result exceeds the preset threshold, it passes, otherwise continue to adjust the parameters of the first prediction model. The experimental results are as follows Figure 5 shown.

[0095] The training of the second prediction model includes: adjusting the BiGRU model features from unidirectional output to bidirectional output; adjusting the CNN convolution logic from a unidirectional output system to a bidirectional loop output system; setting the initial learning rate to 0.02 and the number of learning times to 700; establishing a regression hierarchy, adjusting the convolution kernel to [3,1], setting the step size to [2,1] and the number of channels to 64; adjusting the Adam gradient descent algorithm, verifying the descent factor, and adjusting the parameter settings; verifying the reliability of the network structure, comparing the output results of the training set and the test set, outputting the algorithm process diagram and the iteration number diagram, and printing the output indicators.

[0096] The above is a prediction method for source rock TOC provided by an embodiment of this embodiment. Based on the same idea, this embodiment also provides a corresponding prediction system for source rock TOC. The specific definition of the prediction system for source rock TOC can be found in the definition of the prediction method for source rock TOC in the above text, which will not be repeated here. Each module in the above-mentioned prediction system for source rock TOC can be implemented in whole or in part through software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.

[0097] This embodiment also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 A prediction method for source rock TOC is provided.

[0098] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0099] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for predicting TOC of source rocks, characterized in that: The following steps are involved: Obtain the measured total organic carbon data and multiple logging data obtained from the core sample experiments of the source rocks in the sampling wells in the study area, and perform normalization and reorganization to obtain the preliminary logging curve values; The processed preliminary logging curve values ​​are judged according to the data fuzzy region model to obtain the effective data prediction interval, and the abnormal values ​​of the data related to the prediction results are removed to obtain the final logging related data; Construct a first prediction model based on a convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm; input the measured total organic carbon data obtained in the experiment and the final logging-related data into the trained first prediction model to obtain the predicted total organic carbon data of the untested part of the sampling well; wherein, the first prediction model is based on a gated recurrent unit network model, an adaptive reverse layer is added to the forward layer of the gated recurrent unit network model, and the front and back hidden vectors are concatenated to form a bidirectional gated recurrent unit network model; when the bidirectional gated recurrent unit network model is trained, each part is embedded with an attention mechanism module to participate in the training; Also includes: Perform fuzzy prediction on the logging data of non-sampling wells according to the data fuzzy region model, remove abnormal values ​​of data related to the prediction results, and obtain the logging related data of non-sampling wells; A second prediction model based on a convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm is constructed; based on the measured total organic carbon data and predicted total organic carbon data obtained from the experiment of the sampling wells, and the final logging-related data of the sampling wells and non-sampling wells, the predicted total organic carbon data of the non-sampling wells are obtained through the trained second prediction model; A third prediction model based on a convolutional neural network-bidirectional gated recurrent unit-attention mechanism algorithm was constructed; based on the predicted total organic carbon data of sampling wells and non-sampling wells, the total organic carbon and total organic carbon plane distribution data of source rock sections of other non-sampling wells in the study area were obtained through the trained third prediction model; Among them, the second prediction model and the third prediction model are both based on the gated recurrent unit network model, an adaptive reverse layer is added to the forward layer of the gated recurrent unit network model, and the front and back hidden vectors are spliced ​​to form a bidirectional gated recurrent unit network model; when the bidirectional gated recurrent unit network model is trained, each part is nested with an attention mechanism module to participate in the training.

2. The method for predicting source rock TOC according to claim 1, characterized in that: The method of obtaining the measured total organic carbon data obtained from the core sample experiment of the hydrocarbon source rock in the sampling well in the study area includes the following steps: Collect multiple core samples of source rocks from source rock development intervals of sampling wells in the study area according to preset sampling intervals and preset sampling quantities; The core samples of each source rock are placed in a total organic carbon analyzer for analysis to obtain the measured total organic carbon data of the core samples of each source rock.

3. The method for predicting source rock TOC according to claim 1, characterized in that: The multiple logging data include the natural gamma GR curve, acoustic time difference AC curve, resistivity RT curve, density DEN curve, compensated thermal neutron CNL curve, uranium U curve and potassium K curve of the sampling well; The acquisition of the preliminary logging curve value includes: The natural gamma ray GR curve, acoustic time difference AC curve, resistivity RT curve, density DEN curve, compensated thermal neutron CNL curve, uranium U curve and potassium K curve of the sampling well are integrated, and the window data are interpolated and omitted.

4. The method for predicting source rock TOC according to claim 1, characterized in that: The method of judging the processed preliminary logging curve value according to the data fuzzy region model to obtain the effective data prediction interval, removing the abnormal values ​​of the data related to the prediction result, and obtaining the final logging related data includes the following steps: Analyze the precision and recall of well logging data; The precision is as follows: In the formula, TP is a true positive example and FP is a false positive example; The recall rate Recall is shown as follows: Where, FN is a false negative example; The number of positive samples is flattened to verify the mutual restriction parameters between precision and recall, and to measure the accuracy of the model input data: When the value of F1 is smaller, the true positive TP increases relatively, and the false positive FP decreases relatively, that is, the precision Precision and recall Recall both increase relatively. F1 weights both the precision Precision and recall Recall to determine the effective data prediction interval; According to the effective data prediction interval, the abnormal values ​​of the data related to the prediction results are removed to obtain the final logging related data.

5. The method for predicting source rock TOC according to claim 1, characterized in that: The first prediction model is trained by dividing the training set and the test set by the selection cross-validation method, including the following steps: The union of k-1 subsets is used as the training set of the first prediction model, and the remaining subsets are used as the test set. K training tests are performed, and the mean of the k test results is taken to adjust the feature dimension parameters. Set the initial learning rate to 0.01 and the number of learning times to 350; Establish a regression layer sequence, adjust the convolution kernel to [2,1], set the step size to [1,1], and the number of channels to 32; Adjust the data tiling method, establish the data tiling layer, global average pooling layer and sequence folding layer; adjust the Adam gradient descent algorithm, verify the descent factor, and adjust the parameter settings; Compare the output results of the training set and the test set. If the accuracy of the result exceeds the preset threshold, it passes. Otherwise, continue to adjust the parameters of the first prediction model.

6. The method for predicting source rock TOC according to claim 1, characterized in that: Also includes: All the predicted total organic carbon data were analyzed for correlation, and the overall distribution of total organic carbon in the study area was obtained. Based on the core data and logging data of the sampling wells in the study area, the distribution sections of high total organic carbon content in the source rock development sections of the sampling wells are determined; according to the changing characteristics of the natural gamma GR curve, the source rock development sections of the sampling wells are determined; according to the vertical distribution of the predicted total organic carbon data, the distribution sections of the total organic carbon content in the unsampled wells are determined.

7. A prediction system for source rock TOC, characterized in that: The system comprises: processor; a memory having stored thereon a computer program executable on the processor; Wherein, when the computer program is executed by the processor, the steps of the method for predicting source rock TOC as described in any one of claims 1 to 6 are implemented.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a data processing program, and when the data processing program is executed by a processor, the steps of the method for predicting source rock TOC as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Method and device of quantitatively predicting organic carbon mass fraction of hydrocarbon source rock based on earthquake information

    CN107976711A

  • Methods, systems, and equipment for predicting total organic carbon in shale based on graph neural networks.

    CN113837501B

  • Method and system for predicting TOC (total organic carbon) of source rock in sparse well area, electronic equipment and medium

    CN115629414A

  • Shale TOC content prediction method and device based on machine learning

    CN118050821A