Methods and apparatus for determining the main controlling factors of CO2 huff and puff in shale oil reservoirs

CN117386333BActive Publication Date: 2026-09-11XI'AN PETROLEUM UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311543849.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2026-09-11
Estimated Expiration
2043-11-20

AI Technical Summary

Technical Problem

[0004]本申请实施例通过提供一种页岩油藏注CO2吞吐的主控因素确定方法及装置,解决了现有主控因素分析方法中的不同影响因素相互影响,导致分析结果存在偶然性的问题,实现了对不同影响因素进行量化排序和敏感性分析

Benefits of technology

本申请实施例通过对影响因素进行抽样、模拟,进行循环校验并优化学习模型,有效解决了现有的主控因素分析方法中的不同影响因素相互影响,导致分析结果存在偶然性的问题,进而实现了对不同影响因素进行量化排序和敏感性分析,能够快速、准确地确定二氧化碳吞吐的主控因素。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117386333B_ABST
    Figure CN117386333B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for determining main control factors of CO2 injection output in shale oil reservoirs, which comprises the following steps: determining the influencing factors of CO2 injection output and their value ranges; sampling the influencing factors to obtain a parameter vector, and performing numerical simulation according to the parameter vector to obtain simulation results; determining the storage rate and oil replacement rate according to the simulation results; performing a cyclic verification step according to the data set; the cyclic verification step comprises the following steps: determining a sample proportionality coefficient, grouping the data set according to the sample proportionality coefficient; determining the key hyperparameters of a learning model, training the learning model by taking the grouped data set as a sample, and obtaining an initial prediction result; optimizing the learning model to obtain an optimized learning model prediction result. The application solves the problem of accidental analysis results of the existing main control factor analysis method, realizes quantitative ordering and sensitivity analysis of different influencing factors, and can quickly and accurately determine the main control factors of CO2 injection output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of unconventional reservoir development technology, and in particular to a method and apparatus for determining the main controlling factors of CO2 injection huff and puff in shale oil reservoirs. Background Technology

[0002] Carbon dioxide huff and puff (CO2 huff and puff) refers to the injection of a certain amount of carbon dioxide into the oil reservoir under a certain pressure, followed by a period of well simmering to allow the carbon dioxide to diffuse and dissolve in the crude oil before well extraction. The mechanism by which CO2 huff and puff technology improves oil recovery is as follows: after carbon dioxide is injected into the oil reservoir, it dissolves in the crude oil and formation water, causing carbonation of the formation water. This reduces the viscosity of the crude oil, increases the mobility ratio of crude oil to formation water, and expands the swept volume of carbon dioxide in the formation, thereby improving the oil recovery rate.

[0003] Oil exchange rate and burial rate are important indicators for evaluating the CO2 injection and sumption development effect of shale oil reservoirs. However, these two indicators are affected by many factors, including CO2 injection rate, CO2 gas purity, formation pressure, oil layer thickness, injection and sumption timing, and well shut-in time, and the degree of influence of each factor varies for different projects. Currently, the main controlling factors for CO2 injection and sumption in shale oil reservoirs are mostly determined using univariate comparative analysis. This involves adjusting the value of one parameter while keeping other parameters constant, and then using numerical simulation results to predict and compare the results of the indicators (oil exchange rate and burial rate) corresponding to different parameter values. This method is simple to operate and yields clear results. However, different parameters may influence each other, and the univariate analysis method is affected by the values ​​of other "control variables," leading to randomness in the results. In addition, the univariate analysis method cannot uniformly quantify and rank the degree of influence of different factors. Summary of the Invention

[0004] This application provides a method and apparatus for determining the main controlling factors of CO2 injection and purge in shale oil reservoirs. This solves the problem that different influencing factors interact with each other in existing main controlling factor analysis methods, leading to randomness in the analysis results. It enables quantitative ranking and sensitivity analysis of different influencing factors.

[0005] In a first aspect, embodiments of this application provide a method for determining the main controlling factors of CO2 injection and purge in shale oil reservoirs, comprising: determining the influencing factors of CO2 injection and purge in shale oil reservoirs and the value range of the influencing factors; sampling the influencing factors to obtain a parameter vector, and performing numerical simulation based on the parameter vector to obtain simulation results; wherein the simulation results include oil production, gas production, and gas injection volume; determining the burial rate and oil exchange rate based on the simulation results, and establishing a comprehensive index of the burial rate and oil exchange rate; performing a cyclic verification step based on the acquired dataset; wherein the dataset includes the influencing factors and the corresponding values ​​of the influencing factors, burial rate, oil exchange rate, and the comprehensive index; the cyclic verification step includes: determining the sample proportion coefficient, and grouping the dataset according to the sample proportion coefficient; determining the key hyperparameters of the learning model, and using the grouped dataset as samples to train the learning model to obtain the initial prediction results of the main controlling factors; optimizing the learning model to obtain the optimized prediction results of the main controlling factors predicted by the optimized learning model.

[0006] In conjunction with the first aspect, in the first possible implementation, the influencing factors include gas injection volume, gas injection rate, well shut-in time, huff and puff cycles, gas injection pressure, porosity factor, average permeability factor, and production pressure.

[0007] In conjunction with the first aspect, in the second possible implementation, the sampling of the influencing factors to obtain the parameter vector includes: determining the total number of samples to be sampled and performing a sampling step; the sampling step includes: randomly selecting random parameters from all the influencing factors and extracting corresponding values ​​for the random parameters; and constructing the parameter vector based on the results of multiple samplings.

[0008] In conjunction with the second possible implementation of the first aspect, in the third possible implementation, the sampling of the influencing factors to obtain the parameter vector further includes: performing the sampling step multiple times to obtain multiple parameter vectors.

[0009] In conjunction with the first aspect, the fourth possible implementation, after obtaining the simulation results by performing numerical simulation based on the parameter vector, further includes: obtaining the extraction parameters for non-CO2 huff and puff, and performing numerical simulation to obtain conventional simulation results; wherein, the simulation results include conventional oil production, conventional gas injection, and conventional gas production.

[0010] In conjunction with the first aspect, in the fifth possible implementation, optimizing the learning model includes: calculating the root mean square error (RMSE) of the learning model respectively; using the sample proportion parameter and the key hyperparameter corresponding to the minimum value of the RMSE as the optimized sample proportion parameter and the optimized key hyperparameter; and performing the iterative verification step based on the optimized sample proportion parameter and the optimized key hyperparameter to obtain the optimized learning model.

[0011] In conjunction with the fifth possible implementation of the first aspect, in the sixth possible implementation, the step of calculating the root mean square error of the learning model includes: the key hyperparameters include the minimum number of leaves and the number of decision trees; establishing multiple ordered arrays for storing the minimum number of leaves, the number of decision trees, and the root mean square error; setting a set of values ​​for the minimum number of leaves and the number of decision trees; randomly selecting values ​​multiple times from the set of values, and substituting the selected values ​​into the learning model to obtain multiple root mean square errors.

[0012] In conjunction with the sixth possible implementation of the first aspect, in the seventh possible implementation, the step of calculating the root mean square error of the learning model includes: establishing an ordered coefficient array for storing the sample proportion coefficients; performing arithmetic operations on the sample proportion coefficients and storing the results in the ordered coefficient array; performing the cyclic verification step based on the ordered coefficient array and determining the root mean square error of the learning model.

[0013] Secondly, embodiments of this application provide a device for determining the main controlling factors of CO2 injection in shale oil reservoirs, comprising: a determination module for determining the influencing factors of CO2 injection in shale oil reservoirs and the value range of the influencing factors; a simulation module for sampling the influencing factors to obtain a parameter vector and performing numerical simulation based on the parameter vector to obtain simulation results; wherein the simulation results include oil production, gas production, and gas injection volume; a calculation module for determining the burial rate and oil exchange rate based on the simulation results and establishing a comprehensive index of the burial rate and oil exchange rate; and a verification module for performing a cyclic verification step based on the acquired dataset; wherein the dataset includes the influencing factors and the corresponding values, burial rate, oil exchange rate, and comprehensive index of the influencing factors; the cyclic verification step includes: determining a sample proportion coefficient and grouping the dataset according to the sample proportion coefficient; determining the key hyperparameters of the learning model and using the grouped dataset as samples to train the learning model to obtain initial prediction results of the main controlling factors; and an optimization module for optimizing the learning model to obtain optimized prediction results of the main controlling factors predicted by the optimized learning model.

[0014] In conjunction with the second aspect, in the first possible implementation, the influencing factors include gas injection volume, gas injection rate, well shut-in time, huff and puff cycles, gas injection pressure, porosity factor, average permeability factor, and production pressure.

[0015] In conjunction with the second aspect, in the second possible implementation, the step of sampling the influencing factors to obtain a parameter vector includes: determining the total number of samples to be sampled and performing a sampling step; the sampling step includes: randomly selecting random parameters from all the influencing factors and extracting corresponding values ​​for the random parameters; and constructing the parameter vector based on the results of multiple samplings.

[0016] In conjunction with the second possible implementation of the second aspect, in the third possible implementation, the sampling of the influencing factors to obtain the parameter vector further includes: performing the sampling step multiple times to obtain multiple parameter vectors.

[0017] In conjunction with the second aspect, the fourth possible implementation, after obtaining the simulation results by performing numerical simulation based on the parameter vector, further includes: obtaining the extraction parameters for non-CO2 huff and puff, and performing numerical simulation to obtain conventional simulation results; wherein, the simulation results include conventional oil production, conventional gas injection, and conventional gas production.

[0018] In conjunction with the second aspect, in the fifth possible implementation, optimizing the learning model includes: calculating the root mean square error (RMSE) of the learning model respectively; using the sample proportion parameter and the key hyperparameter corresponding to the minimum value of the RMSE as the optimized sample proportion parameter and the optimized key hyperparameter; and performing the iterative verification step based on the optimized sample proportion parameter and the optimized key hyperparameter to obtain the optimized learning model.

[0019] In conjunction with the fifth possible implementation of the second aspect, in the sixth possible implementation, the step of calculating the root mean square error of the learning model includes: the key hyperparameters include the minimum number of leaves and the number of decision trees; establishing multiple ordered arrays for storing the minimum number of leaves, the number of decision trees, and the root mean square error; setting a set of values ​​for the minimum number of leaves and the number of decision trees; randomly selecting values ​​multiple times from the set of values, and substituting the selected values ​​into the learning model to obtain multiple root mean square errors.

[0020] In conjunction with the sixth possible implementation of the second aspect, in the seventh possible implementation, the step of calculating the root mean square error of the learning model includes: establishing an ordered coefficient array for storing the sample proportion coefficients; performing arithmetic operations on the sample proportion coefficients and storing the results in the ordered coefficient array; performing the cyclic verification step based on the ordered coefficient array and determining the root mean square error of the learning model.

[0021] Thirdly, embodiments of this application provide an apparatus comprising: a processor; a memory for storing processor-executable instructions; wherein, when the processor executes the executable instructions, it implements the method as described in the first aspect or any possible implementation of the first aspect.

[0022] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages: This application embodiment effectively solves the problem of randomness in the analysis results caused by the mutual influence of different influencing factors in existing main control factor analysis methods by sampling and simulating influencing factors, performing cyclic verification and optimizing the learning model. It then realizes the quantitative ranking and sensitivity analysis of different influencing factors, and can quickly and accurately determine the main control factors of carbon dioxide throughput. Attached Figure Description

[0023] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A flowchart illustrating the method for determining the main controlling factors of CO2 huff and puff in shale oil reservoirs provided in this application embodiment; Figure 2 This is a schematic diagram of the device for determining the main controlling factors of CO2 injection and purge in shale oil reservoirs provided in the embodiments of this application; Figure 3 A statistical chart of oil change rates corresponding to the numerical simulation examples provided in this application; Figure 4 A statistical chart of the burial rate corresponding to the numerical simulation examples provided in this application embodiment; Figure 5 The initial prediction results of the learning model provided in the embodiments of this application; Figure 6 The optimized prediction results of the optimized learning model provided in the embodiments of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0026] The following description of some technologies involved in the embodiments of this application is provided to aid understanding and should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, some descriptions of well-known functions and structures are omitted in the following description.

[0027] Figure 1 This is a flowchart of a method for determining the main controlling factors of CO2 huff and puff in shale oil reservoirs, as provided in an embodiment of this application, including steps 101 to 106. Figure 1 This is merely one execution order shown in the embodiments of this application and does not represent the only execution order of the method for determining the main controlling factors of CO2 injection and purge in shale oil reservoirs. Where the final result can be achieved, Figure 1 The steps shown can be performed in parallel or in reverse order.

[0028] Step 101: Determine the influencing factors and their value ranges for CO2 huff and puff in shale oil reservoirs. Specifically, the oil exchange rate and burial rate of CO2 huff and puff development in shale oil reservoirs are affected by a series of factors, including the reservoir's geology and development conditions, as well as the various measures implemented. Based on literature review and the specific conditions of the actual reservoir, determine the influencing factors and their value ranges.

[0029] In this embodiment, the influencing factors include gas injection volume, gas injection rate, well shut-in time, number of injection cycles, gas injection pressure, porosity multiple, average permeability multiple, and production pressure. For example, the gas injection volume ranges from 2000 to 10000 tons, the gas injection rate ranges from 50 to 90 tons per day, the well shut-in time ranges from 30 to 90 days, the number of injection cycles ranges from 3 to 7 cycles, the gas injection pressure ranges from 160 to 200 bar, the porosity multiple ranges from 0.6 to 1.4, the average permeability multiple ranges from 0.5 to 2.5, and the production pressure ranges from 200 to 300 bar.

[0030] Those skilled in the art should realize that the above-mentioned influencing factors and their value ranges are only one embodiment of this application and are not intended to limit the scope of protection of this application. Those skilled in the art can make various modifications and changes to them according to the actual situation.

[0031] Step 102: Sample the influencing factors to obtain parameter vectors, and perform numerical simulations based on the parameter vectors to obtain simulation results. Specifically, analyzing the controlling factors using a learning model requires a large sample size, and large-sample random sampling is performed based on a uniform distribution probability density function.

[0032] In this embodiment, the total number of samples to be sampled is determined, and a sampling step is performed. Specifically, the identified influencing factors are placed into a set, where the number of influencing factors is the total number of samples. This application identifies eight influencing factors, so the total number of samples is eight.

[0033] The sampling step includes: randomly selecting random parameters from all influencing factors and selecting corresponding values ​​for the random parameters. Specifically, random sampling is performed within the set of influencing factors. For example, non-repeating sampling is used when selecting influencing factors; that is, once an influencing factor is selected, it is removed from the sampling set, and subsequent sampling will not repeat that influencing factor. For the selected influencing factors, it is assumed that they conform to a uniform distribution within the value range, and random values ​​are selected within that range. Alternatively, a random function can be used to select values, as follows: In the formula, x represents the value of the extracted influencing factor. The extracted influencing factor is the minimum value within its range. The maximum value of the extracted influencing factor within its range is r, which is a random number between [0,1].

[0034] A parameter vector is constructed based on the results of multiple samplings. Specifically, multiple samplings are performed within the range of values ​​for the influencing factor, and a parameter vector is constructed based on the results of these multiple samplings. ,in, This is the parameter vector of the extracted influencing factors. This is the sampling result for the nth value.

[0035] Multiple sampling steps are performed to obtain multiple parameter vectors. Specifically, random sampling is repeated multiple times until all influencing factors are sampled. In this embodiment, eight samplings without replacement are performed on the set of influencing factors, and multiple values ​​are sampled for each of the eight influencing factors within their respective value ranges. Based on the sampling results of the multiple values, eight parameter vectors are constructed for each of the eight influencing factors.

[0036] In this embodiment, based on the multiple parameter vectors obtained above, numerical simulation is performed using numerical simulation software such as Eclipse (a Java-based extensible development platform) or CMG (a reservoir numerical simulation software) to obtain simulation results. The simulation results include oil production, gas production, and gas injection volume.

[0037] In addition, an extra simulation without CO2 injection was performed to obtain the extraction parameters without CO2 injection, and numerical simulations were conducted to obtain conventional simulation results. These simulation results include conventional oil production, conventional gas injection, and conventional gas production.

[0038] Step 103: Determine the burial rate and oil change rate based on the simulation results, and establish a comprehensive index for the burial rate and oil change rate. In this embodiment, the calculation formulas for the burial rate and oil change rate are as follows: , In the formula, Indicates the burial rate. GPT represents the oil change rate, GIT represents the gas production rate, GIT represents the gas injection rate, and OPT represents the oil production rate. This indicates the regular oil production volume without CO2 huff and puff. For example... Figure 3 and Figure 4 The figures shown are the relationship between the number of calculation cases and the predicted oil change rate and burial rate.

[0039] The formula for calculating the combined index of retention rate and oil change rate is as follows: In the formula, r represents the comprehensive index. Indicates the burial rate. This indicates the oil change rate.

[0040] Based on the simulation results of each set of parameter vectors using the above method, the corresponding burial rate, oil change rate, and comprehensive index are calculated.

[0041] Step 104: Determine the sample proportion coefficient and group the dataset according to the sample proportion coefficient. In this embodiment, the dataset includes influencing factors and their corresponding values, burial rate, oil change rate, and comprehensive indicators.

[0042] In addition, the datasets were standardized before grouping. Different units of measurement can affect the results of data analysis. To eliminate the influence of different units of measurement between indicators, data normalization is required to improve the comparability between data indicators and reduce the risk of the model overlearning the data.

[0043] The data in the dataset is mapped to the range [0,1] using the following formula: In the formula, This represents the value of the i-th data point in the standardized dataset. This represents the value of the i-th data point obtained through numerical simulation software. and This represents the maximum and minimum values ​​of each group of influencing factors found. Furthermore, penetration rates are generally highly heterogeneous; the data distribution after normalization using the above formula is too concentrated at the two ends of [0,1], leading to unreasonable errors and distortions. Therefore, the following formula is used to standardize the penetration rate data: In the formula, This represents the standardized penetration rate. This represents the i-th value of the penetration rate obtained through numerical simulation software. and This represents the maximum and minimum values ​​of the permeability.

[0044] In this embodiment, the dataset is divided equally into N parts, with the sample proportion coefficient determined to be N parts. One part is used as the test dataset, and the remaining N-1 parts are used as the training dataset.

[0045] It is important to note that the training dataset is a data matrix containing parameters and results, which can be further divided into a parameter matrix and a result vector.

[0046] The training performance of a learned model is evaluated by its root mean square error (RMSE). A smaller RMSSE indicates a more reasonable partitioning of the dataset, while a larger RMSSE indicates a poorer partitioning.

[0047] Step 105: Determine the key hyperparameters of the learning model, and use the grouped dataset as samples to train the learning model to obtain the initial prediction results of the controlling factors. The key hyperparameters include the minimum number of leaves and the number of decision trees in the learning model.

[0048] In this embodiment, the learning model used is a random forest model. The random forest training model is built using the TreeBagger function provided by Matlab. The minimum number of leaves and the number of decision trees are key factors for each decision tree in the random forest to learn from the training dataset. A larger minimum number of leaves makes it easier for the decision tree to undergo "pruning" during growth, preventing it from growing fully. Conversely, a smaller minimum number of leaves can lead to overlearning of the training dataset by the learning model. Similarly, the number of decision trees directly affects the model's fitting ability (its ability to learn from data) and generalization ability (its predictive ability on new data). More decision trees result in stronger generalization ability but higher computational cost. Fewer decision trees result in weaker generalization ability. Initially, those skilled in the art set an initial minimum number of leaves and a specific number of decision trees for the learning model based on experience.

[0049] The divided training dataset is used as samples and fed into the learning model to obtain initial prediction results. For example... Figure 5 The results shown are the initial predictions, which demonstrate the importance of each influencing factor for CO2 throughput development. The factors with higher importance are the controlling factors.

[0050] Step 106: Optimize the learning model to obtain the optimized prediction results of the controlling factors predicted by the optimized learning model. To train a learning model with the best performance, it is common practice to manually adjust hyperparameters (hyperparameters are parameters set before the learning process begins, not parameters obtained through training) and dataset partitioning, mechanically achieving model performance. This static approach can lead to the risk of overfitting or underfitting the model when learning from the data.

[0051] In this embodiment, the learning model is optimized by using a cyclic validation method to optimize hyperparameters such as the model dataset partitioning (sample ratio parameter), minimum number of leaves, and number of decision trees, as detailed below.

[0052] The root mean square error (RMSE) of the learning model is calculated, and the sample proportion parameter and key hyperparameter corresponding to the minimum RMSE value are used as the optimization sample proportion parameter and optimization key hyperparameter. Specifically, multiple ordered arrays are created to store the minimum number of leaves, the number of decision trees, and the RMSE. For example, three ordered arrays L, T, and S of length M*N are created to store the minimum number of leaves, the number of decision trees, and the RMSE under the corresponding parameters, respectively. Here, N is the sample proportion parameter, i.e., the average number of divisions in the dataset, and M is the number of possible values ​​for the minimum number of leaves. L The number of parameter values ​​M in the decision tree T The product of, i.e., M=M L *M T .

[0053] Set a set of values ​​for the minimum number of leaves and the number of decision trees. Specifically, set the minimum number of leaves to 1, and select M in increments of 2. L The number of values ​​S represents the set of values ​​for the minimum number of leaves. L Let the decision tree have a base of 5, with the exponent starting from 1 and increasing sequentially, and select M in turn. T The set of possible values ​​S for the number of decision trees T Set S L and set S T After cross-combination, values ​​are selected, that is, S is randomly selected from each combination. L One of the values ​​L i and S T A value T j Combine them and store them in L and T respectively.

[0054] Multiple random values ​​are selected from the set of values, and the results are substituted into the learning model to obtain multiple root mean square errors. In this embodiment, the values ​​of the minimum number of leaves and the number of decision trees are selected. and Substituting these values ​​into the established learning model, we obtain the root mean square error of the learning model under the corresponding parameters. Store in the corresponding sorted array The calculation formula is as follows: In the formula, RMSE represents the root mean square error of the learned model. This represents the true value at the i-th verification point. This represents the predicted value of the learning model at the i-th validation point, and n represents the number of validation points.

[0055] The root mean square error (RMSE) is calculated each time a randomly selected value from the set of key hyperparameters of the learning model is input into the model, until the ordered array S is filled (i.e., increased to M). This process is repeated M times to obtain M RMSE values. The minimum number of leaves and the number of decision trees corresponding to the minimum RMSE are selected as the key hyperparameters for optimizing the learning model.

[0056] In addition, an ordered coefficient array is created to store the sample proportion coefficients. Specifically, an ordered coefficient array is created. ordered coefficient array Used to store sample proportion coefficients.

[0057] The sample proportion coefficient is then arithmetically assigned values, and the results are stored in an ordered coefficient array. Specifically, let... Starting with 2 With a tolerance of 3 , For the number of terms, Perform arithmetic progression and store the values ​​in an ordered coefficient array. In Chinese, the formula is as follows: In the formula, The value of the sample proportion coefficient. As the first item, For tolerance, The number of terms.

[0058] A cyclic verification step is performed based on the ordered coefficient array to determine the root mean square error of the learned model. Specifically, based on the stored ordered coefficient array... The values ​​in the range are used to modify the sample proportion coefficients in turn, and the root mean square error of the learning model after the sample proportion coefficients are modified is calculated. The sample proportion coefficient corresponding to the minimum value of the root mean square error is taken as the optimized sample proportion coefficient.

[0059] Based on the optimized sample proportion parameter and optimized key hyperparameters, a cyclical verification step is performed to obtain the optimized learning model. Specifically, the optimized key hyperparameters are set in the newly established learning model, and the dataset is re-partitioned using the optimized sample proportion coefficient. The learning model is then trained using the optimized training dataset to obtain the optimized learning model.

[0060] The optimized prediction results of the controlling factors predicted using the optimized learning model are as follows: Figure 6 As shown, the importance of each influencing factor for carbon dioxide throughput development is demonstrated, and the factors with higher importance are the main controlling factors.

[0061] This application calculated the root mean square error (RMSE) of the training and test datasets in the learning model before and after optimization. The RMSE of the training dataset in the learning model before optimization was 0.0336, while the RMSE of the optimized learning model was 0.0159, a reduction of 0.0177. The RMSE of the test dataset in the learning model before optimization was 0.0416, while the RMSE of the optimized learning model was 0.0200, a reduction of 0.0216. This demonstrates that the optimized sample proportion coefficient and key hyperparameters are more reasonable.

[0062] While this application provides the method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive labor. The order of steps listed in this embodiment is merely one possible execution order among many and does not represent the only execution order. In actual device or client product execution, the methods shown in this embodiment or the accompanying drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0063] like Figure 2 As shown in the embodiment of this application, a device 200 for determining the main controlling factors of CO2 injection into shale oil reservoirs is also provided. This device includes: a determination module 201, a simulation module 202, a calculation module 203, a verification module 204, and an optimization module 205, as detailed below.

[0064] Module 201 is used to determine the influencing factors and their value ranges for CO2 injection in shale oil reservoirs. Specifically, module 201 is used because the oil exchange rate and burial rate of CO2 injection development in shale oil reservoirs are affected by a series of factors, including the reservoir's geology and development conditions, as well as the various measures implemented. Based on literature review and the specific conditions of the actual reservoir, the influencing factors and their value ranges are determined.

[0065] In this embodiment, the influencing factors include gas injection volume, gas injection rate, well shut-in time, number of injection cycles, gas injection pressure, porosity multiple, average permeability multiple, and production pressure. For example, the gas injection volume ranges from 2000 to 10000 tons, the gas injection rate ranges from 50 to 90 tons per day, the well shut-in time ranges from 30 to 90 days, the number of injection cycles ranges from 3 to 7 cycles, the gas injection pressure ranges from 160 to 200 bar, the porosity multiple ranges from 0.6 to 1.4, the average permeability multiple ranges from 0.5 to 2.5, and the production pressure ranges from 200 to 300 bar.

[0066] The simulation module 202 is used to sample influencing factors to obtain parameter vectors, and to perform numerical simulations based on these parameter vectors to obtain simulation results. These simulation results include oil production, gas production, and gas injection volume. Specifically, the simulation module 202 is used to perform large-sample random sampling based on a uniform distribution probability density function, which requires a large sample size for analyzing the controlling factors using a learning model.

[0067] In this embodiment, the total number of samples to be sampled is determined, and a sampling step is performed. Specifically, the identified influencing factors are placed into a set, where the number of influencing factors is the total number of samples. This application identifies eight influencing factors, so the total number of samples is eight.

[0068] The sampling step includes: randomly selecting random parameters multiple times from all influencing factors and selecting corresponding values ​​for the random parameters. Specifically, random sampling is performed within the set of influencing factors. For example, non-repeating sampling is used when selecting influencing factors; that is, once an influencing factor is selected, it is removed from the sampling set, and subsequent sampling will not repeat that influencing factor. For the selected influencing factors, it is assumed that they conform to a uniform distribution within the value range, and random values ​​are selected within that range. Alternatively, a random function can be used to select values, as follows: In the formula, x represents the value of the extracted influencing factor. The extracted influencing factor is the minimum value within its range. The maximum value of the extracted influencing factor within its range is r, which is a random number between [0,1].

[0069] A parameter vector is constructed based on the results of multiple samplings. Specifically, multiple samplings are performed within the range of values ​​for the influencing factor, and a parameter vector is constructed based on the results of these multiple samplings. ,in, This is the parameter vector of the extracted influencing factors. This is the sampling result for the nth value.

[0070] The sampling process is repeated multiple times to obtain multiple parameter vectors. Specifically, random sampling is repeated multiple times until all influencing factors are sampled. In this embodiment, eight samplings without replacement are performed on the set of influencing factors, and multiple values ​​are sampled for each of the eight selected influencing factors within their respective value ranges. Based on the sampling results of the multiple values, eight parameter vectors are constructed for each influencing factor.

[0071] In this embodiment, based on the multiple parameter vectors obtained above, numerical simulation is performed using numerical simulation software such as Eclipse (a Java-based extensible development platform) or CMG (a reservoir numerical simulation software) to obtain simulation results. The simulation results include oil production, gas production, and gas injection volume.

[0072] In addition, an extra simulation without CO2 huff and puff was performed to obtain the extraction parameters without CO2 huff and puff, and numerical simulation was conducted to obtain the conventional oil production.

[0073] The calculation module 203 is used to determine the burial rate and oil change rate based on the simulation results, and to establish a comprehensive index for the burial rate and oil change rate. Specifically, the calculation module 203 is used to calculate the burial rate and oil change rate using the following formulas: , In the formula, Indicates the burial rate. GPT represents the oil change rate, GIT represents the gas production rate, GIT represents the gas injection rate, and OPT represents the oil production rate. This indicates the regular oil production volume without CO2 huff and puff. For example... Figure 3 and Figure 4 The figures shown are the relationship between the number of calculation cases and the predicted oil change rate and burial rate.

[0074] The formula for calculating the combined index of retention rate and oil change rate is as follows: In the formula, r represents the comprehensive index. Indicates the burial rate. This indicates the oil change rate.

[0075] Based on the simulation results of each set of parameter vectors using the above method, the corresponding burial rate, oil change rate, and comprehensive index are calculated.

[0076] The verification module 204 is used to perform a cyclical verification step based on the acquired dataset. The dataset includes influencing factors and their corresponding values, burial rates, oil change rates, and comprehensive indicators. The cyclical verification step includes: determining the sample proportion coefficient and grouping the dataset according to the sample proportion coefficient; determining the key hyperparameters of the learning model and using the grouped dataset as samples to train the learning model to obtain the initial prediction results of the controlling factors. Specifically, the verification module 204 is used to perform data standardization on the dataset before grouping. Different units of measurement can affect the results of data analysis. To eliminate the influence of different units of measurement between indicators, data normalization is required to improve the comparability between data indicators and reduce the risk of overlearning by the model.

[0077] The data in the dataset is mapped to the range [0,1] using the following formula: In the formula, This represents the value of the i-th data point in the standardized dataset. This represents the value of the i-th data point obtained through numerical simulation software. and This represents the maximum and minimum values ​​of each group of influencing factors found. Furthermore, penetration rates are generally highly heterogeneous; the data distribution after normalization using the above formula is too concentrated at the two ends of [0,1], leading to unreasonable errors and distortions. Therefore, the following formula is used to standardize the penetration rate data: In the formula, This represents the standardized penetration rate. This represents the i-th value of the penetration rate obtained through numerical simulation software. and This represents the maximum and minimum values ​​of the permeability.

[0078] In this embodiment, the dataset is divided equally into N parts, with the sample proportion coefficient determined to be N parts. One part is used as the test dataset, and the remaining N-1 parts are used as the training dataset.

[0079] It is important to note that the training dataset is a data matrix containing parameters and results, which can be further divided into a parameter matrix and a result vector.

[0080] The training performance of a learned model is evaluated by its root mean square error (RMSE). A smaller RMSSE indicates a more reasonable partitioning of the dataset, while a larger RMSSE indicates a poorer partitioning.

[0081] Key hyperparameters include the minimum number of leaves in the learning model and the number of decision trees.

[0082] In this embodiment, the learning model used is a random forest model. The random forest training model is built using the TreeBagger function provided by Matlab. The minimum number of leaves and the number of decision trees are key factors for each decision tree in the random forest to learn from the training dataset. A larger minimum number of leaves makes it easier for the decision tree to undergo "pruning" during growth, preventing it from growing fully. Conversely, a smaller minimum number of leaves can lead to overlearning of the training dataset by the learning model. Similarly, the number of decision trees directly affects the model's fitting ability (its ability to learn from data) and generalization ability (its predictive ability on new data). More decision trees result in stronger generalization ability but higher computational cost. Fewer decision trees result in weaker generalization ability. Initially, those skilled in the art set an initial minimum number of leaves and a specific number of decision trees for the learning model based on experience.

[0083] The divided training dataset is used as samples and fed into the learning model to obtain initial prediction results. For example... Figure 5 The results shown are the initial predictions, which demonstrate the importance of each influencing factor for CO2 throughput development. The factors with higher importance are the controlling factors.

[0084] The optimization module 205 is used to optimize the learning model and obtain the optimized prediction results of the main controlling factors predicted by the optimized learning model. Specifically, the optimization module 205 is used to optimize the learning model. The optimization adopts a cyclical validation method to optimize hyperparameters such as model dataset partitioning (sample ratio parameter), minimum number of leaves, and number of decision trees, as detailed below.

[0085] The root mean square error (RMSE) of the learning model is calculated, and the sample proportion parameter and key hyperparameter corresponding to the minimum RMSE value are used as the optimization sample proportion parameter and optimization key hyperparameter. Specifically, multiple ordered arrays are created to store the minimum number of leaves, the number of decision trees, and the RMSE. For example, three ordered arrays L, T, and S of length M*N are created to store the minimum number of leaves, the number of decision trees, and the RMSE under the corresponding parameters, respectively. Here, N is the sample proportion parameter, i.e., the average number of divisions in the dataset, and M is the number of possible values ​​for the minimum number of leaves. L The number of parameter values ​​M in the decision tree T The product of, i.e., M=ML *M T .

[0086] Set a set of values ​​for the minimum number of leaves and the number of decision trees. Specifically, set the minimum number of leaves to 1, and select M in increments of 2. L The number of values ​​S represents the set of values ​​for the minimum number of leaves. L Let the decision tree have a base of 5, with the exponent starting from 1 and increasing sequentially, and select M in turn. T The set of possible values ​​S for the number of decision trees T Set S L and set S T After cross-combination, values ​​are selected, that is, S is randomly selected from each combination. L One of the values ​​L i and S T A value T j Combine them and store them in L and T respectively.

[0087] Multiple random values ​​are selected from the set of values, and the results are substituted into the learning model to obtain multiple root mean square errors. In this embodiment, the values ​​of the minimum number of leaves and the number of decision trees are selected. and Substituting these values ​​into the established learning model, we obtain the root mean square error of the learning model under the corresponding parameters. Store in the corresponding sorted array The calculation formula is as follows: In the formula, RMSE represents the root mean square error of the learned model. This represents the true value at the i-th verification point. This represents the predicted value of the learning model at the i-th validation point, and n represents the number of validation points.

[0088] The root mean square error (RMSE) is calculated each time a randomly selected value from the set of key hyperparameters of the learning model is input into the model, until the ordered array S is filled (i.e., increased to M). This process is repeated M times to obtain M RMSE values. The minimum number of leaves and the number of decision trees corresponding to the minimum RMSE are selected as the key hyperparameters for optimizing the learning model.

[0089] In addition, an ordered coefficient array is created to store the sample proportion coefficients. Specifically, an ordered coefficient array is created. ordered coefficient array Used to store sample proportion coefficients.

[0090] The sample proportion coefficient is then arithmetically assigned values, and the results are stored in an ordered coefficient array. Specifically, let... Starting with 2 With a tolerance of 3 , For the number of terms, Perform arithmetic progression and store the values ​​in an ordered coefficient array. In Chinese, the formula is as follows: In the formula, The value of the sample proportion coefficient. As the first item, For tolerance, The number of terms.

[0091] A cyclic verification step is performed based on the ordered coefficient array to determine the root mean square error of the learned model. Specifically, based on the stored ordered coefficient array... The values ​​in the range are used to modify the sample proportion coefficients in turn, and the root mean square error of the learning model after the sample proportion coefficients are modified is calculated. The sample proportion coefficient corresponding to the minimum value of the root mean square error is taken as the optimized sample proportion coefficient.

[0092] Based on the optimized sample proportion parameter and optimized key hyperparameters, a cyclical verification step is performed to obtain the optimized learning model. Specifically, the optimized key hyperparameters are set in the newly established learning model, and the dataset is re-partitioned using the optimized sample proportion coefficient. The learning model is then trained using the optimized training dataset to obtain the optimized learning model.

[0093] The optimized prediction results of the controlling factors predicted using the optimized learning model are as follows: Figure 6 As shown, the importance of each influencing factor for carbon dioxide throughput development is demonstrated, and the factors with higher importance are the main controlling factors.

[0094] Some modules in the apparatus described in this application can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0095] The apparatus or module described in the above embodiments can be implemented by a computer chip or physical entity, or by a product with a certain function. For ease of description, the above apparatus is described by dividing it into various modules according to their functions. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware. Of course, a module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.

[0096] The methods, apparatus, or modules described in this application can be implemented in a computer-readable program code manner. The controller can be implemented in any suitable manner, such as a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of a memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code manner, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included within it for implementing various functions can also be considered as structures within the hardware component. Alternatively, the device used to implement various functions can be viewed as either a software module implementing the method or a structure within a hardware component.

[0097] This application also provides an apparatus, the apparatus comprising: a processor; a memory for storing processor-executable instructions; wherein, when the processor executes the executable instructions, it implements the method described in this application.

[0098] Furthermore, in the various embodiments of the present invention, each functional module can be integrated into a processing module, or each module can exist independently, or two or more modules can be integrated into a single module.

[0099] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, or it can be embodied in the process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0100] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this application can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multiprocessor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.

[0101] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of this application.

Claims

1. A method for determining the main controlling factor of CO2 huff and puff in shale reservoirs, characterized in that, include: Determine the influencing factors of CO2 injection and purge in shale oil reservoirs and the range of values ​​for these factors; A parameter vector is obtained by sampling the influencing factors, and simulation results are obtained by performing numerical simulations based on the parameter vectors; wherein, the simulation results include oil production, gas production, and gas injection; the sampling of the influencing factors to obtain the parameter vector includes: performing non-repeating sampling of the influencing factors, and randomly selecting multiple values ​​for the selected influencing factors within their value range using a random function to construct the parameter vector of the selected influencing factors. ,in, This is the parameter vector of the extracted influencing factors. This is the sampling result of the nth random selection; Based on the simulation results, the burial rate and oil change rate are determined, and a comprehensive index of the burial rate and oil change rate is established; wherein, the calculation formula for the comprehensive index of the burial rate and oil change rate is as follows: In the formula, r represents the comprehensive index. Indicates the burial rate. Indicates the oil change rate; A cyclical verification step is performed based on the acquired dataset; wherein, the dataset includes the influencing factors and the corresponding values, burial rate, oil change rate and comprehensive index of the influencing factors; The cyclic verification step includes: Determine the sample proportion coefficients, and group the dataset according to the sample proportion coefficients; The key hyperparameters of the learning model are determined, and the grouped dataset is used as a sample to train the learning model to obtain the initial prediction results of the main control factor. Optimizing the learning model to obtain the optimized prediction results of the main control factors predicted by the optimized learning model includes: calculating the root mean square error (RMSE) of the learning model respectively; using the sample proportion coefficient corresponding to the minimum value of the RMSE and the key hyperparameter as the optimized sample proportion parameter and optimized key hyperparameter; performing the iterative verification step based on the optimized sample proportion parameter and the optimized key hyperparameter to obtain the optimized learning model; wherein, calculating the RMSE of the learning model includes: the key hyperparameter includes the minimum number of leaves and the number of decision trees; establishing multiple ordered arrays for storing the minimum number of leaves, the number of decision trees, and the RMSE; setting a set of values ​​for the minimum number of leaves and the number of decision trees; randomly selecting values ​​multiple times from the set of values, and substituting the selected values ​​into the learning model to obtain multiple RMSEs.

2. The method according to claim 1, characterized in that, The influencing factors include gas injection volume, gas injection rate, well shut-in time, number of injection cycles, gas injection pressure, porosity factor, average permeability factor, and production pressure.

3. The method according to claim 1, characterized in that, The sampling of the influencing factors to obtain the parameter vector includes: Determine the total number of samples to be sampled and proceed with the sampling steps; The sampling step includes: randomly selecting random parameters from all the influencing factors, and selecting corresponding values ​​for the random parameters; The parameter vector is constructed based on the results of multiple samplings.

4. The method according to claim 3, characterized in that, The process of sampling the influencing factors to obtain the parameter vector further includes: The sampling step is performed multiple times to obtain multiple parameter vectors.

5. The method according to claim 1, characterized in that, After obtaining the simulation results by performing numerical simulation based on the parameter vector, the process further includes: The extraction parameters without CO2 injection are obtained, and numerical simulations are performed to obtain conventional simulation results; wherein, the simulation results include conventional oil production, conventional gas injection, and conventional gas production.

6. The method according to claim 1, characterized in that, The calculation of the root mean square error of the learning model includes: Establish an ordered coefficient array for storing the sample proportion coefficients; The sample proportion coefficient is subjected to arithmetic progression, and the results are stored in the ordered coefficient array. The cyclic verification step is performed based on the ordered coefficient array, and the root mean square error of the learning model is determined.

7. A device for determining the main controlling factors of CO2 huff and puff in shale oil reservoirs, characterized in that, include: The determination module is used to determine the influencing factors of CO2 injection and purge in shale oil reservoirs and the value range of the influencing factors; The simulation module is used to sample the influencing factors to obtain a parameter vector, and perform numerical simulation based on the parameter vector to obtain simulation results; wherein, the simulation results include oil production, gas production, and gas injection volume; the sampling of the influencing factors to obtain the parameter vector includes: performing non-repeating sampling of the influencing factors, and using a random function to randomly select multiple values ​​for the selected influencing factors within their value range to construct the parameter vector of the selected influencing factors. ,in, This is the parameter vector of the extracted influencing factors. This is the sampling result of the nth random selection; The calculation module is used to determine the burial rate and oil change rate based on the simulation results, and to establish a comprehensive index of the burial rate and oil change rate; wherein, the calculation formula for the comprehensive index of the burial rate and oil change rate is as follows: In the formula, r represents the comprehensive index. Indicates the burial rate. Indicates the oil change rate; The verification module is used to perform a cyclic verification step based on the acquired dataset; wherein, the dataset includes the influencing factors and the corresponding values, burial rate, oil change rate and comprehensive index of the influencing factors; The cyclic verification step includes: Determine the sample proportion coefficients, and group the dataset according to the sample proportion coefficients; The key hyperparameters of the learning model are determined, and the grouped dataset is used as a sample to train the learning model to obtain the initial prediction results of the main control factor. An optimization module is used to optimize the learning model to obtain optimized prediction results of the main control factors predicted by the optimized learning model. This includes: calculating the root mean square error (RMSE) of the learning model; using the sample proportion coefficient corresponding to the minimum RMSE and the key hyperparameter as the optimized sample proportion parameter and optimized key hyperparameter; performing the iterative verification step based on the optimized sample proportion parameter and the optimized key hyperparameter to obtain the optimized learning model; wherein calculating the RMSE of the learning model includes: the key hyperparameter includes the minimum number of leaves and the number of decision trees; establishing multiple ordered arrays for storing the minimum number of leaves, the number of decision trees, and the RMSE; setting a set of values ​​for the minimum number of leaves and the number of decision trees; randomly selecting values ​​multiple times from the set of values, and substituting the selected values ​​into the learning model to obtain multiple RMSEs.

8. An apparatus for determining the main controlling factors of CO2 huff and puff in shale oil reservoirs, characterized in that, include: processor; Memory used to store processor-executable instructions; When the processor executes the executable instructions, it implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Sensitivity analysis method for dense reservoir carbon dioxide huff and puff influence factors, and application of sensitivity analysis method

    CN108252688A

  • Method for optimizing CO2 huff-puff cycle injection volume by numerical simulation

    CN112145137A