Oil and gas resource evaluation method and device based on ensemble learning, equipment and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2026-08-11
AI Technical Summary
[0017]本发明实施例的技术方案,通过获取油藏样本数据,对油藏样本数据进行预处理,得到目标油藏样本数据,然后基于目标油藏样本数据构建油气资源评价模型,并基于油气资源评价模型对油气资源进行评价。本技术方案,通过集成学习方法构建油气资源评价模型,能够实现油气资源分布的精细化预测,降低勘探风险,提高资源利用效率,为油气勘探与开发提供科学决策支持。
Smart Images

Figure CN122549698A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oil and gas resource exploration and development technology, and in particular to oil and gas resource evaluation methods, devices, equipment and media based on ensemble learning. Background Technology
[0002] Oil and gas resources are a crucial cornerstone of global economic development, and their exploration and development have always been a research hotspot in the energy sector. With technological advancements, the upstream oil and gas sector has accumulated a vast amount of data. How to effectively process and utilize this data has become a major challenge in the field of oil and gas resource assessment.
[0003] Traditional oil and gas resource evaluation methods mainly include analogy method, volumetric method, material balance method and production decline method.
[0004] These methods each have their limitations at different exploration and development stages. The analogy method, based on regional geological and geophysical data, often presents significant uncertainty in its predictions when statistical and probabilistic data are lacking. The volumetric method, a primary method for calculating oil and gas reservoir geological reserves, heavily relies on the accuracy and completeness of the geological model, posing challenges for oil and gas resource evaluation under complex geological conditions. The mass balance method and the diminishing returns method are more commonly used in the oil and gas reservoir extraction stage, using production data and yield figures to infer dynamic geological reserves; however, they are difficult to apply in the early stages of exploration or development. Summary of the Invention
[0005] This invention provides an oil and gas resource evaluation method, apparatus, equipment, and medium based on ensemble learning. By constructing an oil and gas resource evaluation model through ensemble learning, it is possible to achieve refined prediction of oil and gas resource distribution, reduce exploration risks, improve resource utilization efficiency, and provide scientific decision support for oil and gas exploration and development.
[0006] According to one aspect of the present invention, an oil and gas resource evaluation method based on ensemble learning is provided, the method comprising:
[0007] Obtain reservoir sample data;
[0008] The reservoir sample data is preprocessed to obtain the target reservoir sample data;
[0009] An oil and gas resource evaluation model is constructed based on the target reservoir sample data, and the oil and gas resources are evaluated based on the oil and gas resource evaluation model; wherein, the oil and gas resource evaluation model is an oil and gas resource evaluation model based on ensemble learning.
[0010] According to another aspect of the present invention, an oil and gas resource evaluation device based on ensemble learning is provided, the device comprising:
[0011] The reservoir sample data acquisition module is used to acquire reservoir sample data;
[0012] The target reservoir sample data acquisition module is used to preprocess the reservoir sample data to obtain the target reservoir sample data;
[0013] The oil and gas resource evaluation model construction module is used to construct an oil and gas resource evaluation model based on the target reservoir sample data, and to evaluate the oil and gas resources based on the oil and gas resource evaluation model; wherein, the oil and gas resource evaluation model is an oil and gas resource evaluation model based on ensemble learning.
[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0015] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to execute the oil and gas resource evaluation method based on ensemble learning as described in any embodiment of the present invention.
[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the oil and gas resource evaluation method based on ensemble learning as described in any embodiment of the present invention.
[0017] The technical solution of this invention involves acquiring reservoir sample data, preprocessing the reservoir sample data to obtain target reservoir sample data, constructing an oil and gas resource evaluation model based on the target reservoir sample data, and evaluating oil and gas resources based on the oil and gas resource evaluation model. This technical solution, by constructing an oil and gas resource evaluation model using an ensemble learning method, can achieve refined prediction of oil and gas resource distribution, reduce exploration risks, improve resource utilization efficiency, and provide scientific decision support for oil and gas exploration and development.
[0018] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart of the oil and gas resource evaluation method based on ensemble learning provided in Embodiment 1 of the present invention;
[0021] Figure 2 The flowchart of the oil and gas resource evaluation process based on ensemble learning provided in Embodiment 2 of the present invention;
[0022] Figure 3 A schematic diagram of the oil and gas resource evaluation process based on ensemble learning provided in Embodiment 2 of this application;
[0023] Figure 4 This is a schematic diagram showing the distribution of different feature data provided in Embodiment 2 of this application;
[0024] Figure 5 Statistical results for the feature data types provided in Embodiment 2 of this application;
[0025] Figure 6 This is a schematic diagram illustrating the feature contribution rate analysis of the LightGBM resource evaluation model provided in Embodiment 2 of this application;
[0026] Figure 7 This is a schematic diagram of the structure of the oil and gas resource evaluation device based on ensemble learning provided in Embodiment 3 of the present invention;
[0027] Figure 8 This is a schematic diagram of the structure of an electronic device that implements the oil and gas resource evaluation method based on ensemble learning according to embodiments of the present invention. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0029] It should be noted that the terms "target," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] Example 1
[0031] Figure 1 This is a flowchart of an oil and gas resource evaluation method based on ensemble learning according to Embodiment 1 of the present invention. This embodiment is applicable to the evaluation of oil and gas resources. The method can be executed by an oil and gas resource evaluation device based on ensemble learning, which can be implemented in hardware and / or software and can be configured in an equipment. For example, the equipment can be a backend server or other device with communication and computing capabilities. Figure 1 As shown, the method includes:
[0032] S110. Obtain reservoir sample data.
[0033] In this scheme, reservoir sample data refers to various representative data obtained from the reservoir. Reservoir sample data may include core data, well logging data, seismic data, etc. Preferably, reservoir sample data includes rift basin data.
[0034] In this embodiment, comprehensive reservoir sample data can be collected from multiple sources, including geological exploration and production data. Specifically, reservoir sample data can be collected from aspects such as geological structure, rock type, area, and reserve abundance.
[0035] S120. Preprocess the reservoir sample data to obtain the target reservoir sample data.
[0036] In this embodiment, the purpose of preprocessing the reservoir sample data is mainly to improve the quality of the reservoir sample data and transform it into target reservoir sample data that is more suitable for subsequent analysis and modeling.
[0037] Preprocessing includes data collection and understanding, data cleaning, data integration, data transformation, data dimensionality reduction, feature selection, data balancing and enhancement, etc.
[0038] Optionally, the reservoir sample data is preprocessed to obtain target reservoir sample data, including steps A1-A3:
[0039] Step A1: Perform data cleaning, data standardization, data missing value imputation, and data imbalance processing on the reservoir sample data to obtain the first reservoir sample data;
[0040] In this approach, data cleaning is used to remove outliers and duplicate records from the reservoir sample dataset, thereby ensuring data accuracy.
[0041] Data standardization is a process of converting data of different dimensions or magnitudes into a unified standard. Because the extreme value differences and variances of various features in reservoir sample data are large, it is necessary to standardize numerical features to eliminate dimensional differences between different features.
[0042] In this embodiment, data imputation can be performed using the mean, median, mode, previous and next data, or custom data imputation. When the amount of missing data is large, regression modeling is used for data imputation.
[0043] Furthermore, data imbalance can be addressed using methods such as undersampling, oversampling, and SMOTE (Synthetic Minority Oversampling). When the data is training data, data imbalance is used to balance the data; when the data is test data, data imbalance is used to maintain the original distribution of the data.
[0044] Step A2: Perform feature selection on the first reservoir sample data to obtain the second reservoir sample data;
[0045] In this scheme, feature selection aims to identify key features that have a significant impact on oil and gas resource assessment, thereby reducing model complexity and improving model generalization ability.
[0046] Specifically, feature selection methods can be used to select features from the first reservoir sample data to filter out the second reservoir sample data. These feature selection methods include filtering, wrapping, and embedding. Filtering is based on the properties of the features themselves, such as variance, information gain, and mutual information. Wrapping evaluates the performance of feature subsets using a model, such as recursive feature elimination. Embedding performs feature selection during model training, such as regularization methods.
[0047] Step A3: Encode the second reservoir sample data using features to obtain the target reservoir sample data.
[0048] In this scheme, feature encoding is a process of labeling data so that the model can process it.
[0049] Preprocessing reservoir sample data can improve its quality, making it easier to adapt to the needs of subsequent model training.
[0050] Optionally, feature selection is performed on the first reservoir sample data to obtain the second reservoir sample data, including step B1:
[0051] Step B1: Based on statistical methods and correlation analysis, feature selection is performed on the first reservoir sample data to obtain the second reservoir sample data; wherein, the second reservoir sample data includes the basin area, maximum effective thickness of the net producing layer, maximum net thickness, maximum total thickness, maximum permeability, average porosity, geological age, trap type, lithology, sedimentary environment type, and geological period.
[0052] Among these, the maximum effective thickness of the net producing layer refers to the maximum thickness achievable vertically of the portion of the reservoir layer in an oil and gas reservoir that has industrial exploitation value and can produce movable oil and gas; the maximum net thickness refers to the maximum true thickness of a rock layer after removing any thin layers of other materials that may be intercalated within it; permeability refers to the ability of a rock to allow fluids to pass through under a certain pressure difference, and is a parameter characterizing the ability of soil or rock itself to conduct liquids; porosity refers to the ratio of the sum of the volumes of all pore spaces in a rock sample to the volume of the rock sample; geological age refers to the time and sequence of the formation of rocks and strata at different periods in the Earth's crust; trap type is the main basis for classifying oil and gas reservoir types, reflecting the genesis of oil and gas reservoirs; lithology is a way of classifying rock properties in geology, mainly based on the material composition, structure, and texture of rocks; sedimentary environment is a geomorphic unit with unique physical, chemical, and biological characteristics where sedimentation occurs; and geological period refers to a long period in Earth's history with stratigraphic records.
[0053] In this scheme, statistical methods refer to the methods used to collect, organize, analyze, and interpret statistical data, and to draw certain conclusions about the issues reflected in it. Statistical methods include the chi-square test, Poisson distribution, and mathematical induction. Specifically, the chi-square test can be used for feature selection of the first reservoir sample data.
[0054] In this embodiment, the correlation analysis method is mainly used to identify the interrelationships among factors that affect the evaluation of oil and gas resources.
[0055] Using statistical methods and correlation analysis to select features from the first oil reservoir sample data can improve the accuracy and stability of oil and gas resource evaluation.
[0056] S130. Construct an oil and gas resource evaluation model based on the target reservoir sample data, and evaluate the oil and gas resources based on the oil and gas resource evaluation model; wherein, the oil and gas resource evaluation model is an oil and gas resource evaluation model based on ensemble learning.
[0057] In this scheme, the oil and gas resource evaluation model is used to analyze sample data of the target oil reservoir, thereby realizing the evaluation of oil and gas resources.
[0058] Specifically, the target reservoir sample data is input into the oil and gas resource evaluation model to be trained, and the oil and gas resource evaluation model is trained and updated to obtain the updated oil and gas resource evaluation model.
[0059] By constructing an oil and gas resource evaluation model, deep features can be automatically extracted from massive production data, effectively integrating expert knowledge and data-driven methods, significantly improving the accuracy and efficiency of oil and gas resource potential assessment. This not only provides more scientific decision support for oil and gas exploration and development, but also helps reduce exploration risks, optimize resource allocation, and promote the sustainable development of the oil and gas industry.
[0060] Optionally, an oil and gas resource evaluation model is constructed based on the target reservoir sample data, including steps C1-C3:
[0061] Step C1: Divide the target reservoir sample data into a training set and a test set;
[0062] Specifically, the target reservoir sample data can be divided into a training set and a test set in a 75:25 ratio for training and testing of the oil and gas resource evaluation model.
[0063] Step C2: Train the oil and gas resource evaluation model to be trained based on the training set to obtain the initial oil and gas resource evaluation model; wherein, the oil and gas resource evaluation model to be trained includes the LightGBM model, the CatBoost model, and the XGBoost model;
[0064] In this scheme, XGBoost uses decision trees as base learners and reduces the loss function by iteratively adding new trees. Regularization techniques are used to control model complexity and prevent overfitting. LightGBM is an improved version of XGBoost, featuring high training and prediction speeds, low memory consumption, and improved model accuracy through algorithm optimization and feature selection. CatBoost excels at handling categorical features, allowing direct input of column identifiers for categorical features, with the model automatically performing one-hot encoding or efficient encoding processing.
[0065] In this embodiment, the training set can be input into the oil and gas resource evaluation model to be trained, thereby obtaining the initial oil and gas resource evaluation model. During the training process, strategies such as K-fold cross-validation (5-fold) are used to further verify the stability and generalization ability of the model.
[0066] Step C3: Validate the initial oil and gas resource evaluation model using the test set and obtain the validation results. If the validation results meet the preset constraints, then the initial oil and gas resource evaluation model is adopted as the oil and gas resource evaluation model.
[0067] Among these, constraints refer to the conditions that constrain the performance of the oil and gas resource evaluation model. By constraining the performance of the oil and gas resource evaluation model, the accuracy of model predictions can be improved.
[0068] Specifically, the test set is input into the initial oil and gas resource evaluation model to obtain the validation results. Then, the validation results are compared with the actual results. If the error between the validation results and the actual results is small, that is, if the validation results meet the constraints, the initial oil and gas resource evaluation model is adopted as the oil and gas resource evaluation model. If the error between the validation results and the actual results is large, that is, if the validation results do not meet the constraints, the initial oil and gas resource evaluation model is retrained until the validation results meet the constraints.
[0069] By constructing an oil and gas resource evaluation model, we can achieve refined prediction of oil and gas resource distribution, reduce exploration risks, improve resource utilization efficiency, and provide scientific decision support for oil and gas exploration and development.
[0070] Optionally, the oil and gas resource evaluation model to be trained is trained based on the training set to obtain an initial oil and gas resource evaluation model, including steps D1-D2:
[0071] Step D1: Input the training set into the oil and gas resource evaluation model to be trained and execute the oil and gas resource evaluation task;
[0072] In this embodiment, the oil and gas resource assessment task is used to evaluate oil and gas resources. By controlling the oil and gas resource assessment model to be trained through the training set to perform the oil and gas resource assessment task, it is possible to achieve more refined prediction of oil and gas resource distribution, reduce exploration risks, and improve resource utilization efficiency.
[0073] Step D2: Based on the oil and gas resource evaluation task, adjust the network parameters in the oil and gas resource evaluation model to be trained to obtain the initial oil and gas resource evaluation model after training and updating.
[0074] The network parameters include the learning rate and the loss function. The loss function can be the mean squared error loss function, the regression loss function, etc.
[0075] Specifically, the training set can be used as input to perform oil and gas resource evaluation tasks through the oil and gas resource evaluation model to be trained, thereby determining the loss function value corresponding to the oil and gas resource evaluation task. Based on the loss function value, the network parameters of the oil and gas resource evaluation model to be trained are adjusted to optimize the loss function, thus obtaining the initial oil and gas resource evaluation model after training and updating.
[0076] Among these methods, grid search optimization is used to fine-tune the network parameters of the model in order to improve prediction accuracy and training speed.
[0077] By constructing oil and gas resource evaluation models using ensemble learning methods, we can achieve refined prediction of oil and gas resource distribution, reduce exploration risks, and improve resource utilization efficiency.
[0078] The technical solution of this invention involves acquiring reservoir sample data, preprocessing the reservoir sample data to obtain target reservoir sample data, constructing an oil and gas resource evaluation model based on the target reservoir sample data, and evaluating oil and gas resources based on the oil and gas resource evaluation model. By implementing this technical solution and constructing an oil and gas resource evaluation model using an ensemble learning method, it is possible to achieve refined prediction of oil and gas resource distribution, reduce exploration risks, improve resource utilization efficiency, and provide scientific decision support for oil and gas exploration and development.
[0079] Example 2
[0080] Figure 2 This is a flowchart of the oil and gas resource evaluation process based on ensemble learning provided in Embodiment 2 of the present invention. This embodiment supplements the above embodiments in its role in constructing the oil and gas resource evaluation model. Figure 2 As shown, the method includes:
[0081] S210. Obtain reservoir sample data.
[0082] S220. Conduct correlation analysis on the variables in the reservoir sample data to obtain the correlation coefficients; wherein, the variables are the characteristics in the reservoir sample data that affect the evaluation of oil and gas resources.
[0083] In this scheme, correlation coefficients are generally statistical indicators used to measure the degree of linear relationship between two or more oil-related variables. Correlation coefficients include Pearson correlation coefficient and Spearman's rank correlation coefficient.
[0084] Specifically, correlation analysis can be used to analyze the correlation between variables in the reservoir sample data and calculate the correlation coefficient between the variables in the reservoir sample data.
[0085] Identifying the interrelationships among factors influencing oil and gas resource assessment is crucial for subsequent data preprocessing and model building. These correlations can be used to select features, thereby improving the accuracy and stability of the model.
[0086] S230. Display the reservoir sample data and the correlation coefficient.
[0087] Furthermore, descriptive charts are used to display the basic characteristics and distribution of reservoir sample data. These basic characteristics include the data range (maximum and minimum values), central tendency (mean, median, etc.), and dispersion. In reservoir sample data for oil and gas resource evaluation, these basic characteristics help to initially understand the overall situation and distribution patterns of the data. The basic characteristics and distribution of data are typically displayed using descriptive charts (such as histograms), which can visually show information such as the central tendency, dispersion, and distribution pattern of the data. The correlation coefficients between variables are mainly represented by charts such as heatmaps. Heatmaps use color intensity to indicate the magnitude of the correlation coefficients between variables, thus visually demonstrating the strength of the correlation between them.
[0088] By displaying these charts and tables in the reservoir sample data for oil and gas resource assessment, we can gain a deeper understanding of the interactions and influences between various variables, providing strong support for subsequent data preprocessing and model building.
[0089] In this plan, Figure 3 This is a schematic diagram of the oil and gas resource evaluation process based on ensemble learning provided in Embodiment 2 of this application, as shown below. Figure 3 As shown, exploratory analysis of reservoir sample data was conducted to gain a deeper understanding of data characteristics and distribution patterns.
[0090] S240. Preprocess the reservoir sample data to obtain the target reservoir sample data.
[0091] In this embodiment, as Figure 3 As shown, the collected reservoir sample data are preprocessed to ensure data quality.
[0092] S250. Construct an oil and gas resource evaluation model based on the target reservoir sample data, and evaluate the oil and gas resources based on the oil and gas resource evaluation model; wherein, the oil and gas resource evaluation model is an oil and gas resource evaluation model based on ensemble learning.
[0093] Furthermore, such as Figure 3 As shown, training and testing sample datasets are created based on the preprocessed target reservoir sample data to provide a solid data foundation for the model. Then, an ensemble learning algorithm is used to construct an oil and gas resource evaluation model to achieve a more robust and accurate assessment of oil and gas resource potential.
[0094] S260. The performance of the oil and gas resource evaluation model is evaluated based on the evaluation indicators; wherein the evaluation indicators include mean square error, root mean square error, mean absolute error, mean absolute percentage error, and coefficient of determination.
[0095] In this scheme, MSE (Mean Squared Error), RMSE (Root Mean Squared Error), MAE (Mean Absolute Error), MAPE (Mean Absolute Percentage Error), and R are used. 2 Coefficient of Determination (CDE) and other similar metrics are key evaluation indicators for oil and gas resource assessment models. These indicators quantify the predictive performance of the model, thereby verifying its effectiveness and reliability.
[0096] Specifically, the mean squared error measures the average of the squared differences between the predicted and actual values of an oil and gas resource assessment model, and is highly sensitive to outliers. The calculation formula is as follows:
[0097]
[0098] Where n is the number of reservoir samples, y is the actual value of the reservoir abundance of the i-th reservoir sample. i It is the predicted value of the reserve abundance of the i-th reservoir sample.
[0099] In this embodiment, the root mean square error (RMSE) is the square root of the mean square error (MSE), which maintains the same dimensions as the original data and more intuitively reflects the accuracy of the oil and gas resource assessment model's predictions. The calculation formula is as follows:
[0100]
[0101] In this scheme, the mean absolute error (MAE) measures the average of the absolute values of the differences between the model's predicted values and the actual values. The calculation formula is as follows:
[0102]
[0103] Mean absolute percentage error (MAPE) is another metric for measuring prediction accuracy; it expresses the error as a percentage of the actual value. The calculation formula is as follows:
[0104]
[0105] In this scheme, the coefficient of determination, also known as the goodness of fit, represents the proportion of data variability that the model can explain. R 2 The closer the value of R is to 1, the better the fit of the oil and gas resource evaluation model; if R... 2 A value less than 0 indicates that the model performs even worse than simply taking the average of the actual values. The calculation formula is as follows:
[0106]
[0107] in, It is the average of the actual values of the reservoir abundance of the oil reservoir sample.
[0108] In this embodiment, as Figure 3 As shown, the performance of the constructed oil and gas resource evaluation model is evaluated, and the effectiveness and reliability of the model are verified by comparing the prediction results with the actual data.
[0109] The technical solution of this invention involves acquiring reservoir sample data, performing correlation analysis on the variables in the reservoir sample data to obtain correlation coefficients, and then displaying the reservoir sample data and correlation coefficients. Next, the reservoir sample data is preprocessed to obtain target reservoir sample data. Based on the target reservoir sample data, an oil and gas resource evaluation model is constructed, and the oil and gas resources are evaluated based on this model. Finally, the performance of the oil and gas resource evaluation model is evaluated according to assessment indicators. By implementing this technical solution and constructing an oil and gas resource evaluation model using an ensemble learning method, it is possible to achieve refined prediction of oil and gas resource distribution, reduce exploration risks, improve resource utilization efficiency, and provide scientific decision support for oil and gas exploration and development.
[0110] In this scheme, the oil and gas resource assessment dataset aggregates over 200,000 detailed reservoir records within the rift basin. Each record covers 15 key geological parameters, including geological structural features, stratigraphic thickness variations, and rock type diversity. The dataset not only spatially spans multiple regions within the rift basin, achieving broad geographical coverage, but also includes exploration data from different stratigraphic levels, ensuring the comprehensiveness and representativeness of stratigraphic distribution. This makes it highly valuable for oil and gas resource potential assessment based on ensemble learning techniques and possesses significant practical application potential.
[0111] Furthermore, Figure 4 This is a schematic diagram illustrating the distribution of different feature data provided in Embodiment 2 of this application, such as... Figure 4 As shown, the distribution of four different characteristic data is displayed. The distribution plot can intuitively show the central location of the data, and these indicators help to summarize the central trend of the data. Figure 5 Statistical results for the feature data types provided in Embodiment 2 of this application, such as Figure 5As shown, the proportion of shallow marine is as high as 40%, indicating an imbalance in the data for this characteristic.
[0112] In this embodiment, Table 1 shows a performance comparison of the three algorithms, LightGBM, CatBoost, and XgBoost, on the same input dataset. It is clear that LightGBM performs better than XgBoost in R... 2 It achieved the best performance in terms of score (0.8534), with the highest good fit between predicted and actual values and the highest model prediction accuracy. LightGBM also boasts the fastest training time (404 seconds) due to its efficient computational performance. In comparison, CatBoost's R... 2 The score was lower (0.7658) and the training time was the longest (556 seconds), while XgBoost, although R... 2 The score (0.8256) is close to LightGBM, but it has the longest training time (1070 seconds) and is less efficient. In conclusion, LightGBM outperforms CatBoost and XgBoost in both prediction accuracy and training efficiency, making it the best performing of the three algorithms.
[0113] Table 1
[0114]
[0115] Furthermore, Table 2 shows the performance of the LightGBM resource assessment model in predicting reservoir abundance on the training and test sets. The results indicate that the model outperforms the training set on the test set, particularly in R... 2 The score (increased from 0.759 to 0.853) and error metrics (MSE, RMSE, MAE, and MAPE all decreased) showed good generalization ability.
[0116] Table 2
[0117] MSE RMSE MAE MAPE <![CDATA[R 2 ]]> training set 26.347 5.133 3.999 6.093 0.759 test set 19.017 4.361 3.51 5.586 0.835
[0118] In this plan, Figure 6 This is a schematic diagram illustrating the feature contribution rate analysis of the LightGBM resource evaluation model provided in Embodiment 2 of this application, as shown below. Figure 6 As shown, the Field_sqkm (area) feature contributes the most to the model's prediction results. This finding aligns well with the business context, as resource abundance is often directly related to the size of the area in which it is located; larger areas may contain more resource reserves or development potential. This analysis not only validates the model's effectiveness but also provides strong data support for subsequent business decisions and resource management.
[0119] Example 3
[0120] Figure 7 This is a schematic diagram of the oil and gas resource evaluation device based on ensemble learning provided in Embodiment 3 of the present invention. Figure 7 As shown, the device includes:
[0121] The reservoir sample data acquisition module 710 is used to acquire reservoir sample data;
[0122] The target reservoir sample data acquisition module 720 is used to preprocess the reservoir sample data to obtain the target reservoir sample data;
[0123] The oil and gas resource evaluation model construction module 730 is used to construct an oil and gas resource evaluation model based on the target reservoir sample data, and to evaluate the oil and gas resources based on the oil and gas resource evaluation model; wherein, the oil and gas resource evaluation model is an oil and gas resource evaluation model based on ensemble learning.
[0124] Optionally, the target reservoir sample data acquisition module 720 includes:
[0125] The first reservoir sample data acquisition unit is used to perform data cleaning, data standardization, data missing value imputation and data imbalance processing on the reservoir sample data to obtain the first reservoir sample data.
[0126] The second reservoir sample data acquisition unit is used to perform feature selection on the first reservoir sample data to obtain the second reservoir sample data.
[0127] The target reservoir sample data acquisition unit is used to perform feature encoding on the second reservoir sample data to obtain the target reservoir sample data.
[0128] Optionally, the second reservoir sample data acquisition unit is specifically used for:
[0129] Based on statistical methods and correlation analysis, feature selection was performed on the first reservoir sample data to obtain the second reservoir sample data. The second reservoir sample data includes the basin area, maximum effective thickness of the net producing layer, maximum net thickness, maximum total thickness, maximum permeability, average porosity, geological age, trap type, lithology, sedimentary environment type, and geological period.
[0130] Optional, the oil and gas resource evaluation model construction module 730 includes:
[0131] The training set and test set partitioning unit is used to partition the target reservoir sample data into a training set and a test set;
[0132] The initial oil and gas resource evaluation model is obtained by using the training set to train the oil and gas resource evaluation model to be trained, thereby obtaining the initial oil and gas resource evaluation model; wherein, the oil and gas resource evaluation model to be trained includes the LightGBM model, the CatBoost model, and the XGBoost model.
[0133] The oil and gas resource evaluation model determination unit is used to verify the initial oil and gas resource evaluation model using the test set and obtain the verification result. If the verification result meets the preset constraints, the initial oil and gas resource evaluation model is used as the oil and gas resource evaluation model.
[0134] Optionally, the initial oil and gas resource evaluation model yields cells, specifically used for:
[0135] The training set is input into the oil and gas resource evaluation model to be trained to perform the oil and gas resource evaluation task.
[0136] Based on the oil and gas resource evaluation task, the network parameters in the oil and gas resource evaluation model to be trained are adjusted to obtain the initial oil and gas resource evaluation model after training and updating.
[0137] Optionally, the device further includes:
[0138] The performance evaluation module is used to evaluate the performance of the oil and gas resource evaluation model based on evaluation indicators; wherein, the evaluation indicators include mean square error, root mean square error, mean absolute error, mean absolute percentage error, and coefficient of determination.
[0139] Optionally, the device further includes:
[0140] The correlation coefficient acquisition module is used to perform correlation analysis on various variables in the reservoir sample data and obtain the correlation coefficient; wherein, the variables are the characteristics in the reservoir sample data that affect the evaluation of oil and gas resources;
[0141] The display module is used to display the reservoir sample data and the correlation coefficient.
[0142] The oil and gas resource evaluation device provided in this embodiment of the invention can execute an oil and gas resource evaluation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0143] Example 4
[0144] Figure 8A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0145] like Figure 8 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0146] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0147] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as oil and gas resource evaluation methods based on ensemble learning.
[0148] In some embodiments, the ensemble learning-based oil and gas resource assessment method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the ensemble learning-based oil and gas resource assessment method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to execute the ensemble learning-based oil and gas resource assessment method by any other suitable means (e.g., by means of firmware).
[0149] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0150] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0151] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0152] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0153] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0154] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0155] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0156] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for oil and gas resource assessment based on ensemble learning, characterized in that, include: Obtain reservoir sample data; The reservoir sample data is preprocessed to obtain the target reservoir sample data; An oil and gas resource evaluation model is constructed based on the target reservoir sample data, and the oil and gas resources are evaluated based on the oil and gas resource evaluation model; wherein, the oil and gas resource evaluation model is an oil and gas resource evaluation model based on ensemble learning.
2. The method according to claim 1, characterized in that, The reservoir sample data is preprocessed to obtain target reservoir sample data, including: The reservoir sample data is cleaned, standardized, filled with missing data, and imbalanced to obtain the first reservoir sample data. Feature selection is performed on the first reservoir sample data to obtain the second reservoir sample data; The second reservoir sample data is feature-encoded to obtain the target reservoir sample data.
3. The method according to claim 2, characterized in that, Feature selection is performed on the first reservoir sample data to obtain the second reservoir sample data, including: Based on statistical methods and correlation analysis, feature selection was performed on the first reservoir sample data to obtain the second reservoir sample data. The second reservoir sample data includes the basin area, maximum effective thickness of the net producing layer, maximum net thickness, maximum total thickness, maximum permeability, average porosity, geological age, trap type, lithology, sedimentary environment type, and geological period.
4. The method according to claim 1, characterized in that, Based on the target reservoir sample data, an oil and gas resource evaluation model is constructed, including: The target reservoir sample data is divided into a training set and a test set; The oil and gas resource evaluation model to be trained is trained based on the training set to obtain an initial oil and gas resource evaluation model; wherein, the oil and gas resource evaluation model to be trained includes the LightGBM model, the CatBoost model, and the XGBoost model; The initial oil and gas resource evaluation model is validated using the test set to obtain validation results. If the validation results meet the preset constraints, the initial oil and gas resource evaluation model is adopted as the oil and gas resource evaluation model.
5. The method according to claim 4, characterized in that, The initial oil and gas resource evaluation model is obtained by training the model based on the training set, including: The training set is input into the oil and gas resource evaluation model to be trained to perform the oil and gas resource evaluation task. Based on the oil and gas resource evaluation task, the network parameters in the oil and gas resource evaluation model to be trained are adjusted to obtain the initial oil and gas resource evaluation model after training and updating.
6. The method according to claim 1, characterized in that, After constructing an oil and gas resource evaluation model based on the target reservoir sample data, the method further includes: The performance of the oil and gas resource evaluation model is evaluated based on the evaluation indicators, which include mean square error, root mean square error, mean absolute error, mean absolute percentage error, and coefficient of determination.
7. The method according to claim 1, characterized in that, After acquiring reservoir sample data, the method further includes: Correlation analysis was performed on the variables in the reservoir sample data to obtain the correlation coefficients; wherein, the variables are the characteristics in the reservoir sample data that affect the evaluation of oil and gas resources. The reservoir sample data and the correlation coefficients are displayed.
8. An oil and gas resource evaluation device based on ensemble learning, characterized in that, include: The reservoir sample data acquisition module is used to acquire reservoir sample data; The target reservoir sample data acquisition module is used to preprocess the reservoir sample data to obtain the target reservoir sample data; The oil and gas resource evaluation model construction module is used to construct an oil and gas resource evaluation model based on the target reservoir sample data, and to evaluate the oil and gas resources based on the oil and gas resource evaluation model; wherein, the oil and gas resource evaluation model is an oil and gas resource evaluation model based on ensemble learning.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the oil and gas resource evaluation method based on ensemble learning as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the oil and gas resource evaluation method based on ensemble learning as described in any one of claims 1-7.