Method and apparatus for determining hydrocarbon parameters
An oil and gas parameter recommendation system was constructed using the Gaussian kernel KNN algorithm, which solved the problem of missing parameters in oil and gas projects, achieved efficient and accurate evaluation of oil and gas parameters, improved the accuracy of oil and gas reserve and production prediction, and supported oil and gas field development and investment decisions.
Patent Information
- Application Number
- CN202411059902.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-02
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2044-08-02
AI Technical Summary
In oil and gas project transactions, due to problems such as incomplete oil and gas field data, inconsistent formats, and missing parameters, existing technologies are unable to achieve efficient and accurate evaluation of oil and gas parameters. Moreover, they rely on the experience and judgment of petroleum geology and oil and gas reservoir engineering experts, resulting in high deviation of results, long processing time, and significant influence of subjective factors.
We employ a Gaussian kernel-based KNN algorithm to screen training samples through correlation analysis, and use the Gaussian kernel distance correlation relationship and KNN algorithm to determine missing data values, thereby constructing an oil and gas parameter recommendation system to improve the accuracy and efficiency of parameter determination.
It improves the accuracy and efficiency of oil and gas parameter determination, enhances the effectiveness and reliability of new oil and gas project evaluation, improves the accuracy of oil and gas reserve and production forecasting, and supports oil and gas field development and investment decisions.
Smart Images

Figure CN119167011B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oil and gas project geology and oil and gas reservoir evaluation technology, and in particular to a method and apparatus for determining oil and gas parameters. Background Technology
[0002] This section is intended to provide background or context for embodiments of the present invention. The description herein is not intended to imply that it is prior art simply because it is included in this section.
[0003] When oil and gas companies conduct oil and gas project transactions, they need to assess and evaluate target blocks and oil and gas fields. During the technical and economic evaluation of new oil and gas projects, various problems arise, including differing levels of exploration in the acquired blocks, incomplete oil and gas field data, inconsistent data formats and standards, missing data items, and difficulties in data interpretation. Among all these problems, missing parameters have a significant impact on the evaluation. Currently, missing parameters are mainly judged based on the experience of petroleum geology and reservoir engineering experts. This process is not only time-consuming and prone to deviation from reality, but also highly susceptible to subjective factors, making it difficult to meet the needs of efficient and accurate evaluation. Summary of the Invention
[0004] This invention provides a method for determining oil and gas parameters to improve the accuracy and efficiency of determining missing oil and gas parameters. The method includes:
[0005] Acquire various historical oil and gas parameter data and samples to be determined; the samples to be determined are real-time oil and gas parameter data with at least one missing data value;
[0006] Based on the sample to be determined, a correlation analysis algorithm is used to screen the first oil and gas parameter data from a variety of historical oil and gas parameter data, and the first oil and gas parameter data is determined as the training sample; the first oil and gas parameter data is a variety of oil and gas parameter data whose correlation with the sample to be determined is higher than a preset correlation threshold.
[0007] Labels are added to the training samples based on the categories of the data in the training samples;
[0008] Using a pre-defined Gaussian kernel distance correlation, the distance between the sample to be determined and the training samples after adding labels is calculated;
[0009] The first preset number of training samples sorted by distance from smallest to largest are determined as neighbor samples;
[0010] Based on neighbor samples, the sample to be determined, and the preset Gaussian kernel standard deviation, the KNN (K Nearest Neighbors) algorithm is used to determine the missing data values in the sample to be determined.
[0011] This invention also provides an oil and gas parameter determination device to improve the accuracy and efficiency of determining missing oil and gas parameters. The device includes:
[0012] The acquisition module is used to acquire various historical oil and gas parameter data and samples to be determined; the samples to be determined are real-time oil and gas parameter data with at least one missing data value.
[0013] The training sample determination module is used to select first oil and gas parameter data from a variety of historical oil and gas parameter data based on the sample to be determined, using a correlation analysis algorithm, and determine the first oil and gas parameter data as the training sample; the first oil and gas parameter data is a variety of oil and gas parameter data whose correlation with the sample to be determined is higher than a preset correlation threshold;
[0014] The labeling module is used to add labels to training samples based on the categories of data in the training samples;
[0015] The distance determination module is used to calculate the distance between the sample to be determined and the labeled training samples using a preset Gaussian kernel distance correlation relationship.
[0016] The neighbor sample determination module is used to determine the first preset number of training samples sorted by distance from smallest to largest as neighbor samples.
[0017] The missing data determination module is used to determine the missing data values in the sample to be determined using the KNN algorithm based on neighbor samples, the sample to be determined, and the preset Gaussian kernel standard deviation.
[0018] Compared with existing technologies where petroleum geology and reservoir engineering experts rely on experience-based judgment, this invention improves the accuracy and efficiency of determining missing oil and gas parameters by acquiring multiple historical oil and gas parameter data and samples to be determined. The samples to be determined are real-time oil and gas parameter data with at least one missing data value. Based on these samples, a correlation analysis algorithm is used to filter out first oil and gas parameter data from the various historical oil and gas parameter data, identifying this first oil and gas parameter data as training samples. The first oil and gas parameter data consists of multiple oil and gas parameter data whose correlation with the samples to be determined is higher than a preset correlation threshold. Labels are added to the training samples based on the data categories within them. The distance between the samples to be determined and the labeled training samples is calculated using a preset Gaussian kernel distance correlation relationship. The first preset number of training samples, sorted by distance from smallest to largest, are identified as neighbor samples. Based on the neighbor samples, the samples to be determined, and the preset Gaussian kernel standard deviation, the KNN algorithm is used to determine the missing data values in the samples to be determined. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0020] Figure 1 This is a flowchart of a method for determining oil and gas parameters provided in an embodiment of the present invention;
[0021] Figure 2 A flowchart illustrating a specific example of a method for determining oil and gas parameters provided in this embodiment of the invention;
[0022] Figure 3 This is a schematic diagram of the standardized training samples in dimensional space provided in this embodiment of the invention. Figure 1 ;
[0023] Figure 4 This is a schematic diagram of the standardized training samples in dimensional space provided in this embodiment of the invention. Figure 2 ;
[0024] Figure 5 This is a flowchart of a KNN algorithm calculation provided in an embodiment of the present invention;
[0025] Figure 6 This is a schematic diagram of an oil and gas parameter determination device provided in an embodiment of the present invention;
[0026] Figure 7 This is a schematic diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0028] The acquisition, storage, use, and processing of data in this application all comply with the relevant provisions of national laws and regulations.
[0029] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0030] In the description of this specification, the terms "comprising," "including," "having," and "containing" are open-ended terms, meaning that they include but are not limited to. The terms "an embodiment," "a specific embodiment," "some embodiments," and "for example," etc., refer to specific features, structures, or characteristics described in connection with that embodiment or example that are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. The order of steps involved in the various embodiments is used to illustrate the implementation of this application, and the order of steps is not limited and can be adjusted appropriately as needed.
[0031] The purpose of this invention is to provide an oil and gas parameter recommendation method based on the Gaussian kernel KNN algorithm. This method provides resource and reserve evaluation parameters such as favorable area, reservoir thickness, resource abundance, and reservoir lithology for oil and gas exploration, and development parameters such as porosity, permeability, oil saturation, effective thickness, and geosaturation pressure difference for oil and gas field development. This significantly improves the accuracy of oil and gas reserve and production prediction, enhances the objectivity and reliability of oil and gas field economic evaluation, and improves the reliability and scientific nature of investment decisions.
[0032] When oil and gas companies conduct oil and gas project transactions, they need to assess and evaluate target blocks and oil and gas fields. During the techno-economic evaluation of new oil and gas projects, various problems arise, including differing levels of exploration in the acquired blocks, incomplete oil and gas field data, inconsistent data formats and standards, missing data items, and difficulties in data interpretation. Among all these problems, missing parameters have a significant impact on the evaluation. Currently, the experience-based judgments of petroleum geology and reservoir engineering experts are not only time-consuming and prone to deviations from reality, but also heavily influenced by subjective factors, making it difficult to meet the needs of efficient and accurate evaluation. Therefore, how to utilize modern computing technology to improve the accuracy of parameter prediction has become a critical technical problem that urgently needs to be solved.
[0033] In current data analysis and forecasting, to roughly estimate an individual's monthly income, one can find individuals in a database with similar family backgrounds, education levels, work locations, job content, and years of employment. By weighting the monthly income of these individuals according to the importance of each condition, a rough estimate of the individual's monthly income range can be obtained. This calculation, based on real-world data, is far more reliable than subjective guesses. Similarly, based on historical experience, in new oil and gas resource assessments, especially when evaluating reservoirs with similar genesis, sedimentary strata, and development methods, their relevant parameters also provide a certain degree of reference value.
[0034] This invention is based on the above ideas and proposes a method for recommending parameter values based on the nearest neighbor algorithm, using the KNN algorithm based on Gaussian kernels in artificial intelligence as its core. This algorithm, based on a large amount of existing historical data, first identifies parameters that are correlated with the missing parameters and substitutes them into the model to construct a feature space. Then, it assigns the majority of the K most similar (i.e., nearest neighbors) samples in the feature space to a certain category. Finally, it uses the Gaussian kernel function to calculate the weight of each feature (parameter) in this category and provides recommended values based on the weights from parameters of the same category.
[0035] Figure 1 This is a flowchart of a method for determining oil and gas parameters provided in an embodiment of the present invention, such as... Figure 1 As shown, the method may include:
[0036] Step 101: Obtain various historical oil and gas parameter data and samples to be determined; the samples to be determined are real-time oil and gas parameter data with at least one missing data value.
[0037] Step 102: Based on the sample to be determined, use a correlation analysis algorithm to screen the first oil and gas parameter data from a variety of historical oil and gas parameter data, and determine the first oil and gas parameter data as the training sample; the first oil and gas parameter data is a variety of oil and gas parameter data whose correlation with the sample to be determined is higher than a preset correlation threshold.
[0038] Step 103: Add labels to the training samples according to the categories of the data in the training samples;
[0039] Step 104: Calculate the distance between the sample to be determined and the training samples after adding labels using the preset Gaussian kernel distance correlation relationship;
[0040] Step 105: The first preset number of training samples sorted by distance from smallest to largest are determined as neighbor samples;
[0041] Step 106: Based on the neighbor samples, the sample to be determined, and the preset Gaussian kernel standard deviation, use the KNN algorithm to determine the missing data values in the sample to be determined.
[0042] When conducting oil and gas project evaluations, due to limited exploration of the evaluation target or limitations in data acquisition sources, some parameters are often missing, and even for parameters that have been obtained, it is impossible to determine the types of parameters that need to be provided. During the evaluation process, it is often necessary to infer and determine the possible values of the missing qualitative parameters based on the existing parameters.
[0043] This method employs a widely used supervised learning approach in modern artificial intelligence: the Gaussian kernel-based KNN algorithm. It proposes a parameter recommendation method based on the nearest neighbor algorithm to address the problem of missing parameter inference. By applying the Gaussian kernel-based KNN algorithm to data analysis in oil and gas exploration projects, a parameter recommendation system based on geological features and known oil and gas reservoir distribution is constructed, effectively improving the effectiveness and accuracy of new oil and gas project evaluation.
[0044] Figure 2 A flowchart illustrating a specific example of a method for determining oil and gas parameters provided in this embodiment of the invention is shown below. Figure 2 As shown, the oil and gas exploration data analysis workflow involves the following steps: Data set establishment: Collect basic data, including geological and geophysical data such as seismic attributes, lithology, porosity, and permeability, to form a data set, providing the data foundation for the method; Data standardization: Perform standardization processing to ensure that the data can be compared. Standardization processing typically includes normalization, scaling, etc., to make the data on the same scale for subsequent calculations; Data numerical transformation: Convert the data's type values, Boolean values (yes / no, male / female), location, geological layers, and other information into numerical values; Model construction: Select data items closely related to the recommended parameters as data dimensions to build the model. Feature selection can be performed through correlation analysis, analysis of variance, chi-square test, etc.; Model optimization: Check the completeness of dimensional data and data item data in the model, and remove data without significance; K-nearest neighbor calculation: For each area to be explored, calculate its Euclidean distance to historically known oil and gas reservoir samples. Select the K nearest samples as its K nearest neighbors; Recommendation parameter value calculation: Use the Gaussian kernel function calculation results for ranking evaluation. The Gaussian kernel function can effectively handle nonlinear problems by mapping the data to a high-dimensional feature space.
[0045] Specifically, the process involves seven steps: First, integrating multi-source geological and geophysical data, including seismic attributes, lithology, porosity, and permeability, to form a comprehensive database. Second, data standardization, such as normalization and scaling, is performed to ensure the accuracy of cross-data comparisons. Third, the data undergoes numerical conversion, encoding various categories and Boolean information into digital formats. Fourth, key variables are selected through correlation analysis to build a model, and feature selection is optimized using analysis of variance and chi-square tests. Fifth, the model is reviewed, and insignificant features are removed for simplification. Sixth, the K-nearest neighbor method is applied to calculate the Euclidean distance between each exploration area and known oil and gas reservoir samples, selecting the K nearest samples as neighbor references. Seventh, a Gaussian kernel function is adopted for ranking and evaluation, leveraging its advantage in handling nonlinearity to map the data to a high-dimensional space, deeply exploring the potential of oil and gas reservoirs, and achieving an efficient recommendation and evaluation system.
[0046] Each step will be described in detail below:
[0047] The entire algorithm is built on a large amount of data collected in real evaluation processes. This basic data includes geological and geophysical data, such as seismic properties, lithology, porosity, and permeability. These data constitute a dataset that provides the data foundation for this method.
[0048] In one embodiment, after acquiring various historical oil and gas parameter data, the method may further include: standardizing the various historical oil and gas parameter data using a normalization algorithm, a unit conversion algorithm, and a classification algorithm, respectively.
[0049] First, oil and gas parameter data, including historical oil and gas parameter data, needs to be standardized to ensure comparability between different characteristics. During data preprocessing, these basic data come from different sources and require standardization to ensure comparability. Standardization typically includes steps such as normalization, unit conversion, and classification unification to ensure the data is on the same standard for subsequent calculations. The basic data types differ, including numbers, types, Boolean values (yes / no, male / female), location, geological layers, etc. To ensure comparability, all this data needs to be converted into comparable numbers, such as converting addresses to postal codes and whether a value is 1 or 0.
[0050] In this step, all oil and gas field parameters obtained from the new project data are taken as the research object. The parameters of historical oil and gas field projects, such as burial depth, permeability, thickness, porosity, saturation, etc., are sorted out and entered into the oil and gas field evaluation model.
[0051] In one embodiment, in step 102, based on the sample to be determined, a correlation analysis algorithm is used to screen first oil and gas parameter data from multiple historical oil and gas parameter data, and the first oil and gas parameter data is determined as a training sample. This includes: based on the sample to be determined, using a correlation analysis algorithm, determining the correlation between multiple historical oil and gas parameter data and the sample to be determined; determining the data completeness of multiple oil and gas parameter data whose correlation is higher than a preset correlation threshold; determining the oil and gas parameter data whose data completeness is higher than a preset data completeness threshold as the first oil and gas parameter data; and determining the first oil and gas parameter data as a training sample.
[0052] In one embodiment, the correlation analysis algorithm includes an analysis of variance algorithm and / or a chi-square test algorithm.
[0053] In step 102, for different oil and gas reservoir project data, based on domain knowledge, data items closely related to the recommended parameters are selected as data dimensions to construct the model. This helps reduce the curse of dimensionality and improves model efficiency. Feature selection can be performed through correlation analysis, analysis of variance, chi-square test, etc.
[0054] After selecting data items closely related to the recommended parameters, it's necessary to check the data completeness in the database. If the completeness of a dimension is below 30% across all data, indicating significant data loss, then that dimension is unreliable and its data needs to be removed from the model. If a data item is missing from the recommended parameter column, then that data item is invalid and needs to be deleted.
[0055] By selecting data items closely related to the recommended parameters based on domain knowledge, the curse of dimensionality is reduced, and model efficiency is improved. The selected features are then fed into the model to complete model construction.
[0056] In this step, firstly, based on geological theory and prior knowledge, features that are obviously unrelated to hydrocarbon accumulation are excluded. Statistical methods are used to evaluate the correlation between each feature and the presence or absence of hydrocarbons, and features with low correlation are eliminated. Finally, the set of selected features such as permeability, thickness, porosity, and saturation is submitted to geological experts for review to ensure that the feature selection is in line with geological principles and can effectively support the prediction objectives of the Gaussian kernel KNN artificial intelligence algorithm.
[0057] Therefore, four parameters were extracted from 20 data points in this experiment, namely burial depth, permeability, thickness, and porosity, as training samples. The data are shown in Table 1 below. The permeability, thickness, and porosity were normalized using the LN function.
[0058] Table 1
[0059] Serial Number Burial depth Penetration thickness Porosity 1 2000 7.13089883 4.605170186 3.135494216 2 4250 6.224558429 2.708050201 2.397895273 3 1750 7.016609684 3.295836866 3.020424886 4 1590 4.337290741 3.36729583 3.113515309 5 1700 6.620073207 3.555348061 3.218875825 6 2595 3.63758616 2.116255515 2.841998174 7 2415.8 3.401197382 1.193922468 3.135494216 8 5000 8.565983356 2.862200881 3.295836866 9 2500 3.248434627 4.927253685 1.791759469 10 5500 5.220355825 7.60090246 2.602689685 11 2956.5 6.91770561 5.010635294 2.833213344 12 4113 6.891117638 5.135798437 2.862200881 13 5025 4.537961436 4.442651256 2.708050201 14 1540 1.203972804 2.140066163 2.351375257 15 3150 3.80666249 2.541601993 2.803360381 16 3203 8.565983356 3.597312261 2.602689685 17 3293 0.916290732 3.63758616 3.401197382 18 33.4 4.259152537 3.828641396 2.429217744 19 1850 1.609437912 3.583518938 2.602689685 20 1500 4.991112628 3.401197382 2.841998174 21 1540 4.875197323 3.583518938 2.793616089 22 6000 0.916290732 3.713572067 0.09531018 23 2032 3.988984047 5.811140993 3.141994781 24 632 4.094344562 3.258096538 2.094330154 25 5426 3.258096538 5.583496309 3.938470175 26 2467 3.891820298 5.857933154 0.879626748 27 (Missing data) 3.401197382 4.178992036 0.955511445 28 (Missing data) 3.258096538 5.996452089 0.378436436 29 1245 4.959342 4.532599493 3.80666249 31 5643 3.526360525 4.882801923 3.135494216 32 1346 4.17438727 4.718498871 1.196948189 33 2467 6.045005314 8.114324709 3.970291914 34 2456 3.951243719 4.812184355 1.974081026 35 3457 4.718498871 3.881563798 3.218875825 36 1436 6.137727054 6.333457232 3.549617387 37 232 1.238374231 5.780743516 0.631271777 38 587 2.442347035 2.541601993 1.043804052 39 3135 1.994700313 4.638604962 0.500775288 40 4326 4.025351691 4.294560609 3.135494216 41 2356 4.493120682 4.007333185 2.060513532 42 4536 3.157000421 5.834810737 1.667706821 43 2446 3.135494216 3.958906591 2.674148649 44 6423 4.828313737 3.17805383 3.772760938 45 234 1.280933845 4.143134726 0.405465108 46 341.4 3.931825633 3.148453361 1.547562509 47 4213 6.11368218 2.151762203 3.072693315 48 3538 5.926926026 3.19047635 3.433987204 49 2461 3.401197382 3.465735903 2.766319109 50 2314 3.63758616 2.791165108 2.272125886 51 1453 3.135494216 4.537961436 1.556037136 52 394 5.648974238 2.753660712 3.314186005 53 542 5.791488055 3.808882247 3.951243719 54 1346 1.029619417 2.890371758 0.262364264 55 3475 5.834810737 4.143134726 3.33220451 56 1246 3.214867803 3.737669618 1.750937475 57 2590 4.304065093 3.526360525 2.451005098 58 3212 6.11368218 3.555348061 3.433987204 59 379 5.837730447 2.850706502 1.829376333 60 2534 3.465735903 4.399375273 1.291983682 61 2548 6.584791392 2.251291799 3.970291914 62 3095 3.33220451 3.526360525 3.663561646 63 4532 3.044522438 2.48490665 0.470003629 64 2314 4.127134385 6.12905021 1.481604541 65 4525 5.659482216 5.749392986 2.517696473
[0060] For each area to be explored, the Euclidean distance between it and historically known oil and gas reservoir samples is calculated. This is obtained by calculating the difference between two samples across all selected features. Then, the K nearest samples are selected as its K nearest neighbors. Based on the labels of these K samples, a ranking evaluation is performed using the Gaussian kernel function. The Gaussian kernel function can effectively handle nonlinear problems, making previously indistinguishable patterns separable by mapping the data to a high-dimensional feature space. Here, the label is the target variable or category in the dataset.
[0061] The K value in the KNN model is determined to be the smaller of the training sample size and 20. This avoids overfitting on small datasets while maintaining computational efficiency on large datasets. A Gaussian kernel is introduced, which is a method to measure the similarity between two samples. It calculates a similarity score based on the Euclidean distance between the samples and a preset standard deviation (100 in this case).
[0062] The Gaussian kernel can take into account the non-linear relationship of distance in the feature space, making similarity calculation more flexible and adaptable to complex data distributions. Using the Gaussian kernel-based KNN algorithm in conjunction with a custom similarity calculation method (here, the Gaussian kernel function), all samples in the training set are sorted according to their similarity to the sample to be predicted. This way, the most similar sample will be placed at the top of the list.
[0063] For each area to be explored, the Euclidean distance between it and historically known oil and gas reservoir samples is calculated. The nearest neighbor model, after normalizing the parameters in a multi-dimensional space, projects them onto a coordinate system to find neighboring points, such as... Figure 3 , Figure 4 The data is categorized after normalization based on porosity and thickness, with each data point weighted and calculated once according to each of the four dimensions.
[0064] Figure 3 This is a schematic diagram of the standardized training samples in dimensional space provided in this embodiment of the invention. Figure 1 , Figure 3 This image shows a scatter plot with a 45-degree rotation after the data provided by this invention has undergone dimensional enhancement processing and is categorized. The rotation angle is 30 degrees, and the pitch angle is 30 degrees. Figure 4 This is a schematic diagram of the standardized training samples in dimensional space provided in this embodiment of the invention. Figure 2 , Figure 4 The diagram shows a scatter plot of the data provided by this invention after dimensional enhancement processing and classification at 30-degree rotation and 30-degree pitch angles.
[0065] Figure 5 This is a flowchart of a KNN algorithm calculation provided in an embodiment of the present invention, such as... Figure 5As shown, after starting, the data is preprocessed and transformed into a multidimensional dataset in memory. The distances between the test data and each data point in the model are calculated, and the data are sorted according to increasing distances. The error between the expected output and the actual output is calculated, and the recommended parameters are checked for suitability. If suitable, the process ends; otherwise, the K value is adjusted to improve the range, and the process returns to the step of calculating the distances between the test data and each data point in the model, until the recommended parameters are deemed suitable.
[0066] After preprocessing, the data has been transformed into a multidimensional dataset in memory. At this point, when evaluating a new project M, if there is a missing qualitative parameter X, the subsurface temperature can be calculated based on experience, such as formation thickness, surface air temperature, surface air temperature, and lithology, so that the temperature of the oil field can be known. Select several parameters related to the missing parameter from the acquired data and obtain recommended values by following the steps below.
[0067] In one embodiment, the preset Gaussian kernel distance correlation is:
[0068] ;
[0069] Among them, X t For the sample to be determined, X i For the first i The training samples, the samples to be determined, and the dimensions of the training samples are determined by the number of categories in the first oil and gas parameter data; d(X t X i X represents the distance between the sample to be determined and the training samples in the dimensional space; the dimension of the dimensional space is determined by the number of categories in the first oil and gas parameter data; tj X is the sample to be determined for the j-th label; ij For the j-th label i One training sample; σ represents the sum of distances in the dimensional space between the sample to be determined and all training samples on the j-th label; σ is the bandwidth of the Gaussian kernel function or the standard deviation of the Gaussian kernel; N is the number of training samples, and m is the number of labels.
[0070] In this example, the original data has a total of n samples, N training samples, and a d-dimensional space. The training samples can be represented as X = (X1, X2, ..., X...). n )∈R N×d ; n represents the total number of samples, N represents the number of training samples, and d represents the dimension of the dimensional space. Let C be the set of the labeled training set, and C has four labels, i.e., C = (C1, C2, ..., C4). Then each data point in the dataset can be represented as follows: X j =(X j1 X j2 ..., X jd ), j∈N; Xj The tag is C i , i=1,2,...,20. The algorithm proposed in this paper will be described in detail in the following sections.
[0071] In this method, Euclidean distance is used to calculate spatial distance. Assume... For four-dimensional test data, then Let i = 1, 2, ..., 20 be the four-dimensional training data. The distance calculation formula in the existing technology is as follows:
[0072] ;
[0073] Among them, X t For the sample to be determined, X i For the first i One training sample; X tk For the first k The first tag t One training sample; X ij For the j-th label i One training sample; j , k= 1, 2, ……,m .
[0074] As can be seen from the distance formulas introduced above, existing distance measurement methods calculate distances based on all features of the sample. Since it cannot be guaranteed that all features are consistent with the classification, if no weights are assigned to features when calculating distance, the distance between nearest neighbors becomes controlled by irrelevant features. The nearest neighbor method is highly sensitive to this situation. Therefore, the affinity distance function, a local distance function, has been proposed. It considers both distance and local learning of features, and its distance formula is:
[0075] ;
[0076] The formula for the weight function W in the formula is as follows:
[0077]
[0078] For the formula This represents the sum of distances between the test point and all training points on the j-th feature; This represents the training point X on the j-th feature. i The sum of distances to the remaining training points.
[0079] For weighting functions, some are... As a weighting function, the Gaussian kernel function is expressed by the following formula:
[0080]
[0081] Where Z is for making ;d(X t X i ) is X t The sample to be determined and X i No. i The distance between two points in a training sample, K(.), represents the kernel function.
[0082] This invention proposes a distance function based on a Gaussian kernel, which is represented in 3D space as follows:
[0083]
[0084] In the above equation, the first term on the right-hand side is the absolute distance between the test point and the training point, and the second term is relative to the weight function in the equation. X in the formula... t For the sample to be determined, X i For the first i The training samples, the samples to be determined, and the dimensions of the training samples are determined by the number of categories in the first oil and gas parameter data; d(X t X i X represents the distance between the sample to be determined and the training samples in the dimensional space; the dimension of the dimensional space is determined by the number of categories in the first oil and gas parameter data; tj For the first j The undetermined sample with 1 label; X ij For the first j The first tag i One training sample; Indicates the first j The sum of the distances in dimensional space between the sample to be determined on each label and all training samples; σ is the bandwidth of the Gaussian kernel function or the standard deviation of the Gaussian kernel.
[0085] The Gaussian kernel function in this paper uses the absolute distance function for calculation, and the distance in the numerator of the second term on the right-hand side of the equation is... Considered as the sample X to be determined t With training sample X i In the j The projection values of each feature onto the dimensional space. It is in the j The sample X to be determined has one feature. t The projection values of all training samples in the dimensional space, and the projection values of one training sample with the rest of the training samples in the dimensional space. j The projection values of the features (j = 1, 2, ..., d) onto the high-dimensional space are Therefore, the numerator in the weighting function can be seen as the influence of the test data on the entire training dataset in the dimensional space, while the denominator can be seen as the influence between the relevant test data and the relevant training data in the high-dimensional space.
[0086] Select the K projects closest to project M and obtain the parameter X values from these projects. Compare the parameter X values of the K projects, and assign the parameter X value of project M to the one with the highest percentage among the K projects, according to the principle of majority rule. This yields the recommended parameter.
[0087] Based on the labels of these k samples, the recommended parameter values are predicted using the K-Nearest Neighbors (KNN) artificial intelligence algorithm with Gaussian kernels.
[0088] After obtaining 20 sample data points, feature vectors, and the corresponding formulas for the Gaussian kernel KNN artificial intelligence algorithm through the three steps described above, the values of the parameters to be tested are predicted programmatically:
[0089] In step one, a training dataset was prepared, which contained multiple instances with features and labels. The test sample to be predicted was defined, and the K value and the standard deviation of the Gaussian kernel were set. The training set was sorted according to its similarity to the test sample using the Gaussian kernel function to obtain a list of stored data. The label values of the top K neighbors were selected from the sorted neighbors and stored in the list.
[0090] The KNN algorithm with a Gaussian kernel is called, taking the neighbor labels, a list of neighbor samples, the standard deviation of the Gaussian kernel, and the test sample itself as input, to calculate the weighted average predicted label. Finally, the predicted label value is output: the inferred porosity value i: 18.377862427129372.
[0091] This invention also proposes an oil and gas parameter determination device, the principle of which is similar to the oil and gas parameter determination method, and will not be described in detail here.
[0092] Figure 6 This is a schematic diagram of an oil and gas parameter determination device provided in an embodiment of the present invention, as shown below. Figure 6 As shown, the oil and gas parameter determination device may include:
[0093] The acquisition module 601 is used to acquire various historical oil and gas parameter data and samples to be determined; the samples to be determined are real-time oil and gas parameter data with at least one missing data value.
[0094] The training sample determination module 602 is used to select first oil and gas parameter data from a variety of historical oil and gas parameter data based on the sample to be determined using a correlation analysis algorithm, and determine the first oil and gas parameter data as the training sample; the first oil and gas parameter data is a variety of oil and gas parameter data whose correlation with the sample to be determined is higher than a preset correlation threshold.
[0095] The labeling module 603 is used to add labels to training samples based on the categories of data in the training samples;
[0096] The distance determination module 604 is used to calculate the distance between the sample to be determined and the training samples after adding labels by using a preset Gaussian kernel distance correlation relationship;
[0097] The neighbor sample determination module 605 is used to determine the first preset number of training samples sorted by distance from smallest to largest as neighbor samples.
[0098] The missing data determination module 606 is used to determine the missing data values in the sample to be determined by using the KNN algorithm based on the neighbor samples, the sample to be determined and the preset Gaussian kernel standard deviation.
[0099] In one embodiment, the training sample determination module 602 is specifically used for:
[0100] Based on the sample to be determined, correlation analysis algorithms are used to determine the correlation between various historical oil and gas parameter data and the sample to be determined.
[0101] Determine the data integrity of multiple oil and gas parameter data whose correlation exceeds a preset correlation threshold;
[0102] Oil and gas parameter data with a data integrity level higher than the preset data integrity level are identified as the first oil and gas parameter data;
[0103] The first oil and gas parameter data was selected as the training sample.
[0104] In one embodiment, the preset Gaussian kernel distance correlation is:
[0105] ;
[0106] Among them, X t For the sample to be determined, X i For the first i The training samples, the samples to be determined, and the dimensions of the training samples are determined by the number of categories in the first oil and gas parameter data; d(X t X i X represents the distance between the sample to be determined and the training samples in the dimensional space; the dimension of the dimensional space is determined by the number of categories in the first oil and gas parameter data; tj X is the sample to be determined for the j-th label;ij For the j-th label i One training sample; σ represents the sum of distances in the dimensional space between the sample to be determined and all training samples on the j-th label; σ is the bandwidth of the Gaussian kernel function or the standard deviation of the Gaussian kernel; N is the number of training samples, and m is the number of labels.
[0107] In one embodiment, the oil and gas parameter determination device further includes: a standardization processing module, used for:
[0108] Various historical oil and gas parameter data were standardized using normalization, unit conversion, and classification algorithms, respectively.
[0109] In one embodiment, the correlation analysis algorithm includes an analysis of variance algorithm and / or a chi-square test algorithm.
[0110] Compared with existing technologies where petroleum geology and reservoir engineering experts rely on experience-based judgment, this invention improves the accuracy and efficiency of determining missing oil and gas parameters by acquiring multiple historical oil and gas parameter data and samples to be determined. The samples to be determined are real-time oil and gas parameter data with at least one missing data value. Based on these samples, a correlation analysis algorithm is used to filter out first oil and gas parameter data from the various historical oil and gas parameter data, identifying this first oil and gas parameter data as training samples. The first oil and gas parameter data consists of multiple oil and gas parameter data whose correlation with the samples to be determined is higher than a preset correlation threshold. Labels are added to the training samples based on the data categories within them. The distance between the samples to be determined and the labeled training samples is calculated using a preset Gaussian kernel distance correlation relationship. The first preset number of training samples, sorted by distance from smallest to largest, are identified as neighbor samples. Based on the neighbor samples, the samples to be determined, and the preset Gaussian kernel standard deviation, the KNN algorithm is used to determine the missing data values in the samples to be determined.
[0111] This method employs a widely used supervised learning approach in modern artificial intelligence: the Gaussian kernel-based KNN algorithm. It proposes a parameter recommendation method based on the nearest neighbor algorithm to address the problem of missing parameter inference. Its key feature is the application of the Gaussian kernel-based KNN algorithm to data analysis in oil and gas exploration projects. This constructs a parameter recommendation system based on geological features and known oil and gas reservoir distribution, thereby effectively improving the effectiveness and accuracy of evaluating new oil and gas projects.
[0112] The results of this invention provide key parameters for oil and gas resource evaluation, oil and gas reserve evaluation, and oil and gas field development planning. Compared with current methods that rely on manual judgment based on experience, the prediction accuracy of these key parameters is significantly improved. Using this method to predict parameters of known oil and gas fields, the results show that parameters such as porosity, oil saturation, and oil and gas production have a consistency rate of over 90% with actual operating parameters of the oil and gas fields. This indicates that the method is reliable, easily applicable, and can serve as an important basis for oil and gas field investment decisions.
[0113] This invention also provides a computer device. Figure 7 This is a schematic diagram of a computer device in an embodiment of the present invention. The computer device 700 includes a memory 710, a processor 720, and a computer program 730 stored in the memory 710 and executable on the processor 720. When the processor 720 executes the computer program 730, it implements the above-mentioned method for determining oil and gas parameters.
[0114] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for determining oil and gas parameters.
[0115] This invention also provides a computer program product, which includes a computer program that, when executed by a processor, implements the above-described method for determining oil and gas parameters.
[0116] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0117] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0118] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0119] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0120] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method of determining a hydrocarbon parameter, characterized by, The method comprises the following steps: Obtain a plurality of historical oil and gas parameter data and a to-be-determined sample; The to-be-determined sample is real-time oil and gas parameter data with at least one missing data value; the historical oil and gas parameter data include multi-source geological and geophysical data of burial depth, thickness, seismic attribute, lithology, porosity and permeability; the to-be-determined sample includes burial depth, thickness, porosity and permeability; According to the geological theory, exclude characteristics irrelevant to oil and gas accumulation from the historical oil and gas parameter data; According to the to-be-determined sample, use a correlation analysis algorithm to screen first oil and gas parameter data from the plurality of historical oil and gas parameter data after excluding irrelevant characteristics, audit the first oil and gas parameter data, and determine the audited first oil and gas parameter data as training samples; The first oil and gas parameter data is a plurality of oil and gas parameter data with a correlation higher than a preset correlation threshold with the to-be-determined sample; Add labels to the training samples according to the categories of the data in the training samples; Calculate the distance between the to-be-determined sample and the training samples with added labels by using a preset Gaussian kernel distance correlation; Determine the neighbor samples as the first preset number of training samples sorted in ascending order of distance; Determine the missing data value in the to-be-determined sample by using a K-nearest neighbor (KNN) algorithm according to the neighbor samples, the to-be-determined sample and a preset Gaussian kernel standard deviation; In a multi-dimensional space, normalize each parameter, project it into a coordinate system, find the nearest point, and classify each data according to the normalized porosity and thickness. The correlation analysis algorithm includes a variance analysis algorithm and / or a chi-square test algorithm. The preset Gaussian kernel distance correlation is: ; Among them, X t For the sample to be determined, X i For the first i The training samples, the samples to be determined, and the dimensions of the training samples are determined by the number of categories in the first oil and gas parameter data; d(X t X i X represents the distance between the sample to be determined and the training samples in the dimensional space; the dimension of the dimensional space is determined by the number of categories in the first oil and gas parameter data; tj X is the sample to be determined for the j-th label; ij For the j-th label i One training sample; Let represent the sum of distances in the dimensional space between the sample to be determined and all training samples at the j-th label; σ is the bandwidth of the Gaussian kernel function or the standard deviation of the Gaussian kernel; N is the number of training samples, and m is the number of labels; the label is the target variable or category in the dataset; the first term on the right-hand side of the equation is the absolute distance between the test point and the training point, and the second term is the weight function; the numerator of the second term on the right-hand side... The sample X to be determined t With training sample X i The projection value in the dimensional space on the j-th feature. The sample X to be determined on the j-th feature t The projection values of a training sample onto the j-th feature in the high-dimensional space are the projection values of the training samples onto the j-th feature in the high-dimensional space. The numerator in the weighting function is the influence of the test data on the entire training dataset in the dimensional space, and the denominator is the influence between the relevant test data and the relevant training data in the high-dimensional space.
2. The method of claim 1, wherein, After obtaining the plurality of historical oil and gas parameter data, the method further comprises: Standardize the plurality of historical oil and gas parameter data by using a normalization algorithm, a unit conversion algorithm and a classification algorithm.
3. The method of claim 1, wherein, After the plurality of historical oil and gas parameter data and the to-be-determined sample, the method further comprises standardizing the plurality of historical oil and gas parameter data, including scaling to make the data on the same scale; converting type values, Boolean values, positions and geological layer information in the data set into numbers, converting addresses into codes, and converting whether into 1 or 0.
4. The method of claim 1, wherein, Determine the missing data value in the to-be-determined sample by using a K-nearest neighbor (KNN) algorithm according to the neighbor samples, the to-be-determined sample and a preset Gaussian kernel standard deviation, including: Determine the initial K value in the KNN model as the smaller one of the size of the to-be-determined sample and 20; sort the distances between the to-be-determined sample and each data in the KNN model in ascending order, calculate the error between the expected output and the actual output, determine whether the recommended parameter is appropriate, and if so, end the process, if not, adjust the K value, improve the range, and return to the step of calculating the distances between the to-be-determined sample and each data in the KNN model until the recommended parameter is appropriate.
5. The method of claim 1, wherein, Determine the missing data value in the to-be-determined sample by using a K-nearest neighbor (KNN) algorithm according to the neighbor samples, the to-be-determined sample and a preset Gaussian kernel standard deviation, including: Based on the sample to be determined, the correlation between various historical oil and gas parameter data and the sample to be determined is determined using a correlation analysis algorithm; the data completeness of various oil and gas parameter data with correlations higher than a preset correlation threshold is determined; the oil and gas parameter data with data completeness higher than a preset data completeness threshold is determined as the first oil and gas parameter data; the first oil and gas parameter data is determined as the training sample.
6. An apparatus for determining a hydrocarbon parameter, the apparatus comprising: include: The acquisition module is used to acquire various historical oil and gas parameter data and samples to be determined; The samples to be determined are real-time oil and gas parameter data with at least one missing data value; historical oil and gas parameter data include multi-source geological and geophysical data of burial depth, thickness, seismic attributes, lithology, porosity, and permeability; the samples to be determined include burial depth, thickness, porosity, and permeability. The training sample determination module is used to exclude features unrelated to hydrocarbon accumulation from historical hydrocarbon parameter data based on geological theory; based on the sample to be determined, a correlation analysis algorithm is used to screen the first hydrocarbon parameter data from various historical hydrocarbon parameter data after excluding irrelevant features, the first hydrocarbon parameter data is reviewed, and the reviewed first hydrocarbon parameter data is determined as the training sample. The first oil and gas parameter data consists of multiple oil and gas parameter data whose correlation with the sample to be determined is higher than a preset correlation threshold; The labeling module is used to add labels to training samples based on the categories of data in the training samples; The distance determination module is used to calculate the distance between the sample to be determined and the labeled training samples using a preset Gaussian kernel distance correlation relationship. The neighbor sample determination module is used to determine the first preset number of training samples sorted by distance from smallest to largest as neighbor samples. The missing data determination module is used to determine the missing data values in the sample to be determined by using the K-nearest neighbor (KNN) algorithm based on neighbor samples, the sample to be determined, and the preset Gaussian kernel standard deviation. In multidimensional space, after normalizing each parameter, it is projected onto the coordinate system to find neighboring points. After normalization based on porosity and thickness, it is classified. Each data point is weighted and calculated once according to the dimension of the sample to be determined. The correlation analysis algorithm includes an analysis of variance algorithm and / or a chi-square test algorithm; The preset Gaussian kernel distance correlation is as follows: ; Among them, X t For the sample to be determined, X i For the first i The training samples, the samples to be determined, and the dimensions of the training samples are determined by the number of categories in the first oil and gas parameter data; d(X t X i X represents the distance between the sample to be determined and the training samples in the dimensional space; the dimension of the dimensional space is determined by the number of categories in the first oil and gas parameter data; tj X is the sample to be determined for the j-th label; ij For the j-th label i One training sample; Let represent the sum of distances in the dimensional space between the sample to be determined and all training samples at the j-th label; σ is the bandwidth of the Gaussian kernel function or the standard deviation of the Gaussian kernel; N is the number of training samples, and m is the number of labels; the label is the target variable or category in the dataset; the first term on the right-hand side of the equation is the absolute distance between the test point and the training point, and the second term is the weight function; the numerator of the second term on the right-hand side... The sample X to be determined t With training sample X i The projection value in the dimensional space on the j-th feature. The sample X to be determined on the j-th feature t The projection values of a training sample onto the j-th feature in the high-dimensional space are the projection values of the training samples onto the j-th feature in the high-dimensional space. ; The numerator of the weighting function is the influence of the test data on the entire training dataset in the dimensional space, and the denominator is the influence between the relevant test data and the relevant training data in the high-dimensional space.
7. The apparatus of claim 6, wherein, Also includes: The standardized processing module is used for: Various historical oil and gas parameter data were standardized using normalization, unit conversion, and classification algorithms, respectively.
8. The apparatus of claim 6, wherein, Also includes: The standardized processing module is used for: Standardize various historical oil and gas parameter data, including scaling to make the data on the same scale; convert data type values, Boolean values, location, and geological layer information into numbers, addresses into codes, and whether they are converted into 1 or 0.
9. The apparatus of claim 6, wherein, The missing data determination module is specifically used for: The initial K value in the KNN model is determined to be the smaller of the sample size to be determined and 20. The distance between the sample to be determined and each data point in the KNN model is calculated and sorted according to the increasing distance relationship. The error between the expected output and the actual output is calculated. The recommended parameters are judged to be appropriate. If they are, the process ends. If not, the K value is adjusted to improve the range. The process returns to the step of calculating the distance between the sample to be determined and each data point in the KNN model until the recommended parameters are judged to be appropriate.
10. The apparatus of claim 6, wherein, The training sample determination module is specifically used for: Based on the sample to be determined, the correlation between various historical oil and gas parameter data and the sample to be determined is determined using a correlation analysis algorithm; the data completeness of various oil and gas parameter data with correlations higher than a preset correlation threshold is determined; the oil and gas parameter data with data completeness higher than a preset data completeness threshold is determined as the first oil and gas parameter data; the first oil and gas parameter data is determined as the training sample.
11. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 5.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1 to 5.
13. A computer program product, characterised in that, The computer program product includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Chip defect detection method based on Gaussian kernel mean value local linear embedding
CN112580568A
Shale oil reservoir rock mechanical parameter value prediction method
CN118052030A