A method for studying the automatic classification and prediction of natural gas exploration workload

By using multiple regression models to establish an automatic classification model in natural gas exploration, the problem of lack of systematicity and accuracy of existing prediction methods is solved, and more accurate and systematic workload prediction is achieved.

CN119226850BActive Publication Date: 2025-06-20SOUTHWEST PETROLEUM UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411283617.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-13
Publication Date
2025-06-20
Estimated Expiration
2044-09-13

AI Technical Summary

Technical Problem

The existing gas exploration workload prediction methods are based on experience or simple statistical analysis and lack systematicity and accuracy.

Method used

The multivariate regression model is used to establish an automatic classification model. Through automatic classification, automatic optimization and automatic prediction, the classification of sample data is identified and the fitted value of the exploration workload is calculated.

Benefits of technology

It improves the accuracy and systematicity of natural gas exploration workload prediction, can provide scientific basis for exploration decisions, and helps understand the distribution and changing trends of workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119226850B_ABST
    Figure CN119226850B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for automatically classifying and predicting natural gas exploration workload, which method comprises: establishing an automatic classification model based on multiple regression: assuming that all sample data are respectively x 1n , x 2n ... x pn , the number of influencing indicators is p, the number of all sample data is n, and the exploration workload is y. Classification is carried out according to different linear relationships satisfied between all exploration workloads and all sample data; constructing prediction sample data, and calculating the exploration workload of the prediction sample data by the automatic classification model. The present invention solves the problems that the existing prediction methods are based on experience or simple statistical analysis and lack systematicness and accuracy. The method principle of the present invention is simple, easy to implement, short in calculation time-consuming, and adopts automatic classification, automatic optimization and automatic prediction, and has a very wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for classifying natural gas exploration workloads, and more particularly to a method for automatically classifying and predicting natural gas exploration workloads. Background Art

[0002] Natural gas exploration is an important part of energy development, and the accuracy of workload prediction directly affects exploration decisions and investment benefits. Traditional prediction methods often rely on experience or simple statistical analysis, lacking systematicness and accuracy. Therefore, developing a method that can automatically classify and predict natural gas exploration workloads is of great significance for improving exploration efficiency and reducing costs. Summary of the Invention

[0003] The object of the present invention is to provide a method for automatically classifying and predicting natural gas exploration workloads, which solves the problems of existing prediction methods being based on experience or simple statistical analysis and lacking systematicness and accuracy. By adopting automatic classification, automatic optimization, and automatic prediction, it has a very broad application prospect.

[0004] To achieve the above object, the present invention provides a method for automatically classifying and predicting natural gas exploration workloads, which includes:

[0005] Step 1: Based on multiple regression, establish an automatic classification model

[0006] Assume that all sample data are respectively x 1n , x 2n ... x pn , the influencing indicators are p, the number of all sample data is n, and the exploration workload is y. If there is a linear relationship between all exploration workloads and all sample data, all sample data are classified into one category, as shown in the following formula:

[0007] y j = a0 + α1x 1j + a2x 2j + a3x 3j +... + a p x pj (1)

[0008] In formula (1): j represents the serial number of the jth sample data randomly selected from n samples, j = 0, 1... n; n represents the total number of all sample data; p represents the number of influencing indicators; a p represents the coefficient of the pth influencing indicator; x pj represents the jth sample data in the pth influencing indicator; y j represents the jth exploration workload.

[0009] If different linear relationships are satisfied between all exploration workloads and all sample data, the data points satisfying the same linear relationship are classified into one category. Suppose all sample data can be classified into k categories, and the number of samples in each category is n k , n k ≥ p + 1, then each category satisfies the following formula:

[0010] y j,k = a 0,k + a 1,k x 1j,k + a 2j,k x 2j , k +..+ a pj,k x pj,k (2)

[0011] In formula (2): k represents the total number of classifications, k ∈ [1, n); j represents the number of the jth sample data in the kth classification, j = 0, 1……nk; nk represents the number of sample data in the kth classification; p represents the number of influencing indicators, y j,k represents the fitted value of the jth exploration workload in the kth category; x pj,k represents the pth influencing indicator in the jth sample data in the kth category; a p,k represents the regression parameter of the pth influencing indicator in the kth category;

[0012] According to the least squares principle, the optimal solution of the regression parameter a in formula (2) is obtained as follows:

[0013]

[0014] In formula (3):

[0015]

[0016] The steps to identify which sample data in all sample data can be classified into one category are as follows:

[0017] (1) Set the number of sample data n in each category k ≥ p + 1;

[0018] (2) Through times, the more the number of samples n, the greater the number of loops. Therefore, the number of random selection sample loops is set to 1000 times. Randomly select non-repeating n k sample data from all sample data for multiple regression analysis to obtain the regression parameter to minimize the sum of squared fitting errors between the fitted value and the actual value of the exploration workload of the sample data Calculate the maximum value R of the absolute value of the relative error between the fitted value and the actual value of the exploration workload of the sample data in the k-th classification max = max{|R 1,k |, |R 2,k |,..., |R j,k |}; where the relative error of the exploration workload of the k-classification sample data is calculated as follows: where the y jk all represent the fitted values of the exploration workload calculated by the formula (2) for the k-th classification; the all represent the actual values of the j-th exploration workload in the k-th classification

[0019] (3) Use the regression parameters in step (2) to calculate the absolute value of the relative error between the fitted value and the actual value of the exploration workload of the remaining sample data |R o |, where o in |R o | is the number of remaining samples, o = 1, 2,..., nk.

[0020] If |R o | < R max , classify this sample data into this class, and at this time n k = n k +1; when all the remaining sample data are classified, the number of sample data that cannot be classified is n - n k , if n - n k ≥ p, repeat steps (1) to (3) until n - n k < p and the classification is completed, then each sample data that satisfies n - n k < p is a separate large class;

[0021] Step 2: Construct prediction sample data, and calculate the exploration workload of the prediction sample data by the automatic classification model

[0022] Construct prediction sample data, calculate the classification distance between the sample data and the prediction sample data, and obtain the minimum value of the classification distance; according to the minimum value of the classification distance, obtain the number of the sample data, query the regression parameters in all classifications in the automatic classification model by the number of the sample data, and calculate the predicted value of the prediction sample data and the sum of squared fitting errors of the prediction sample data.

[0023] Preferably, assume that the number of all sample data n = 20 and the influencing index p = 7. In step 1, assume that n k = p + 2 = 9, and the sample data is classified into four large classes, among which two large classes satisfy multiple regression, denoted as the first large class and the second large class.

[0024] More preferably, the regression parameters of the first large class are The sample data of the first major category obtained is n1 = 9, and the maximum value of the absolute value of the relative error is 0.017328571%.

[0025] More preferably, among the remaining 20 - 9 = 11 sample data, steps (1) to (3) are repeated to obtain the sample data of the second major category n2 = 9; the regression parameters of the second major category are The maximum value of the absolute value of the relative error is 0.107075%.

[0026] More preferably, the number of remaining samples is 20 - 18 = 2, satisfying n k <p + 1, so the remaining data are each divided into one major category. A total of four major categories are divided.

[0027] Preferably, assuming that the number of all sample data n = 20 and the influencing index p = 7, in step one, assume n k = p + 3 = 10, and the sample data is classified into 2 categories, denoted as the first major category and the second major category.

[0028] More preferably, the regression parameters of the first major category are The sample data of the first major category obtained is n1 = 10, and the maximum value of the absolute value of the relative error is 0.93307%.

[0029] More preferably, the regression parameters of the second major category are The sample data of the second major category obtained is n2 = 10, and the maximum value of the absolute value of the relative error is 32.67453333%.

[0030] More preferably, the number of remaining samples is 20 - 20 = 0, so the sample data is just divided into two major categories.

[0031] Preferably, in step two, the calculation formula of the classification distance is as follows:

[0032]

[0033] In formula (4):

[0034]

[0035] In formula (4), x i,n is the nth data of the ith influencing index in the predicted sample data, i = 1, 2,... p; x i,j,k is the jth data of the ith influencing index in the kth category of the sample data, j = 1, 2,..., n k ; D j,k represents the classification distance from the nth data in the predicted sample data to the jth data in the kth category of the sample data.

[0036] The minimum value of the classification distance is calculated as follows:

[0037] D jmin = min{D j,1 , D j,2 , …, D j,k} (5)

[0038] In formula (5), D j,1 represents the classification distance from the nth data in the predicted sample data to the jth data in the first category of the sample data, and D j,k represents the classification distance from the nth data in the predicted sample data to the jth data in the kth category of the sample data;

[0039] The predicted value of the predicted sample data is calculated as follows:

[0040] y j,m = a 0,m + a 1,m x 1j + a 2,m x 2j + a 3,m x 3j + … + a p,m x pj (6)

[0041] In formula (6), m represents the mth category in the sample data, m = 1, 2, ..., k; j represents the jth sample data in the predicted sample data, j = 1, 2 … n k ; n k represents the number of predicted sample data; p represents the number of influencing indicators; y j,m represents the predicted value of the predicted sample data in the mth category; q p,m represents the coefficient of the pth influencing indicator in the mth category; x pj represents the jth predicted sample data of the pth influencing indicator

[0042] Preferably, when the sum of squared fitting errors of the predicted sample data is the smallest, the predicted value of the predicted sample data is the exploration workload of the predicted sample data.

[0043] Preferably, when the number of classified samples n k in the automatic classification model < p + 1, the regression parameters cannot be obtained. Using the principle of proximity, the exploration workload of the predicted sample data is taken as the actual value of the exploration workload of the jth sample data.

[0044] A method for automatically classifying and predicting natural gas exploration workloads according to the present invention solves the problems that existing prediction methods are based on experience or simple statistical analysis and lack systematicness and accuracy, and has the following advantages:

[0045] 1. The method principle of the present invention is simple, easy to implement, and has a short calculation time. It adopts automatic classification, automatic optimization, and automatic prediction, and has a very broad application prospect.

[0046] 2. This method can be extended to non-linear classification and prediction. It can provide a scientific basis for exploration decision-making, help decision-makers better understand the distribution and change trend of exploration workload, so as to formulate a more reasonable exploration plan and investment strategy. At the same time, this method can also be applied to other similar fields, such as oil exploration and mineral resource development. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a key factor diagram affecting the workload in the field of marine conventional gas of the present invention.

[0048] Figure 2 It is a common multiple regression prediction curve diagram of the present invention.

[0049] Figure 3 It is the first major category fitting curve diagram of the present invention with n k = p + 1.

[0050] Figure 4 It is the second major category fitting curve diagram of the present invention with n k = p + 1.

[0051] Figure 5 It is the first major category fitting curve diagram of the present invention with n k = p + 2.

[0052] Figure 6 It is the second major category fitting curve diagram of the present invention with n k = p + 2.

[0053] Figure 7 It is the first major category fitting curve diagram of the present invention with n k = p + 3.

[0054] Figure 8 It is the second major category fitting curve diagram of the present invention with n k = p + 3.

[0055] Figure 9 It is the neural network learning fitting curve diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0056] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0057] Example 1

[0058] A method for automatically classifying and predicting the workload of natural gas exploration, the method comprising:

[0059] 1. Analyze marine conventional gas reservoirs to establish influencing factors of exploration workload

[0060] Through the analysis of multiple marine conventional gas reservoirs, the present invention believes that the deployment of exploration wells is inseparable from the six major conditions of "source, reservoir, cap, trap, migration, and preservation" required for hydrocarbon formation and the main controlling factors of hydrocarbon accumulation. Among them, the source condition, sedimentary microfacies, effective reservoir thickness, effective porosity, reservoir continuity, development of ancient and modern structures, development degree of source faults, gas reservoir trap type, reserve abundance, area of favorable zones, and success rate of exploration wells are key factors, jointly affecting the deployment of exploration wells. Therefore, the main influencing factors of the exploration workload of marine conventional gas fields are established.

[0061] As Figure 1 shown, the key factor diagram of the workload in the field of marine conventional gas affected by the present invention. From Figure 1 it can be seen that the factors affecting the deployment of exploration wells include geological factors and comprehensive factors. Among them, geological factors include reserve abundance, hydrocarbon generation intensity of the source rock, ancient and modern structural traps, development of source faults, sedimentary microfacies, reservoir thickness, reservoir porosity, reservoir distribution continuity, and area of favorable zones; comprehensive factors include the success rate of exploration wells in the gas reservoir (inside + outside) and the number of exploration wells in the gas reservoir (inside + outside).

[0062] 2. Select and establish characteristic index factors from the influencing factors of exploration workload

[0063] Taking structural traps as an example, data are screened out from the influencing factors of exploration workload, and characteristic indicators related to the workload of natural gas exploration are extracted from production. Therefore, the influencing factors of the workload in the field of marine conventional gas in structural traps are simplified to: effective thickness, porosity, success rate of development wells, reserve abundance, area of favorable zones, total number of exploration wells, gas-bearing area, and hydrocarbon generation intensity factors. The statistical results are shown in Table 1. The characteristics that have a significant impact on the prediction results are screened out. The main prediction of the exploration workload is how many wells need to be drilled to meet the corresponding workload. The commonly used prediction method is multiple linear regression prediction.

[0064] Table 1 Table of main factors affecting workload in the structural field

[0065]

[0066]

[0067] As Figure 2 shown, the commonly used multiple regression prediction curve diagram of the present invention. From Figure 2It can be seen that the error between the actual value and the fitted value is relatively large. If this method is used to predict future workload, a large error will be caused. Therefore, it is necessary to establish an automatic classification prediction model based on multiple regression for prediction.

[0068] 3. Construct an automatic classification model based on multiple regression

[0069] At present, many classification models have their own advantages, disadvantages and applicable fields, as shown in Table 2 for details.

[0070] Table 2 Statistical table of the advantages and disadvantages of different classification models

[0071]

[0072]

[0073] As shown in Table 2, these classification methods do not consider the regularity of the influence of each index factor on the workload, such as whether it is linear or non-linear, and these classifications only classify from the discreteness of the data.

[0074] However, for the prediction of exploration workload, both the regular classification and its predictability need to be considered. Therefore, the present invention establishes an automatic classification prediction method based on multiple regression, and the establishment process is as follows:

[0075] Assume that all sample data are respectively x 1n , x 2n ... x pn , the number of influencing indicators is p, the number of all sample data is n, and the exploration workload is y; if there is a linear relationship between all exploration workloads and all sample data, all sample data are divided into one category, as shown in the following formula:

[0076] y j = α0 + a1x 1j + a2x 2j + a3x 3j + … + a p x pj (1)

[0077] In formula (1): j represents the number of the jth sample data randomly selected from n samples, j = 0, 1... n; n represents the total number of all sample data; p represents the number of influencing indicators; a p represents the coefficient of the pth influencing indicator; x pj represents the jth sample data in the pth influencing indicator; y j represents the jth exploration workload.

[0078] If different linear relationships are satisfied between all exploration workloads and all sample data, the data points that satisfy the same linear relationship are grouped into one category. Assume that all sample data can be divided into k categories, and the number of samples in each category is n k , n k ≥ p + 1, then each category satisfies the following formula:

[0079] y j,k = a 0,k + a 1,k x 1j,k + a 2,k x 2j,k + a 3,k x 3j,k +…+ a p,k x pj,k (2)

[0080] In formula (2): k represents the total number of categories, k ∈ [1, n); j represents the number of the jth sample data in the kth category, j = 0, 1......n k ; n k represents the number of sample data in the kth category; p represents the number of influencing indicators, y j,k represents the fitted value of the jth exploration workload in the kth category (fitted value of the total number of exploration wells); x pj,k represents the pth influencing indicator in the jth sample data in the kth category; a p,k represents the regression parameter of the pth influencing indicator in the kth category.

[0081] The solution method of the coefficient matrix for each category is the same as the multiple regression method. According to the principle of the least squares method, the optimal solution of the regression parameter a described in formula (2) is obtained from the existing literature 1 (Liu Fan. Review of the Least Squares Method [J]. Journal of Central China Normal University (Natural Science Edition), 1992, (4): 46 - 57) as follows:

[0082]

[0083] In formula (3):

[0084]

[0085] The steps to identify which sample data can be grouped into one category are as follows:

[0086] (1) Set the number of sample data for classification n k ≥ p + 1, n k is at least p + 1;

[0087] (2) Through the loop times To reduce the number of operations, the number of times of randomly selecting the sample size is 1000 times, and n non-repeated samples are randomly selected from all n sample data k for multiple regression analysis of the sample data to obtain the regression parameters of the corresponding group to minimize the sum of squared fitting errors between the fitted value and the actual value of the exploration workload of the sample data (i.e., At this time, calculate the maximum value R of the absolute value of the relative error between the fitted value and the actual value of the exploration workload for each sample data max = max{|R 1,k |, |R 2,k |,..., |R j,k |}; According to the existing literature 2 (Zhu Jianxin, Li Youfa. Numerical Calculation Method, 3rd Edition [M]. Beijing: Higher Education Press, 2012: 196), calculate the relative error between the fitted value and the actual value of the exploration workload for the k-th classified sample data as where both represent the actual value of the j-th exploration workload of the k-th classification, that is, the total number of exploration wells, and yj,k both represent the fitted value of the exploration workload calculated by the formula (2) for the k-th classification, that is, the fitted value of the total number of exploration wells.

[0088] (3) Use the regression parameters in (2) to calculate the absolute value of the relative error |R o | (where o is the number of remaining samples, o = 1, 2,..., n k ), if |R o | < R max , classify this sample data into this category, and at this time n k = n k + 1; when all the remaining sample data are classified, the number of sample data that cannot be classified is n - n k , if n - n k ≥ p, repeat steps (1) to (3) until n - n k < p classification is completed, and each sample data that satisfies n - n k < p is a separate large category.

[0089] The above calculation process analysis is as follows:[[]]END]]

[0090] According to the previous sample data for classification calculation, it is known from Table 1 that the total number of samples is 20 (excluding the prediction samples), set the number of loop times to 1000 times, the influencing factor p = 7, and the minimum number of classification n k is at least 8.

[0091] Therefore, first assume n k= p + 1 = 8. First, randomly select 8 sample data from the total samples. According to the principle of permutations and combinations, there are a total of ways of selection. Since the calculated data is too large, assume the number of random selection sample cycles is 1000 times. Using the existing multiple regression theory (the above-mentioned existing literature 1), calculate the minimum value of the sum of squared fitting errors between the fitting value and the actual value of the exploration workload of the sample data where both represent the actual value of the jth exploration workload of the kth classification, that is, the total number of exploration wells, y j,k both represent the fitting value of the exploration workload calculated by the formula (2) of the kth classification, that is, the fitting value of the total number of exploration wells, and are specifically shown in Tables 3 - 10) regression parameters (calculated by formula (3)), and calculate the maximum value of the absolute value of the relative error between the fitting value and the actual value of the exploration workload of each sample data. Through the absolute value of the relative error R max = max{|R 1,k |, |R 2,k |,... |R j,k |} (the maximum value of the absolute value of the relative error in the randomly selected 8 samples is calculated to be 0.0000333%). Using this regression parameter (that is, ), substitute x 1j,k , x 2j,k ...... x pj,k which are the sample data corresponding to the effective thickness, porosity, development well success rate, reserve abundance, favorable area, gas-bearing area, and hydrocarbon generation intensity influence indicators in Table 1 in turn, into formula (2) to obtain the fitting value of the exploration workload (the fitting value of the total number of exploration wells), and calculate the absolute value of the relative error |R o | between the fitting value and the actual value of the exploration workload of the remaining sample data. If |R o | < R max , classify the sample data into this category to complete the classification of the sample data, as shown in Tables 3 - 4. As Figures 3 - 4 shown, it can be divided into six categories, among which two categories satisfy the multiple regression law, and the other four individual samples are each a category, specifically as follows:

[0092] The regression parameters of the first category

[0093] , n1 = 8, the maximum value of the absolute value of the relative error is 0.0000333%, and the sum of squared fitting errors where both represent the actual value of the jth exploration workload of the kth classification, that is, the total number of exploration wells, y j,kAll represent the fitted value of the exploration workload calculated by the formula (2) for the k-th classification, that is, the fitted value of the total number of exploration wells) is 2E-12, indicating a good fitting effect. According to the existing literature 3 (Han Ming. Probability Theory and Mathematical Statistics Course, 3rd Edition [M]. Shanghai: Tongji University Press, 2023), the specific results are shown in Table 3 and Figure 3 .

[0094] Table 3 n k = p + 1, the first major category table

[0095]

[0096]

[0097] As Figure 3 shown, for the n k = p + 1 first major category fitting curve graph of the present invention. It can be seen from Figure 3 that the total number of exploration wells value (actual value) and the fitted value of the total number of exploration wells (fitted value) are basically the same, indicating a very good fitting effect.

[0098] Second major category, regression parameter

[0099]

[0100] n2 = 7 + 1 = 8, the absolute value of the maximum relative error reaches 0.000033%, and the sum of the squared fitting errors where All represent the actual value of the j-th exploration workload of the k-th classification, that is, the total number of exploration wells value, y j,k All represent the fitted value of the exploration workload calculated by the formula (2) for the k-th classification, that is, the fitted value of the total number of exploration wells) is 5E-12, indicating a good fitting effect. The specific results are shown in Table 4 and Figure 4 .

[0101] Table 4 n k = p + 1, the second major category table

[0102]

[0103] As Figure 4 shown, for the n k = p + 1 second major category fitting curve graph of the present invention. It can be seen from Figure 4 that with the same calculation steps, the total number of exploration wells value (actual value) and the fitted value of the total number of exploration wells (fitted value) are basically the same, indicating a very good fitting effect.

[0104] The number of remaining samples is 20 - 16 = 4, which does not satisfy n k ≥ p + 1. Therefore, the remaining 4 sample data are separately divided into one major category, and a total of six major categories are divided, as shown in Table 5.

[0105] Table 5 n k = p + 1, Tables of the third, fourth, fifth, and sixth categories

[0106]

[0107] Classify and calculate according to the previous sample data. Assume n k = p + 2 = 9, complete the classification of the sample data, which is divided into four major categories in total. The regression parameter of the first major category n1 = 9, the maximum value of the absolute relative error is 0.017328571%, and the sum of the squared fitting errors where both represent the actual value of the jth exploration workload of the kth classification, that is, the total number of exploration wells, y j,k both represent the fitted value of the exploration workload calculated by the formula (2) for the kth classification, that is, the fitted value of the total number of exploration wells) is 0.000012579, indicating that the fitting effect is good. The specific results are shown in Table 6 and Figure 5 .

[0108] Table 6 n k = p + 2, Table of the first major category

[0109]

[0110]

[0111] As Figure 5 shown, the fitting curve graph of the first major category of the present invention nk = p + 2. From Figure 5 it can be seen that the total number of exploration wells (actual value) and the fitted value of the total number of exploration wells (fitted value) are basically the same, indicating that the fitting effect is very good.

[0112] For the second major category, the regression parameter n2 = 7 + 2 = 9, the maximum absolute value of the relative error reaches 0.107075%, and the sum of the squared fitting errors where both represent the actual value of the jth exploration workload of the kth classification, that is, the total number of exploration wells, y j,k both represent the fitted value of the exploration workload calculated by the formula (2) for the kth classification, that is, the fitted value of the total number of exploration wells) is 0.00005008, indicating that the fitting effect is good. The specific results are shown in Table 7 and Figure 6 .

[0113] Table 7 n k = p + 2, Table of the second major category

[0114]

[0115] AsFigure 6 As shown, in the present invention, n k = p + 2, the second major category of fitting curve graphs. From Figure 6 it can be seen that for the same calculation steps, the total value of the exploration wells (actual value) and the fitted value of the total number of exploration wells (fitted value) are basically the same, indicating that the fitting effect is very good.

[0116] The number of remaining samples is 20 - 18 = 2, which does not satisfy n k ≥ p + 1. Therefore, the remaining 2 sample data are each divided into a separate category, and a total of four major categories are divided, as shown in Table 8.

[0117] Table 8 n k = p + 2, the tables of the third and fourth major categories

[0118]

[0119] According to the classification calculation of the previous sample data, assuming n k = p + 3 = 10, the classification of the sample data is completed, and it is just divided into two major categories. The regression parameter of the first category is n1 = 10, the absolute value of the maximum relative error reaches 0.93307%, and the sum of the squared fitting errors where both represent the actual value of the jth exploration workload of the kth classification, that is, the total value of the exploration wells, and y j,k both represent the fitted value of the exploration workload calculated by the formula (2) of the kth classification, that is, the fitted value of the total number of exploration wells) is 0.0160267, indicating that the fitting effect is better. The classified data is shown in Table 9, and the fitting data is as Figure 7 shown.

[0120] Table 9 n k = p + 3, the table of the first major category

[0121]

[0122] As Figure 7 shown, in the present invention, n k = p + 3, the fitting curve graph of the first major category. From Figure 7 it can be seen that the total value of the exploration wells (actual value) and the fitted value of the total number of exploration wells (fitted value) are basically the same, indicating that the fitting effect is very good.

[0123] The regression parameter of the second major category is n2 = 10, the absolute value of the maximum relative error reaches 32.67453333%, and the sum of the squared fitting errors where both represent the actual value of the jth exploration workload of the kth classification, that is, the total value of the exploration wells, and y j,kAll represent the fitted value of the exploration workload calculated by the formula (2) for the k-th classification, that is, the fitted value of the total number of exploration wells) is 4.583465, indicating a poor fitting effect. The classification data is shown in Table 10, and the fitted data is as Figure 8 shown.

[0124] Table 10 n k = p + 3 The second major category table

[0125]

[0126] As Figure 8 shown, in the present invention, n k = p + 3 The fitting curve graph of the second major category. From Figure 8 it can be seen that the total number of exploration wells value (actual value) and the fitted value of the total number of exploration wells (fitted value) are not very consistent, indicating that the k = p + 3 fitting effect of the second major category is not as good as that of the k = p + 3 first major category.

[0127] The remaining number of samples is 20 - 20 = 0, so the sample data is just divided into two major categories.

[0128] Combined with the previous analysis, it can be obtained that when n k takes different values, the classification is different, and the fitting effect is also different.

[0129] 4. Construct prediction sample data, and calculate the exploration workload of the prediction sample data by the automatic classification model

[0130] Construct prediction sample data, calculate the classification distance between the sample data and the prediction sample data, and obtain the minimum value of the classification distance;

[0131] According to the minimum value of the classification distance, obtain the number of the sample data, query the regression parameters in all classifications in the automatic classification model by the number of the sample data, and calculate the predicted value of the prediction sample data and the sum of squared fitting errors of the prediction sample data.

[0132] The calculation formula of the classification distance is as follows:

[0133]

[0134] In formula (4), x i,n is the n-th data of the i-th influencing index in the prediction sample data, i = 1, 2,... p; x i,j,k is the j-th data of the i-th influencing index in the k-th category of the sample data, j = 1, 2,... n k ; D j,k represents the classification distance from the n-th data in the prediction sample data to the j-th data in the k-th category of the sample data;

[0135] The minimum value of the classification distance is calculated as follows:

[0136] D jmin = min{D j,1 , D j,2 , …, D j,k}} (5)

[0137] In formula (5), D j,1 represents the classification distance from the nth data in the predicted sample data to the jth data in the first category of the sample data, and D j,k represents the classification distance from the nth data in the predicted sample data to the jth data in the kth category of the sample data;

[0138] The predicted value of the predicted sample data is calculated as follows:

[0139] y j,m = a 0,m + a 1,m x 1j + a 2,m x 2j + a 3,m x 3j + … + a p,m x pj (6)

[0140] In formula (6), m represents the mth category in the sample data, m = 1, 2, …, k; j represents the jth sample data in the predicted sample data, j = 1, 2 …… n k ; n k represents the number of predicted sample data; p represents the number of influencing indicators; represents the predicted value of the predicted sample data in the mth category; a p,m represents the coefficient of the pth influencing indicator in the mth category; x pj represents the jth predicted sample data of the pth influencing indicator.

[0141] If the sum of squared fitting errors of the predicted sample data is the smallest, then the predicted value of the predicted sample data is the exploration workload of the predicted sample data. If the number of classified samples nk in the automatic classification model is less than p + 1, the regression parameters cannot be obtained. Using the principle of proximity, the exploration workload of the predicted sample data is taken as the actual value of the exploration workload of the jth sample data, specifically as follows:

[0142]

[0143] In formula (7): represents the exploration workload of the predicted sample data; y j,m represents the actual value of the exploration workload of the jth sample data, which is the total number of exploration wells.

[0144] Experimental Example 1 Verifies the Effectiveness of the Automatic Classification Prediction Method

[0145] To verify the effectiveness of the automatic classification prediction method in Example 1, the average value of all the previous samples (i.e., the prediction sample) was taken for prediction. The result showed that: the calculation result of the prediction sample (the average value shown in Table 1) was obtained through Formula (5), and the calculation result is shown in Table 11. Through the minimum distance in Table 11, it was found that the prediction sample was closest to Sample No. 16 (i.e., Sample No. 16 in Table 1). During the above classification process (i.e., as shown in Tables 3 to 10), it can be seen that Sample No. 16 belongs to the first sample in the fifth major category when n k = p + 1. However, the data of the fifth major category of samples is very few, and the regression parameters of the multiple regression equation cannot be obtained. Therefore, Formula (7) was applied to obtain the fitted value of the total number of exploration wells for the prediction sample which is 12, and its fitting squared error is 12.6025. Sample No. 16 belongs to the 8th sample in the first major category when n k = p + 2. Therefore, the regression parameters of the first major category (i.e., ) were used. Substitute x 1j,m , x 2j,m ......x pj,m which are the sample data corresponding to the effective thickness, porosity, development well success rate, reserve abundance, favorable area, gas-bearing area, and hydrocarbon generation intensity indicators of the prediction sample (the fitted values shown in Table 1) into Formula (6) for calculation, and the fitted value of the total number of exploration wells was obtained as 11.29, and its sum of squared fitting errors is 8.0656. When Sample No. 16 is at n k = p + 3, it belongs to the 10th sample in the first major category. Using the regression parameters of the first major category (i.e., ), substitute x 1j,m , x 2j,m ......x pj,m which are the sample data corresponding to the effective thickness, porosity, development well success rate, reserve abundance, favorable area, gas-bearing area, and hydrocarbon generation intensity indicators of the prediction sample (the prediction sample shown in Table 1) into Formula (6) for calculation, and the fitted value of the total number of exploration wells was obtained as 7.669, and its sum of squared fitting errors is 0.609961.

[0146] Table 11 Classification Distance Calculation Result Table

[0147]

[0148]

[0149] Using the existing Network neural network method, deep learning is performed on the same prediction samples according to the existing literature 4 (Jiao Licheng. Neural Network System Theory [M]. Xi'an: Xidian University Press, 1990: 284), and their learning effects are compared. The learning effects are shown in Figure 9 , using the values of the corresponding indicators of the prediction samples as the prediction input values, the predicted result is 11.32, and the sum of squared fitting errors is 8.2369.

[0150] Comprehensively analyzing the above experimental results, it is known that when n k = p + 3, the prediction accuracy is higher than that when n k = p + 1, n k = p + 2 and the network neural network method, which proves that the method of the present invention is reliable. With the change of randomly sampled data, different classification models are obtained, but generally increasing the sample quantity and sampling quantity will better reflect the overall distribution and improve the final classification accuracy of the model.

[0151] As Figure 9 shown, the neural network learning fitting curve graph of the present invention. It can be seen from Figure 9 that through the comparative analysis of the experimental results, the automatic classification prediction method of Embodiment 1 can achieve the same effect as the neural network prediction. However, the neural network depends on a sufficient amount of sample data. If the prediction samples are not within the range of the learning samples, the prediction results may deviate greatly, and the neural network cannot express the regularity between the indicators.

[0152] Although the content of the present invention has been described in detail through the above preferred embodiments, it should be recognized that the above description should not be considered as a limitation of the present invention. After those skilled in the art have read the above content, various modifications and substitutions to the present invention will be obvious. Therefore, the protection scope of the present invention should be defined by the appended claims.

Claims

1. A method for automatic classification and prediction of natural gas exploration workload, characterized in that: The method includes: Step 1: Establish an automatic classification model based on multivariate regression Assume that all sample data are x 1n , x 2n ……x pn There are p influencing indicators, n samples, and y exploration workload; If all exploration workloads and all sample data satisfy a linear relationship, all sample data are classified into one category, as follows: y j’ =a0+a1x 1j’ +a2x 2j’ +a3x 3j’ +…+a p x pj’ (1) In formula (1): j' represents the number of the j'th sample data randomly selected from n samples, j' = 0, 1...n; n represents the total number of all sample data; p represents the number of influencing indicators; a p represents the coefficient of the pth influencing indicator; x pj’ represents the j'th sample data in the pth influencing indicator; y j’ represents the j'th exploration workload; If all exploration workloads and all sample data satisfy different linear relationships, all sample data are divided into k categories, and the number of sample data in each category is n k , n k ≥p+1, then each class satisfies the following formula: y j,k =a 0,k +a 1,k x 1j,k +a 2j,k x 2j,k +..+a pj,k x pj,k (2) In formula (2): k represents the total number of categories, k∈[1,n); j represents the number of the jth sample data in the kth category, j=0,1……n k ;n k represents the number of sample data of the kth classification; p represents the number of influencing indicators, y j,k represents the fitted value of the j-th exploration workload of the k-th category; x pj,k represents the jth sample data of the pth influencing indicator in the kth category; a p,k represents the regression parameter of the p-th influencing indicator in the k-th category; According to the principle of least squares method, the regression parameter a in formula (2) is obtained as Regression parameters is the optimal solution for the regression parameter a, and the regression parameter for: In formula (3): The steps to identify which sample data can be classified into one category among all sample data are as follows: (1) Set the number of sample data n for each category k ≥p+1; (2) Randomly select non-repeating n from all sample data k Perform multiple regression analysis on sample data and obtain regression parameters Minimize the sum of squares of the fitting errors between the fitting value and the actual value of the exploration workload of the sample data Calculate the maximum value R of the absolute value of the relative error between the fitting value and the actual value of the exploration workload of the sample data in the kth category max =max{|R 1,k |,|R 2,k |…|R j,k |}; The relative error of the exploration workload for calculating k-classified sample data is: j=1,2,...,n k ; Wherein jk All represent the fitting values ​​of the exploration workload of the kth category calculated by the formula (2); All represent the actual value of the jth exploration workload of the kth category; (3) Using the regression parameters in step (2) Then calculate the absolute value of the relative error between the fitting value and the actual value of the exploration workload of the remaining sample data | R o |, wherein the |R o The o in | is the number of remaining samples, o = 1, 2, ..., n k ; If | R o |<R max , the sample data is classified into this category, then n k =n k +1; when all remaining sample data are classified, the number of sample data that cannot be classified is nn k , if nn k Repeat steps (1) to (3) until nn k <p classification is completed, then nn is satisfied k Each sample data of <p is a single large category; Step 2: Construct prediction sample data and calculate the exploration workload of the prediction sample data by the automatic classification model Construct prediction sample data, calculate the classification distance between sample data and prediction sample data, and obtain the minimum value of the classification distance; The number of the sample data is obtained according to the minimum value of the classification distance, and the regression parameters in all classifications in the automatic classification model are queried from the number of the sample data. The predicted value of the predicted sample data and the sum of squares of the fitting errors of the predicted sample data are calculated from the regression parameters.

2. The method according to claim 1, characterized in that Assume that the number of all sample data is n = 20, and the influence index is p = 7. In step 1, assume that n k When =p+2=9, the sample data is classified into four categories, of which there are two categories that meet the requirements of multiple regression, which are recorded as the first category and the second category.

3. The method according to claim 2, characterized in that The regression parameters of the first category are The sample data n1 of the first category is obtained as 9, and the maximum absolute value of the relative error is 0.017328571%.

4. The method according to claim 2, characterized in that: In the remaining 20-9=11 sample data, repeat steps (1) to (3) to obtain the sample data of the second largest category n2=9; the regression parameter of the second largest category is The maximum absolute value of the relative error is 0.107075%.

5. The method according to claim 1, characterized in that Assume that the number of all sample data is n = 20, and the influence index is p = 7. In step 1, assume that n k =p+3=10, the sample data is classified into 2 categories, recorded as the first category and the second category.

6. The method according to claim 5, characterized in that The regression parameters of the first category are The sample data of the first category n1=10 is obtained, and the maximum absolute value of the relative error is 0.93307%.

7. The method according to claim 5, characterized in that The regression parameters for the second category are The sample data of the second largest category n2=10 is obtained, and the maximum absolute value of the relative error is 32.67453333%.

8. The method according to claim 1, characterized in that In step 2, the calculation formula of the classification distance is as follows: In formula (4), x i,n is the nth item of the i-th influencing indicator in the predicted sample data, i = 1, 2, ... p; x i,j,k is the jth item of data of the i-th influencing indicator in the k-th category of the sample data, j = 1, 2, ..., n k ;D j,k Represents the classification distance from the nth item of data in the predicted sample data to the jth item of data in the kth category of the sample data; The minimum value of the classification distance is calculated as follows: D jmin =min{D j,1 ,D j,2 ,…,D j,k } (5) In formula (5), D j,1 represents the classification distance from the nth item in the predicted sample data to the jth item in the first category of the sample data, D j,k Represents the classification distance from the nth item of data in the predicted sample data to the jth item of data in the kth category of the sample data; The calculation of the predicted value of the predicted sample data is as follows: y j”,m =a 0,m +a 1,m x 1j” +a 2,m x 2j” +a 3,m x 3j” +…+a p,m x pj″ (6) In formula (6), m represents the mth class in the sample data, m = 1, 2, ..., k; j' represents the j'th sample data in the predicted sample data, j' = 1, 2 ... n k’ ;n k’ represents the number of predicted sample data; p represents the number of influencing indicators; y j”,m Represents the predicted value of the predicted sample data in the mth class; a p,m represents the coefficient of the pth influencing indicator in the mth category; x pj″ Represents the j'th predicted sample data of the p'th influencing indicator.

9. The method according to claim 1, characterized in that: If the sum of squares of the fitting errors of the predicted sample data is the smallest, the predicted value of the predicted sample data is the exploration workload of the predicted sample data.

10. The method according to claim 1, characterized in that If the number of samples classified in the automatic classification model is n k When <p+1, the regression parameter cannot be obtained. The nearest principle is adopted and the exploration workload of the predicted sample data is the actual value of the exploration workload of the jth sample data.

Citation Information

Patent Citations

  • Unconventional natural gas fracturing effect evaluation and productivity prediction method based on principal component analysis

    CN111046341A

  • Method and system for training yield prediction model and related equipment

    CN114169505A