A Method for Predicting the Effect of Well Fracturing Based on a Kernel Regression Acceleration Algorithm

By using a method based on nuclear regression acceleration algorithm in the prediction of oil well fracturing effect, the problems of high storage pressure, long running time and low accuracy in the prior art are solved, and fast and accurate fracturing effect prediction is achieved.

CN117195140BActive Publication Date: 2025-06-13NORTHEAST GASOLINEEUM UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311206528.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-18
Publication Date
2025-06-13
Estimated Expiration
2043-09-18

AI Technical Summary

Technical Problem

The prior art has problems such as high algorithm storage pressure, long running time and low accuracy in predicting oil well fracturing effect.

Method used

Using a method based on the nuclear regression acceleration algorithm, a low-rank approximation kernel SVR model is constructed to achieve fast and accurate fracturing effect prediction by pre-processing, data division and matrix approximation of fracturing data.

Benefits of technology

It effectively reduces the storage pressure and running time complexity of the algorithm, and improves the prediction accuracy, so as to quickly and accurately predict fracturing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117195140B_ABST
    Figure CN117195140B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of oil well fracturing effect prediction, and particularly relates to a prediction method for oil well fracturing effect based on a kernel regression acceleration algorithm. The method includes obtaining fracturing data and preprocessing the data, and the preprocessing includes cleaning, filling, and conversion; performing logarithmic centering processing on the result after preprocessing, finding a projection vector to perform data partitioning to obtain m subsets; randomly extracting c columns from the kernel matrix to construct a column subset matrix C; constructing a cross matrix W according to the column subset matrix C, thereby obtaining a low-rank approximation matrix and obtaining a kernel SVR model; training the model obtained in step S3 on each region of the m subsets; finally, each kernel SVR model predicts the fracturing data to be identified falling within the same region. Through the present invention, the fracturing effect prediction can be carried out more accurately and quickly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of oil well fracturing effect prediction, and particularly relates to an oil well fracturing effect prediction method based on a kernel regression acceleration algorithm. Background Art

[0002] In oil and gas fields, fracturing refers to a method of forming fractures in oil and gas reservoirs by using hydraulic action during the production process of oil or gas. Hydraulic fracturing is an artificial formation fracture, which improves the underground oil flow environment, increases the oil well production, and plays an important role in improving the bottom flow conditions of the oil well, slowing down the interlayer velocity, and increasing the utilization rate of the oil reservoir. The fracturing process can greatly improve the production capacity of oil wells and the development effect of oil fields, and is a cost-effective measure to increase production.

[0003] In recent years, the research on fracturing measure effect prediction mainly uses statistical methods, generally using numerical simulation methods for prediction. The numerical simulation method requires accurate reservoir parameters and fracturing construction parameters, with complex calculations and large workloads, making it difficult to achieve ideal accuracy and difficult to meet the requirements of simple operation and fast calculation during on-site fracturing well selection. The grey theory and fuzzy neural network methods predict the fracturing effect, but the prediction accuracy is not high enough.

[0004] The current technical background is the SVR regression algorithm, data partitioning method, and kernel approximation algorithm.

[0005] In practical applications, it is found that the existing application technologies mainly use data mining algorithms to establish a model of fracturing effect and influencing factors for quantitative prediction of fracturing effect. However, they do not consider the large scale of fracturing data in actual development, resulting in problems such as large algorithm storage pressure, long running time, and the need to strengthen the accuracy rate. Summary of the Invention

[0006] (I) Technical Problems to be Solved

[0007] The present invention provides an oil well fracturing effect prediction method based on a kernel regression acceleration algorithm to overcome the defects of large algorithm storage pressure, long running time, and low accuracy rate existing in the prior art.

[0008] (II) Technical Solutions

[0009] To solve the above problems, the present invention provides an oil well fracturing effect prediction method based on a kernel regression acceleration algorithm, including:

[0010] Step S1: Obtain fracturing data and perform preprocessing on the logarithm, where the preprocessing includes cleaning, filling, and conversion;

[0011] Step S2: Use the result preprocessed in step S1 to perform logarithmic centering, find the projection vector to divide the data, and obtain m subsets;

[0012] Step S3: Randomly extract c columns from the kernel matrix to construct a column subset matrix C; construct a cross matrix W based on the column subset matrix C to obtain a low-rank approximation matrix , we get the kernel SVR model;

[0013] Step S4: train the model obtained in step S3 on each region of the m subsets;

[0014] Step S5: Finally, each core SVR model predicts the to-be-identified fracturing data falling within the same region.

[0015] Preferably, for data X=(x ij ) n×p Clean, fill and transform, where X is a dataset with n rows and p dimensions, x ij represents the jth feature of the i-th sample.

[0016] Preferably, step S2 specifically includes using a logarithmic centered PCA method to obtain a projection vector:

[0017]

[0018]

[0019] The sample data processed in step S1 is X=(x ij ) n×p , the sample data after logarithmic processing is

[0020]

[0021] w belongs to The feature vector of is also the projection vector; the instances are sorted in ascending order according to their corresponding projection values, and the instance set is decomposed into m non-intersecting instance subsets with approximately equal capacity by segment interception;

[0022] According to the above, we can get the vector w and get an ordered sequence w·z for the data set Z. 1′ ≤w·z 2′ ≤…≤w·z n′ , define b p Point energy segmentation interval [w·z 1′ ,w·z n′ ] to m subintervals,

[0023]

[0024] Among them, is the largest integer close to z;

[0025] Therefore, the data and the feature space where they are located can be divided into m sub-regions {D 1 , D 2 , …, D m}, where |D P | = n p ,

[0026]

[0027] Preferably, in step S3, constructing the column subset matrix C specifically includes: in the first round of sampling, part of the columns c1 are extracted according to the different sampling probabilities of each column of the kernel matrix K, and its pseudo-inverse matrix is calculated to further obtain the residual matrix after the first round of sampling; then, the sampling probability of the i-th column in the second round of sampling is calculated according to the residual matrix, and c2 columns are extracted from the kernel matrix at this probability to construct the column subset matrix C.

[0028] Preferably, in the step S3, the kernel SVR model is

[0029]

[0030] Among them, a i , a i * are Lagrange multipliers; b is the bias, which is a constant term; is the low-rank approximation matrix.

[0031] Preferably, the formula for the low-rank approximation matrix K~ is:

[0032]

[0033] Among them: C is the column subset matrix, W represents the m×n matrix formed by the intersection of m column samples in the matrix K and the corresponding n rows, and C T is the transpose of the column subset matrix.

[0034] Preferably, step S5 includes:

[0035] For the sample data to be recognized, first judge the subset it falls into, and then use the model corresponding to this subset for prediction.

[0036] (III) Beneficial effects

[0037] The oil well fracturing effect prediction method based on the kernel regression acceleration algorithm provided by the present invention fully considers the problem of large-scale fracturing data in actual development, divides the data, and then excellently completes the model training and prediction of large-scale data. At the same time, the matrix approximation algorithm is used to approximate the kernel matrix to the greatest extent under the constraint condition of low rank, reducing the storage pressure and running time complexity of the algorithm while maintaining the accuracy of the learning algorithm, and being able to quickly and accurately predict the fracturing data. Brief Description of the Drawings

[0038] Figure 1 It is a flowchart of the oil well fracturing effect prediction method based on the kernel regression acceleration algorithm in an embodiment of the present invention. Detailed Embodiment

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0040] As Figure 1 shown, the present invention provides an oil well fracturing effect prediction method based on the kernel regression acceleration algorithm, including the following steps:

[0041] Step S1: Obtain fracturing data and perform cleaning, filling, and conversion on the data.

[0042] Specifically, in the preprocessing stage of the fracturing historical data, it is necessary to clean, fill, and convert the data. For the problem that some data in the data set are missing, some missing data need to be removed, or the data need to be filled according to expert experience. Individual feature values of the data need to be converted through interval division. This step includes: performing cleaning, filling, and conversion on the data X=(x ij ) n×p

[0043] Among them, X is a data set with n rows and p dimensions, and x ij represents the j-th feature of the i-th sample

[0044] Step S2: Use the result after preprocessing in Step S1 to perform logarithmic centering processing, find the projection vector to divide the data, and obtain m subsets.

[0045] Specifically, in order to satisfy the maximum variance between the data in each sub-region after data division and ensure that the first principal component obtained contains more information of the data set, the logarithmic centering PCA method is used to obtain the projection vector:

[0046] ​

[0047]

[0048] Among them, the sample data processed in step S1 is X = (x ij ) n×p , and the sample data after logarithmic processing is

[0049]

[0050] w is the eigenvector belonging to , and it is also the projection vector; the instances are sorted in ascending order according to their corresponding projection values, and the instance set is decomposed into m non-overlapping real example subsets with approximately equal capacities by means of segment-by-segment interception;

[0051] According to the vector w obtained above, for the data set Z, an ordered sequence w·z 1′ ≤w·z 2′ ≤…≤w·z n′ can be obtained. Define b p points can divide the interval [w·z 1′ , w·z n′ into m sub-intervals,

[0052]

[0053] Among them, is the largest integer close to z;

[0054] Therefore, the feature space where the data is located can be divided into m sub-regions {D 1 , D 2 , …, D m}, where |D P | = n p ,

[0055]

[0056] Step S3: Randomly extract c columns from the kernel matrix to construct a column subset matrix C, and construct a cross matrix W according to the column subset matrix C, so as to obtain a low-rank approximation matrix , and obtain a kernel SVR model.

[0057] Among them, the construction of the column subset matrix C specifically includes: in the first round of sampling, part of the columns c1 are extracted according to the different sampling probabilities of each column of the kernel matrix K, its pseudo-inverse matrix is calculated, and the residual matrix after the first round of sampling is further obtained; then, according to the residual matrix, the sampling probability of the i-th column in the second round of sampling is calculated, and c2 columns are extracted from the kernel matrix with this probability to construct the column subset matrix C.

[0058] This step specifically includes: introducing the unequal probability adaptive sampling method Combining the idea of unequal probability sampling with adaptive sampling can fully retain the original data information while making the sample columns and sample rows more representative; reducing the sampling error while improving the sampling efficiency. Consider a symmetric positive semi-definite matrix K ∈ R n×m , and generating a low-rank approximation matrix of K based on m << n sample columns randomly selected from the matrix itself Assume that the sample columns have been selected. Let C′ denote the m×n matrix composed of these sample columns, and W denote the m×n matrix in matrix K formed by the intersection of m column sample columns and the corresponding n rows. The method is to approximate K using C′ and W, that is:

[0059]

[0060] where W -1 denotes the inverse matrix of matrix W. It can be proved that as the number of sampled column m increases, converges to K.

[0061] Kernel SVR decision function:

[0062]

[0063] where the kernel function k(X i , X j ) has various forms.

[0064] Algorithm research based on unequal probability adaptive sampling. For an arbitrary matrix, adaptive sampling is a relatively effective sampling method, and the error of only one round of sampling can be reduced by updating the sampling through multiple iterations. General adaptive sampling is two-round sampling: in the first round of sampling, some columns c1 are drawn according to different sampling probabilities of each column of the kernel matrix K, and its pseudo-inverse matrix is calculated to further obtain the residual matrix after the first round of sampling. Then, according to the residual matrix, the sampling probability of the i-th column in the second round of sampling is calculated, and c2 columns are drawn from the kernel matrix with this probability to construct the column subset matrix C. The cross matrix W is constructed according to the sample sub-matrix C, thereby obtaining the low-rank approximation matrix

[0065] First, assume that the number of drawn columns c is the same as the number of rows r 1 in the first round of sampling 2 and the number of rows in the second round of sampling is r 1 = 0.5r; Introduce a superimposed parameter a′. When the rank k is fixed, by changing the superimposed parameter, the number of sampled columns and rows can be changed, and the total number of sample columns and rows can be controlled within the overall range by adjusting the value range of a′, that is, c ≤ n, r = r 2 +r 1 ≤ m, thus ensuring the effectiveness and consistency of sampling.

[0066] The inclusion probability in the first round of sampling:

[0067] According to P i Select c1 columns with relatively large inclusion probabilities from K to construct matrix C1, and calculate the residual matrix:

[0068] A = K - KC 1 + C 1 (8)

[0069] Then calculate the inclusion probability of the i-th column in the second round of sampling based on the residual matrix:

[0070]

[0071] And select c2 columns from the kernel matrix with this probability to construct the column subset matrix C. Based on the sample columns of size c2 generated by sampling, the intersection points of the c2 rows of the corresponding matrix K form the cross matrix W. To sum up, the low-rank approximation matrix is:

[0072]

[0073] where: C is the column subset matrix, W represents the m×n matrix formed by the intersection of m column sample columns and the corresponding n rows in matrix K, C T is the transpose of the column subset matrix;

[0074] The final prediction model obtained in step S3 is:

[0075]

[0076] a i ,a i * are Lagrange multipliers; b is the bias, which is a constant term, is the low-rank approximation matrix.

[0077] Step S4: Train the model obtained in step S3 on each region of the m subsets.

[0078] Step S5: Finally, each kernel SVR model predicts the fracturing data to be recognized that falls into the same region.

[0079] Specifically, for the sample data to be recognized, first determine the subset it falls into, and then use the model corresponding to that subset for prediction. This method can maximize the preservation of local information of the data and improve the execution efficiency of learning algorithms under big data.

[0080] Use the hyperplane-based method to divide the data set into several subsets, and at the same time, based on the method to approximate the kernel matrix to obtain the kernel SVR model, and train the model on each subset. Finally, use each kernel SVR model to predict the instances to be recognized that fall into the same region, and predict the oil well fracturing effect for the oil well fracturing data.

[0081] It can be seen from the above technical solutions that the method of the present invention can perform fracturing effect prediction more accurately and quickly. First, the data is segmented, and then the model training and prediction of large-scale fracturing data are excellently completed. At the same time, the matrix approximation method is used to approximate the kernel matrix of the kernel regression algorithm to improve the efficiency of the optimization algorithm. This patent of the present invention has good adaptability and practicability for fracturing effect prediction, and provides a new method for oil well fracturing effect prediction. Aiming at the problem of large-scale fracturing data in actual development, based on the research of the kernel regression algorithm based on partitioning and matrix approximation, a method for predicting oil well fracturing effect that meets the needs of oilfield sites is explored.

[0082] Use the hyperplane-based method to divide the data set into several subsets, and at the same time, based on the method to approximate the kernel matrix to obtain the kernel SVR model, and train the model on each subset. Finally, use each kernel SVR model to predict the instances to be recognized that fall into the same region, and predict the oil well fracturing effect for the oil well fracturing data. The above embodiments are only used to illustrate the present invention, rather than to limit the present invention. Those of ordinary skill in the relevant technical fields can also make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the present invention. The patent protection scope of the present invention shall be defined by the claims.

Claims

1. A method for predicting the effect of oil well fracturing based on a kernel regression acceleration algorithm, characterized in that, it includes: Step S1: Obtain fracturing data and preprocess the data. The preprocessing includes cleaning, filling, and transformation, which specifically include: For data X = (x ij ) n×p , perform cleaning, filling, and transformation, where X is a dataset with n rows and p dimensions, and x ij represents the j-th feature of the i-th sample; Step S2: Perform logarithmic centering processing using the results after preprocessing in Step S1, find the projection vector for data partitioning, and obtain m subsets, which specifically includes obtaining the projection vector by using the logarithmic centering PCA method: Among them, the sample data processed in step S1 is X = (x ij ) n×p , and the logarithmically processed sample data is w belongs to is the eigenvector and also the projection vector; sort the instances in ascending order according to their corresponding projection values, and decompose the instance set into m non-overlapping real instance subsets with approximately equal capacities by means of segment-by-segment interception; According to the vector w obtained above, for the dataset Z, an ordered sequence w·z can be obtained. 1′ ≤w·z 2′ ≤…≤w·z n′ , define b p The point can divide the interval [w·z 1′ ,w·z n′ into m subintervals. wherein, is the largest integer close to z; Therefore, the data and the feature space where it is located are divided into m sub-regions {D 1 , D 2 , …, D m}, where |D P | = n p , Step S3: Randomly extract c columns from the kernel matrix to construct a column subset matrix C; construct a cross matrix W based on the column subset matrix C, thereby obtaining a low-rank approximation matrix Obtain a kernel SVR model; where the construction of the column subset matrix C specifically includes: in the first round of sampling, extract a part of the columns c1 according to the different sampling probabilities of each column of the kernel matrix K, calculate its pseudo-inverse matrix, and further obtain the residual matrix after the first round of sampling; then calculate the sampling probability of the i-th column in the second round of sampling according to the residual matrix, and extract c2 columns from the kernel matrix with this probability to construct the column subset matrix C; Step S4: Train the model obtained in Step S3 on each region of the m subsets; Step S5: Finally, each kernel SVR model predicts the fracturing data to be identified that falls into the same region.

2. The method for predicting the effect of oil well fracturing based on a kernel regression acceleration algorithm according to claim 1, characterized in that, in the said Step S3, the kernel SVR model is: where a i , a i * are Lagrange multipliers; b is a bias, which is a constant term, is a low-rank approximation matrix.

3. The method for predicting the effect of oil well fracturing based on a kernel regression acceleration algorithm according to claim 2, characterized in that, The low-rank approximation matrix has the following formula: Where: C is a column subset matrix, and W represents an m×n matrix C formed by the intersection of m column sample columns and the corresponding n rows in matrix K T which is the transpose of the column subset matrix.

4. The method for predicting the effect of oil well fracturing based on a kernel regression acceleration algorithm according to claim 1, characterized in that, Step S5 includes: For the sample data to be identified, first judge the subset it falls into, and then use the model corresponding to the subset for prediction.

Citation Information

Patent Citations

  • Low-order representing human body behavior identification method based on irrelevance constraint

    CN104298977A

  • Systems and methods for customizing kernel machines with deep neural networks

    US20190108444A1