A method, system, electronic device and medium for determining a formation classification characteristic factor
By filtering and processing drilling data using local outlier factors and Savitzky-Golay convolutional smoothing algorithms, combined with minimum description length and Pearson correlation analysis, the problem of low drilling data quality was solved, and the accuracy and computational efficiency of the formation classification model were improved.
Patent Information
- Application Number
- CN202310627102.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-30
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2043-05-30
AI Technical Summary
The presence of numerous outliers and missing values in drilling data makes it difficult to guarantee data quality, affecting the classification accuracy of formation classification models and the efficiency of data analysis.
The local outlier factor algorithm and Savitzky-Golay convolutional smoothing algorithm were used to preprocess the data. The stratigraphic classification characteristic factors were determined by combining minimum description length and Pearson correlation analysis, and a deep neural network model was used for training.
It significantly improves the interpretability and utilization of drilling data, reduces feature factor redundancy, and enhances the accuracy and computational efficiency of formation classification models.
Smart Images

Figure CN116644284B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of petroleum drilling engineering, in particular to a formation classification feature factor determination method and system, an electronic device and a medium. BACKGROUND
[0002] With the advent of microcomputers and the continuous improvement of computing performance, comprehensive mud logging technology is constantly improving. Comprehensive mud logging technology is a comprehensive mud logging operation in which circulating drilling fluid is used as a carrier for recording information, various detection instruments are used, and geological, oil and gas, pressure, rock physical properties and other information in the drilling fluid are recorded as a function of depth. In the comprehensive mud logging instrument, the sensor collects a group of data every five seconds, and each group of data has nearly a hundred feature factors. During drilling, a large amount of data is accumulated. However, due to the influence of environment, measurement method and system noise, there are a large number of abnormal values and missing values in the drilling data, the drilling data curve has obvious burrs, the data quality is difficult to guarantee, and the subsequent data analysis and mining are affected. Therefore, drilling data preprocessing is the basis for scientific research.
[0003] The comprehensive mud logging data of a single well has millions of rows, with various data segments under non-drilling conditions in between, which increases the data storage cost and reduces the data analysis and operation efficiency. Moreover, the drilling condition period division accuracy is low, and the data sets corresponding to various formations are not clearly divided, thereby affecting the training of the formation classification model and leading to low classification accuracy of the final formation classification model. SUMMARY
[0004] The purpose of the present application is to provide a formation classification feature factor determination method, system, electronic device and medium to improve the quality of formation classification data.
[0005] To achieve the above purpose, the present application provides the following solutions:
[0006] A formation classification feature factor determination method, comprising:
[0007] obtaining a historical drilling data time series matrix and a formation type corresponding to the historical drilling data; the historical drilling data time series matrix is a time series matrix of order a x b; a is the total number of types of historical drilling data; b is the total number of time points at which drilling data is obtained; the time points at which drilling data is obtained include drilling time points and non-drilling time points;
[0008] filtering the historical drilling data time series matrix to obtain drilling data sets of drilling time points of different formation types;
[0009] for the drilling data set of drilling time points of any formation type:
[0010] The drilling data set at the drilling time is preprocessed based on a local outlier factor algorithm and a Savitzky-Golay convolution smoothing algorithm to obtain a preprocessed drilling data set.
[0011] Based on a minimum description length principle and a Pearson correlation analysis, a plurality of formation classification characteristic factors in the preprocessed drilling data set are determined.
[0012] Based on a plurality of formation classification characteristic factors of different formation types, a deep neural network model is trained to obtain a formation classification model.
[0013] A plurality of formation classification characteristic factors at the current drilling time are obtained, and the formation classification model is used to determine the current formation type.
[0014] Optionally, the historical drilling data time series matrix is filtered to obtain a drilling data set at the drilling time of different formation types, specifically including:
[0015] According to the depth of the historical drilling data, the historical drilling data at the drilling time in the historical drilling data time series matrix is filtered out;
[0016] According to the formation type corresponding to the historical drilling data, the historical drilling data at the drilling time is classified to obtain a drilling data set at the drilling time of different formation types.
[0017] Optionally, the drilling data set at the drilling time is preprocessed based on a local outlier factor algorithm and a Savitzky-Golay convolution smoothing algorithm to obtain a preprocessed drilling data set, specifically including:
[0018] The drilling data set at the drilling time is processed by using a local outlier factor algorithm to obtain a corrected drilling data set;
[0019] The corrected drilling data set is filtered by using a Savitzky-Golay convolution smoothing algorithm to obtain a preprocessed drilling data set.
[0020] Optionally, the drilling data set at the drilling time is processed by using a local outlier factor algorithm to obtain a corrected drilling data set, specifically including:
[0021] The drilling data set at the drilling time is processed by using a local outlier factor algorithm to obtain an outlier point set;
[0022] The drilling data set at the drilling time is processed by using a local outlier factor algorithm to obtain an outlier point set;
[0023] Optionally, a plurality of formation classification characteristic factors in the preprocessed drilling data set are determined based on a minimum description length principle and a Pearson correlation analysis, and specifically include:
[0024] A plurality of formation classification factors in the preprocessed drilling data set are determined based on a minimum description length principle.
[0025] A plurality of formation classification characteristic factors are determined based on a Pearson correlation analysis of the plurality of formation classification factors.
[0026] A system for determining formation classification characteristic factors includes:
[0027] A historical data acquisition module is configured to acquire a historical drilling data time series matrix and a corresponding formation type of historical drilling data, wherein the historical drilling data time series matrix is an a×b order time series matrix, a is a total number of types of historical drilling data, and b is a total number of time points at which drilling data is acquired, and the time points include drilling time points and non-drilling time points.
[0028] A screening module is configured to screen the historical drilling data time series matrix to obtain drilling data sets at drilling time points for different formation types.
[0029] A characteristic factor determination module is configured to:
[0030] For a drilling data set at a drilling time point for any formation type:
[0031] The drilling data set at the drilling time point is preprocessed based on a local outlier factor algorithm and a Savitzky-Golay convolution smoothing algorithm to obtain a preprocessed drilling data set.
[0032] A plurality of formation classification characteristic factors in the preprocessed drilling data set are determined based on a minimum description length principle and a Pearson correlation analysis.
[0033] A model training module is configured to train a deep neural network model based on a plurality of formation classification characteristic factors for different formation types to obtain a formation classification model.
[0034] A classification module is configured to acquire a plurality of formation classification characteristic factors at a current drilling time point and determine a current formation type using the formation classification model.
[0035] An electronic device includes a memory configured to store a computer program and a processor configured to execute the computer program to cause the electronic device to perform the above-described method for determining formation classification characteristic factors.
[0036] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the formation classification characteristic factor determination method.
[0037] According to the specific embodiments of the present application, the following technical effects are disclosed.
[0038] The formation classification characteristic factor determination method, system, electronic device and medium provided by the present application can significantly improve the data interpretability and utilization rate by filtering and reorganizing the historical drilling data time sequence matrix. In addition, the local outlier factor algorithm used can effectively detect outliers in the data set; the Savitzky-Golay smoothing filter method selected can filter drilling data noise while ensuring the shape and width of the signal unchanged, improving the smoothness of the drilling data curve. Finally, combined with the minimum description length method and the Pearson correlation analysis method, the factors in the pretreated drilling data set are selected and the correlation information between the factors is analyzed, and finally the formation classification characteristic factors with reliability and strong independence are obtained, realizing the feature space dimension compression, which is helpful to the selection of subsequent formation classification model and the improvement of model operation efficiency and accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0040] Figure 1 The formation classification characteristic factor determination method flowchart provided by the present application is provided.
[0041] Figure 2 The formation classification characteristic factor determination method flowchart provided by the present application is provided.
[0042] Figure 3 The minimum description length model diagram is provided. DETAILED DESCRIPTION
[0043] The technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0044] The present application aims to provide a formation classification feature factor determination method, system, electronic device and medium to improve formation classification data quality.
[0045] The present application provides a data-driven drilling formation classification factor extraction method (i.e., a formation classification feature factor determination method) to solve the problems of low data quality, low data utilization, weak data interpretability, feature factor redundancy and unclear formation classification related factors in the above drilling engineering.
[0046] To make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below in combination with the drawings and specific embodiments.
[0047] Embodiment one
[0048] The present application relies on a data cutting method to divide the original drilling data (historical drilling data time series matrix) to obtain feature data sets under drilling conditions corresponding to various types of formations (drilling data sets at drilling time), thereby enhancing the interpretability of drilling data. At the same time, since there are nearly one hundred drilling data factors, the correlation between factors is complex, and there are some invalid factors and redundant factors among them. By using an effective formation classification factor extraction method, high-quality related factors of formation types are selected, which can reduce the dimensionality of drilling data, eliminate the redundancy of drilling data feature factors, greatly shorten the data mining time, reduce the storage cost, and improve the accuracy and training speed of the formation classification model.
[0049] As shown in Figure 1 and Figure 2 , the present application provides a formation classification feature factor determination method, which comprises:
[0050] Step 101: obtaining a historical drilling data time series matrix and a formation type corresponding to the historical drilling data; the historical drilling data time series matrix is a time series matrix of order a x b; a is the total number of types of historical drilling data; b is the total number of time points at which drilling data is obtained; the time points at which drilling data is obtained include drilling time points and non-drilling time points.
[0051] Step 102: screening the historical drilling data time series matrix to obtain drilling data sets at drilling time for different formation types.
[0052] As an optional embodiment, step 102 specifically comprises:
[0053] According to the well depth in the historical drilling data, the historical drilling data at drilling time in the historical drilling data time series matrix is screened out.
[0054] According to the stratum type corresponding to the historical drilling data, the historical drilling data at the drilling time is classified to obtain drilling data sets at the drilling time of different stratum types.
[0055] In practical application, the original drilling data is cut and divided by relying on data cutting means: by analyzing comprehensive logging data, combined with drilling daily reports and logging daily reports, drilling engineering conditions have many non-drilling conditions in addition to drilling, and the starting time of the working condition record is not clear. The data segments are interlaced under various stratum types and complex working conditions. Through data cutting and reorganization, the data set corresponding to various drilling conditions under each stratum type can be obtained, the interpretability of the original drilling data is enhanced, the data waste is reduced, the usability of the drilling data is improved, and a solid foundation is laid for data-driven research such as stratum classification and identification. Specifically as follows:
[0056] A comprehensive logging time sequence matrix (historical drilling data time sequence matrix) is established, which contains nearly one hundred characteristic factors such as well depth, drilling pressure, torque, drilling time, gas content and their parameter values.
[0057] Since the well depth is measured every five seconds, the well depth will change during drilling. The repeated section of the well depth data is deleted to obtain an ideal drilling data set (including historical drilling data at the drilling time). The comprehensive logging time sequence matrix is for all time periods, which contains drilling and non-drilling conditions. The well depth will not change in non-drilling conditions. The repeated section of the well depth data is deleted, i.e. the non-drilling section is deleted.
[0058] The data points with time jump in the ideal drilling data set are labeled to facilitate quick confirmation of the drilling and non-drilling time sequence interface.
[0059] According to the drilling daily report and the logging daily report, the drilling condition and the non-drilling condition are extracted. The drilling time is corrected, and the working condition information is recorded. The drilling daily report and the logging daily report generated by the drilling engineering site record feedback contain well site comprehensive information, drilling construction profile and logging construction profile, which contain well dynamic, well depth stratum type, construction profile and other information.
[0060] According to the data law of drilling operation characteristic factors, the working condition attribute of the drilling time sequence data is divided. According to the data law of whether the drilling pressure is zero and whether the well depth changes, the drilling and non-drilling conditions are judged.
[0061] For the same kind of strata, the data under different drilling periods are reorganized and labeled with strata type to obtain a feature factor data set (drilling time drilling data set) under each kind of strata. In practical application, the strata types in a well include Penglaizhen group, Suining group, upper Shaximiao group, lower Shaximiao group, Qianfoya group, Daanzhai segment of Ziliujing group, Ma'anshan segment of Ziliujing group, Dongyue temple segment of Ziliujing group, Zhenzhuchong segment of Ziliujing group, fifth segment of Xujiahe group, fourth segment of Xujiahe group, third segment of Xujiahe group, second segment of Xujiahe group, Xiaotangzi group, Ma'antang group, fourth segment of Leikoupo group, third segment of Leikoupo group, second segment of Leikoupo group, first segment of Leikoupo group, and Jialingjiang group.
[0062] For the drilling time drilling data set of any strata type:
[0063] Step 103: based on the local outlier factor algorithm and the Savitzky-Golay convolution smoothing algorithm, the drilling time drilling data set is preprocessed to obtain a preprocessed drilling data set.
[0064] As an optional implementation, step 103 specifically includes:
[0065] Step 1031: the drilling time drilling data set is processed by using the local outlier factor algorithm to obtain a corrected drilling data set. Step 1031 specifically includes:
[0066] S1: the drilling time drilling data set is processed by using the local outlier factor algorithm to obtain an outlier point set.
[0067] In practical application, the basic idea of the local outlier factor (LOF) algorithm is to first calculate a local reachable density of each data point according to the data density around the data point, and then further calculate an outlier factor of each data point through the local reachable density, which identifies the outlier degree of a data point. The larger the factor value is, the higher the outlier degree is, and the smaller the factor value is, the lower the outlier degree is. Finally, the top (n) points with the largest outlier degree are output. Specifically, it includes:
[0068] Input the drilling data feature factor data point set (a subset composed of each column data point in the feature factor data set).
[0069] Calculate the kth reachable distance of each data point in the kth distance neighborhood of each data point:
[0070] reach_dist k (o,p)=max{d k (o),d(o,p)}。
[0071] Wherein, d k(o) is the kth distance of the neighborhood point o, d(o, p) is the distance from the neighborhood point o to the data point p.
[0072] The kth local reachable density of each data point is calculated:
[0073]
[0074] Where, N k (p) is the kth distance neighborhood of p.
[0075] The kth local outlier factor of each point is calculated:
[0076]
[0077] The outlier point set is outputted for the data points with the largest n local outlier factors:
[0078] O = {o1, o2,..., o n}.
[0079] S2 performs outlier correction on the drilling data set at the drilling time according to the outlier point set, to obtain a corrected drilling data set.
[0080] For the characteristic factor outlier value detected by the above outlier detection method, the average value of the two observation values before and after the outlier value is used to correct the outlier value, or the outlier value is deleted according to the actual situation.
[0081] Step 1032: using the Savitzky-Golay convolution smoothing algorithm, filtering the corrected drilling data set to obtain a pretreated drilling data set.
[0082] The core idea of the Savitzky-Golay convolution smoothing algorithm is to perform p-order polynomial fitting on the data points in a certain length window, so as to obtain the fitting result. After discretization processing, the moving window least square polynomial smoothing filter is actually a moving window weighted average algorithm, but its weighted coefficients are not simple constant windows, but are obtained by least square fitting of a given high-order polynomial in the sliding window.
[0083] The filtering window with a width of q = 2m + 1 is set for the above preliminary data pretreated drilling data set, and each measurement point is x = (-m, -m + 1,..., 0,..., m - 1, m).
[0084] A p-1 order polynomial is used to fit the data points in the window:
[0085] y = a0 + a1x + a2x 2 +... + a p-1 x p-1 .
[0086] A system of q equations consists of p linear equations. To ensure the existence of solutions to the system, n > k.
[0087]
[0088] The drilling data prediction equations can be represented by a matrix as follows:
[0089] Y (2m+1)×1 =X (2m+1)×k ·A k×1 +E (2m+1)×1 .
[0090] The fitting parameters A are determined by fitting using the least squares method:
[0091] Least square solution of A for:
[0092] Therefore, the smoothed filter prediction value of the comprehensive logging feature factor is obtained. This reduces the impact of noise.
[0093] Where, B = X·(X T ·X) -1 ·X T .
[0094] Step 104: Based on the minimum description length principle and Pearson correlation analysis, determine multiple formation classification feature factors in the preprocessed drilling dataset.
[0095] As an optional implementation, step 104 specifically includes:
[0096] Step 1041: Using the minimum description length principle, determine multiple formation classification factors in the preprocessed drilling dataset.
[0097] In practical applications, the basic idea of the Minimum Description Length (MDL) principle is: for a given dataset D, in order to save the most storage space, we attempt to find a model M from a possible models (or programs / or algorithms). i (1≤i≤a), M i It can extract all the patterns in dataset D to the maximum extent, compress the data, and then use model M. i The data itself, including the compressed data C i When stored together, their total storage size is S. i (Size). Since different models have varying compression efficiencies for D, generally, the higher the compression ratio of D, the higher the model complexity. Therefore, from among many feasible compression schemes, we select the one with the smallest S...i Minimum Description Length. The principle of Minimum Description Length is to choose the model M that has the minimum total description length i . The Minimum Description Length model is shown in Figure 3
[0098] Applying MDL principle to the feature selection of stratigraphic classification dataset, MDL algorithm treats each feature in the stratigraphic classification dataset as a simple prediction model of the target attribute (stratigraphic class). These single prediction models are compared and scored using their corresponding MDL measures. Using MDL algorithm, the model selection problem becomes a data communication problem. The attribute score uses two parts of code to transmit data. The first part transmits the model, the model parameters are the target probabilities associated with each prediction value. The second part transmits the original data that is predicted wrong using the model. The formula is as follows:
[0099] S i (MODEL i ,D)=S(MODEL i )+S(C i )。
[0100] S i (MODEL i ,D) is the total size of applying the i-th well data to establish a simple prediction model for the stratigraphic class on the preprocessed well dataset, S(MODEL i ) is the size of applying the i-th well data attribute to establish a simple prediction model (MODEL i ) for the stratigraphic class target attribute, S(C i ) is the total size of all prediction errors after applying MODEL i to the i-th well data attribute.
[0101] Apply a well data in the preprocessed well dataset as a prediction attribute X1, X2, …, X a , respectively, and establish a prediction model with the stratigraphic class "label" column as the target attribute Y.
[0102] Where the prediction accuracy rate (compression rate, the ratio of accurate samples to the total number of samples is the accuracy rate) of X1 is c%, that is, X1 can correctly describe c% of Y data, and the remaining (1-c%) Y data (compressed data) cannot be correctly described by X1, so its total length L1 is:
[0103] Length(Model(X1,Y))+Length(Y)*(1-c%).
[0104] The remaining X2,...,X a the total length of the prediction model is L2,...,L a The minimum description length model is the one with the minimum L1,L2,...,L a The classification attribute (stratigraphic classification factor) corresponding to the minimum one is found.
[0105] For the preprocessed drilling data set D, the minimum S i (MODEL i , D) is taken, and the MDL algorithm is applied to obtain a relatively optimal characteristic of the target attribute, i.e., the stratigraphic classification factor attribute contains the most information related to the target attribute.
[0106] According to the MDL score ranking, the feature scores of different drilling data with respect to the stratigraphic category are obtained in turn, and the drilling data with a higher score is selected as the input of the Pearson correlation analysis.
[0107] Step 1042: According to the plurality of stratigraphic classification factors, the plurality of stratigraphic classification characteristic factors are determined by using the Pearson correlation analysis.
[0108] The set of stratigraphic classification factors obtained in the scoring stage in the above characteristic selection evaluates the prediction importance of each factor attribute in the preprocessed drilling data set with respect to the target attribute, but does not consider the relationship between these stratigraphic classification factor attributes, so it is necessary to explore the correlation between the stratigraphic classification factors based on the Pearson correlation analysis to investigate their independence.
[0109] According to the Pearson correlation analysis, the correlation coefficient between the characteristic factors in the preprocessed drilling data set is obtained, r represents the sample correlation coefficient, and p is the population correlation coefficient, which is unknown and is usually estimated by the sample correlation coefficient r:
[0110]
[0111] where X1 and X2 are two characteristic factors in the stratigraphic classification data set, is the cross product sum of deviations from the mean of X1 and X2, and are the sum of squares of deviations from the mean of X1 and X2, respectively.
[0112] According to the combination of classification factors selected by the MDL algorithm, the Pearson correlation analysis is performed, and finally the combination of stratigraphic classification characteristic factors with good independence is established as the input applied to the deep neural network model.
[0113] Step 105: Based on the plurality of stratigraphic classification characteristic factors of different stratigraphic categories, the deep neural network model is trained to obtain a stratigraphic classification model.
[0114] Step 106: Obtain a plurality of formation classification characteristic factors of a current drilling time, and determine a current formation type by using the formation classification model.
[0115] Compared with the prior art, the present application has the following advantages:
[0116] The present application adopts a data screening and reorganization method to screen, divide and label classify the original comprehensive logging instrument data, significantly improves the data interpretability and utilization, and consolidates the foundation of subsequent data mining.
[0117] The local outlier factor algorithm adopted by the present application considers the local and global properties of the drilling data set at the same time, and determines the outliers relative to the neighborhood point density, and when there are different clusters with different densities in the data set, the LOF can effectively detect outliers; The selected Savitzky-Golay smoothing filter method can filter out drilling data noise while ensuring the shape and width of the signal unchanged, and improve the smoothness of the drilling data curve.
[0118] The present application combines the minimum description length method and factor correlation analysis means to select features in the processed drilling formation classification data set, analyze the correlation information between the features, and finally obtain formation classification factors with reliability and strong independence, realize feature space dimension compression, and help to select the subsequent formation classification model and improve the model operation efficiency and accuracy.
[0119] Example two
[0120] In order to perform the method corresponding to the above-mentioned example one to realize the corresponding functions and technical effects, a formation classification characteristic factor determination system is provided below, comprising:
[0121] A historical data acquisition module is configured to acquire a historical drilling data time sequence matrix and a formation type corresponding to the historical drilling data; the historical drilling data time sequence matrix is a time sequence matrix of order a*b; a is the total number of types of historical drilling data; b is the total number of time points at which drilling data is acquired; the time points at which drilling data is acquired include drilling time points and non-drilling time points.
[0122] A screening module is configured to screen the historical drilling data time sequence matrix to obtain drilling data sets of drilling time points of different formation types.
[0123] A characteristic factor determination module is configured to:
[0124] For any drilling data set of drilling time points of a formation type:
[0125] Based on the local outlier factor algorithm and the Savitzky-Golay convolution smoothing algorithm, the drilling data set of the drilling time points is preprocessed to obtain a preprocessed drilling data set.
[0126] Determine a plurality of formation classification characteristic factors in the preprocessed drilling data set based on the minimum description length principle and Pearson correlation analysis.
[0127] A model training module is configured to train a deep neural network model based on the plurality of formation classification characteristic factors of different formation types, and obtain a formation classification model.
[0128] A classification module is configured to obtain a plurality of formation classification characteristic factors at a current drilling time, and determine a current formation type by using the formation classification model.
[0129] Embodiment three
[0130] An electronic device includes a memory and a processor, the memory is configured to store a computer program, and the processor is configured to run the computer program to make the electronic device execute the formation classification characteristic factor determination method in embodiment one.
[0131] Embodiment four
[0132] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the formation classification characteristic factor determination method in embodiment one.
[0133] In the specification, each embodiment is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts of each embodiment can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part.
[0134] The principles and implementation manners of the present application are described by using specific examples in the present application. The above description of the embodiments is only used to help understand the method of the present application and its core idea. For those skilled in the art, according to the idea of the present application, the specific implementation manner and application range can be changed. In summary, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A method for determining a formation classification characteristic factor, the method comprising: The method comprises the following steps: obtaining a historical drilling data time sequence matrix and stratum types corresponding to historical drilling data; the historical drilling data time sequence matrix is a time sequence matrix of a×b order; a is the total number of types of historical drilling data; b is the total number of time points at which drilling data is obtained; the time points at which drilling data is obtained include drilling time points and non-drilling time points; screening the historical drilling data time sequence matrix to obtain drilling data sets at drilling time points of different stratum types; for a drilling data set at a drilling time point of any stratum type: based on a local outlier factor algorithm and a Savitzky-Golay convolution smoothing algorithm, preprocessing the drilling data set at the drilling time point to obtain a preprocessed drilling data set; based on a minimum description length principle and Pearson correlation analysis, determining a plurality of stratum classification characteristic factors in the preprocessed drilling data set; based on a plurality of stratum classification characteristic factors of different stratum types, training a deep neural network model to obtain a stratum classification model; obtaining a plurality of stratum classification characteristic factors at a current drilling time point, and determining a current stratum type by using the stratum classification model.
2. The formation classification characteristic factor determination method of claim 1, wherein, The screening of the historical drilling data time sequence matrix to obtain the drilling data set at the drilling time point of the different stratum types specifically comprises: screening historical drilling data at drilling time points in the historical drilling data time sequence matrix according to well depth in the historical drilling data; classifying the historical drilling data at the drilling time points according to the stratum types corresponding to the historical drilling data to obtain drilling data sets at drilling time points of different stratum types.
3. The method according to claim 1, wherein, The preprocessing of the drilling data set at the drilling time point based on the local outlier factor algorithm and the Savitzky-Golay convolution smoothing algorithm specifically comprises: processing the drilling data set at the drilling time point by using the local outlier factor algorithm to obtain a corrected drilling data set; filtering the corrected drilling data set by using the Savitzky-Golay convolution smoothing algorithm to obtain a preprocessed drilling data set.
4. The formation classification characteristic factor determination method of claim 3, wherein, The processing of the drilling data set at the drilling time point by using the local outlier factor algorithm specifically comprises: processing the drilling data set at the drilling time point by using the local outlier factor algorithm to obtain an outlier point set; performing outlier correction on the drilling data set at the drilling time point according to the outlier point set to obtain a corrected drilling data set.
5. The method for determining formation classification characteristics factors according to claim 1, characterized in that, The determination of a plurality of stratum classification characteristic factors in the preprocessed drilling data set based on the minimum description length principle and the Pearson correlation analysis specifically comprises: determining a plurality of stratum classification factors in the preprocessed drilling data set by using the minimum description length principle; determining a plurality of stratum classification characteristic factors by using the Pearson correlation analysis according to the plurality of stratum classification factors.
6. A formation classification characteristic factor determination system characterized by, The method comprises the following steps: a historical data obtaining module is configured to obtain a historical drilling data time sequence matrix and stratum types corresponding to historical drilling data; The historical drilling data time sequence matrix is a time sequence matrix of a×b order; a is a total number of categories of the historical drilling data; b is a total number of time points at which the drilling data is acquired; the time points at which the drilling data is acquired include drilling time points and non-drilling time points; The screening module is configured to screen the historical drilling data time sequence matrix to obtain drilling data sets of drilling time points of different formation categories; The feature factor determination module is configured to: For the drilling data set of drilling time points of any formation category: Based on a local outlier factor algorithm and a Savitzky-Golay convolution smoothing algorithm, the drilling data set of drilling time points is preprocessed to obtain a preprocessed drilling data set; Based on a minimum description length principle and a Pearson correlation analysis, a plurality of formation classification feature factors in the preprocessed drilling data set are determined; The model training module is configured to train a deep neural network model based on a plurality of formation classification feature factors of different formation categories to obtain a formation classification model; The classification module is configured to acquire a plurality of formation classification feature factors of a current drilling time point, and determine a current formation category by using the formation classification model.
7. An electronic device, comprising: The method comprises: A memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to make the electronic device execute the formation classification feature factor determination method in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the formation classification feature factor determination method in any one of claims 1-5.
Citation Information
Patent Citations
PUE prediction method and device of data center and storage medium
CN115577307A
Power flow optimization method and system based on dynamic load prediction, and storage medium
CN116031888A