Well leakage associated key parameter selection method and system based on data fusion processing
By integrating multiple screening methods and filtering algorithms in drilling engineering, and extracting key parameters using principal component analysis, the problems of low processing efficiency and data redundancy are solved, and data quality and parameter selection accuracy are improved.
Patent Information
- Application Number
- CN202311589244.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-27
AI Technical Summary
In drilling projects, the direct storage and processing efficiency of massive data is inefficient, data redundancy leads to difficulties in subsequent processing, and there are distortions and abnormalities in the data collected by the sensor in real time, affecting the accuracy of data analysis.
The data fusion processing method is used to filter outliers by fusing at least three screening methods, removing target outliers, and smoothing the data with filtering algorithms. Finally, the key parameters of well leakage association are extracted based on the principal component analysis method.
It effectively reduces the risk of data misjudgment, improves data quality, and significantly improves the accuracy of selecting key parameters for well leakage associations.
Smart Images

Figure CN120045930A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of oil drilling engineering, and specifically, to a method and system for selecting key parameters related to lost circulation based on data fusion processing. Background Art
[0002] In the current information era, in practical problems, the amount of information is increasing, and the dimension of sample parameters is very high, making the data dimensionality disaster inevitable. The number of characteristic parameters of the collected object is dozens or hundreds, making the direct storage and processing of data not very meaningful and inefficient. A large amount of data causes data redundancy, affecting subsequent processing. Therefore, in practical applications, necessary preprocessing is carried out on the data to improve the data quality. Through feature extraction, the data dimension is significantly reduced, key characteristic parameters are extracted, and the essential characteristics of the data are obtained.
[0003] Facing the massive data collected by sensors in real time, choosing the correct method to preprocess the data to obtain a clean and usable data set is a major challenge currently faced. The distorted and abnormal data in the drilling engineering acquisition system will greatly reduce the effectiveness and accuracy of information, easily cause misjudgment in data analysis, and also affect the accuracy of research.
[0004] Data preprocessing plays the role of a "security inspector" for data in the extraction of important characteristic parameters of faults. By means of preset certain condition thresholds and exploring the relationships between data, functions such as classification, screening, and filtering are realized for all sample data passing through this data preprocessing model, laying a solid foundation for the subsequent extraction of fault characteristic parameters.
[0005] The preprocessing of data can affect the selection of key parameters related to lost circulation. For this reason, the present invention provides a method and system for selecting key parameters related to lost circulation based on data fusion processing. Summary of the Invention
[0006] In order to overcome the defects of the prior art, the purpose of the present invention is to provide a method and system for selecting key parameters related to lost circulation based on data fusion processing. By fusing multiple screening methods to process the drilling data set, the risk of data misjudgment is effectively reduced. At the same time, a filtering algorithm is combined to smooth the drilling data set, improving the data quality of the drilling data set. Furthermore, based on the principal component analysis method, the accuracy of selecting key parameters related to lost circulation is greatly improved.
[0007] The present invention provides a method for selecting key parameters related to lost circulation based on data fusion processing, and the method includes:
[0008] S1. For the drilling time series data, use at least three screening methods to screen for outliers, and obtain at least three outlier groups;
[0009] S2. Based on the outlier group, remove the target outliers from the drilling time series data to obtain preprocessed drilling data;
[0010] S3. Smooth the preprocessed drilling data to obtain smoothed drilling data;
[0011] S4. Extract features from the smoothed drilling data to obtain key parameters related to lost circulation.
[0012] According to an embodiment of the present invention, in step S1, the drilling time series data is obtained, wherein the drilling time series data includes multiple sets of column data, and a single column data is a time series data of a drilling characteristic parameter.
[0013] According to an embodiment of the present invention, in step S1, outlier screening is performed by a first screening method to obtain a first outlier group, wherein the first screening method includes the following steps:
[0014] S111. Calculate the first average value of each column data in the drilling time series data;
[0015] S112. For each measurement value in each column data, calculate the residual error corresponding to each measurement value according to the corresponding first average value;
[0016] S113. Calculate the standard deviation of each column data according to Bessel's formula;
[0017] S114. For each measurement value, determine whether the measurement value is an outlier according to the corresponding residual error and the standard deviation;
[0018] S115. Form the first outlier group with the measurement values determined to be the outliers.
[0019] According to an embodiment of the present invention, in step S1, outlier screening is performed by a second screening method to obtain a second outlier group, wherein the second screening method includes the following steps:
[0020] S121. For each column data, regard the measurement values greater than the corresponding preset measurement value in the column data as suspicious values;
[0021] S122. Calculate the second average value and the average deviation of the remaining data, wherein the remaining data is the data in the column data except the suspicious values;
[0022] S123. Determine whether the suspicious value is the outlier according to the second average value and the average deviation;
[0023] S124. Compose the second outlier group with the suspicious values determined to be the outliers.
[0024] According to an embodiment of the present invention, in step S1, outlier screening is performed by a third screening method to obtain a third outlier group, wherein the third screening method includes the following steps:
[0025] S131. For each column of data, randomly select one of the measured values from the column of data, and denote it as the selected measured value;
[0026] S132. Sort the measured values in each column of data according to the magnitude relationship to obtain the sorted data;
[0027] S133. Determine the significance level value of the statistical test for detecting outliers;
[0028] S134. Determine the outlier detection critical value according to the significance level value and the sample size of the measured values in each column of data;
[0029] S135. For each selected measured value in each column of data, calculate the first difference and the third difference or the second difference and the third difference, wherein the first difference is the difference between the selected measured value and the previous adjacent measured value, the second difference is the difference between the second measured value and the first measured value in the sorted data, and the third difference is the difference between the selected measured value and the first measured value;
[0030] S136. Determine whether the current selected measured value is the outlier according to the first difference and the third difference or the second difference and the third difference, in combination with the outlier detection critical value, to obtain a judgment result;
[0031] S137. When the judgment result is yes, then remove the selected measured value from the corresponding column of data, and let the column of data after removal be the column of data, and return to step S131 until all the measured values in the column of data are traversed;
[0032] S138. When the judgment result is no, then return to step S131;
[0033] S139. Compose the third outlier group with all the selected measured values determined to be the outliers.
[0034] According to an embodiment of the present invention, in step S2, compare the outliers in different outlier groups to obtain the target outlier, wherein the target outlier is an outlier that exists in at least any two outlier groups.
[0035] According to an embodiment of the present invention, in step S3, the smoothed drilling data is obtained through the following steps:
[0036] S31. Calculate the predicted value X(k|k - 1) of the drilling system at time k using the optimal estimated value X(k - 1|k - 1) of the drilling system at time k - 1; the optimal estimated value of the drilling system corresponding to the initial time is the identity matrix; the drilling system is the system for collecting the drilling data;
[0037] S32. Calculate the covariance matrix P(k|k - 1) corresponding to the predicted value X(k|k - 1) at time k using the covariance matrix P(k - 1|k - 1) at time k - 1;
[0038] S33. Calculate the optimal estimated value X(k|k) at time k according to the predicted value X(k|k - 1) at time k, the measured value Z(k) at time k, and the Kalman gain at time k;
[0039] S34. Calculate the covariance matrix P(k|k) corresponding to the optimal estimated value X(k|k) at time k according to the Kalman gain at time k and the covariance matrix P(k|k - 1);
[0040] S35. Smooth the measured value at time k using the optimal estimated value X(k|k) at time k;
[0041] S36. Let k = k + 1, and return to step S31 until all the pre - processed drilling data is smoothed.
[0042] According to an embodiment of the present invention, in step S4, the key parameters related to lost circulation are obtained through the following steps:
[0043] S41. Construct a mathematical model for principal component analysis of the lost - circulation data set;
[0044] S42. Input the smoothed drilling data into a preset software;
[0045] S43. Select the column data in the smoothed drilling data where the measured values are not all zero to obtain the selected drilling characteristic parameters;
[0046] S44. Select the correlation test option to confirm whether each of the selected drilling characteristic parameters is suitable for principal component analysis;
[0047] S45. When passing the correlation test, perform principal component analysis on each of the selected drilling characteristic parameters to obtain the scree plot curve, the eigenvalues of each principal component, and the cumulative variance contribution rate;
[0048] S46. Select the target principal components from each of the principal components according to the gravel plot curve, the eigenvalue, and the cumulative variance contribution rate;
[0049] S47. After extracting each of the principal components and the common factors and performing factor rotation analysis, output the total variance explained data and the component score coefficient matrix;
[0050] S48. Substitute the value of the sum of squared rotated loadings in the total variance explained data and the values in the component score coefficient matrix into the principal component analysis mathematical model of the lost circulation dataset;
[0051] S49. Determine the key parameters related to lost circulation according to the principal component analysis mathematical model of the lost circulation dataset after substituting the values.
[0052] According to another aspect of the present invention, there is also provided a storage medium, which contains a series of instructions for executing the method steps described in any one of the above.
[0053] According to another aspect of the present invention, there is also provided a system for selecting key parameters related to lost circulation based on data fusion processing, which executes the method described in any one of the above. The system for selecting key parameters related to lost circulation includes:
[0054] An outlier determination module, which is used to screen outliers for the drilling time series data using at least three screening methods to obtain at least three outlier groups;
[0055] An outlier rejection module, which is used to reject the target outliers from the drilling time series data based on the outlier groups to obtain the preprocessed drilling data;
[0056] A smoothing processing module, which is used to perform smoothing processing on the preprocessed drilling data to obtain the smoothed drilling data;
[0057] A key parameter extraction module, which is used to perform feature extraction processing on the smoothed drilling data to obtain the key parameters related to lost circulation.
[0058] The present invention provides a method and a system for selecting key parameters related to lost circulation based on data fusion processing. Compared with the prior art, the following advantages are achieved:
[0059] By fusing at least three different screening methods to process the drilling dataset and combining a filtering algorithm to perform smoothing processing on the drilling dataset, the present invention improves the data quality of the drilling dataset, and then selects the key parameters related to lost circulation based on the principal component analysis method, greatly improving the accuracy of selecting the key parameters related to lost circulation.
[0060] Other features and advantages of the present invention will be set forth in the following description, and in part will be obvious from the description, or may be learned by practice of the present invention. The objectives and other advantages of the present invention may be realized and attained by the structure particularly pointed out in the specification, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The drawings are provided to further understand the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0062] Figure 1 The flowchart of the method for selecting key parameters related to lost circulation based on data fusion processing according to an embodiment of the present invention is shown;
[0063] Figure 2 The flowchart of the first screening method according to an embodiment of the present invention is shown;
[0064] Figure 3 The flowchart of the second screening method according to an embodiment of the present invention is shown;
[0065] Figure 4 The flowchart of the third screening method according to an embodiment of the present invention is shown;
[0066] Figure 5 The flowchart of the method for obtaining smoothed drilling data according to an embodiment of the present invention is shown;
[0067] Figure 6 The flowchart of the method for obtaining key parameters related to lost circulation according to an embodiment of the present invention is shown;
[0068] Figure 7 The schematic diagram of the method for selecting key parameters related to lost circulation based on data fusion processing according to an embodiment of the present invention is shown.
[0069] In the drawings, the same components are denoted by the same reference numerals. Additionally, the drawings are not drawn to actual scale. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] To make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the embodiments of the present invention in detail with reference to the drawings.
[0071] The object of the present invention is to provide a method and system for selecting key parameters related to lost circulation based on data fusion processing. By fusing at least three screening methods to process the drilling data set, the risk of data misjudgment is effectively reduced. At the same time, a filtering algorithm is combined to smooth the drilling data set, improving the data quality of the drilling data set. Furthermore, based on the principal component analysis method, the accuracy of selecting key parameters related to lost circulation is greatly improved.
[0072] Figure 1 Fig. 4 shows a flowchart of the steps of a method for selecting key parameters related to lost circulation based on data fusion processing according to an embodiment of the present invention.
[0073] As Figure 1 shown, in step S1, for the drilling time series data, at least three screening methods are used to screen out outliers, obtaining at least three outlier groups. Specifically, in step S1, the drilling time series data is acquired, where the drilling time series data includes multiple sets of column data, and a single column data is a time series data of a drilling characteristic parameter.
[0074] As Figure 1 shown, in step S2, based on the outlier groups, the target outliers are removed from the drilling time series data, obtaining the preprocessed drilling data. Specifically, after processing the drilling data with at least three screening methods, each screening method will screen out a certain number of outliers. By combining at least three groups of outliers, the outliers that finally need to be removed from the drilling data are selected. Further, in step S2, the outliers in different outlier groups are compared to obtain the target outliers, where the target outliers are the outliers that exist in at least any two outlier groups.
[0075] In one embodiment, when an outlier is screened as an outlier in two of at least three screening methods, it is determined that the outlier is the outlier that finally needs to be removed. The present invention determines the final outlier by using a weighted average method for each outlier. When performing data preprocessing on the drilling data, at least three screening methods are fused to avoid the problem of data misjudgment to a greater extent and greatly improve the quality of the drilling data.
[0076] As Figure 1 shown, in step S3, the preprocessed drilling data is smoothed to obtain the smoothed drilling data. Specifically, in step S3, a filtering algorithm is used to smooth the preprocessed drilling data to obtain the smoothed drilling data.
[0077] As Figure 1As shown, in step S4, feature extraction is performed on the smoothed drilling data to obtain key parameters related to well leakage. Specifically, the principal component analysis algorithm is used to perform feature extraction on the smoothed drilling data to obtain key parameters related to well leakage.
[0078] The present invention improves the data quality of the drilling data set by integrating at least three different screening methods to process the drilling data set and also combining a filtering algorithm to smooth the drilling data set. Furthermore, based on the principal component analysis method, key parameters related to well leakage are selected, greatly improving the accuracy of selecting key parameters related to well leakage.
[0079] Figure 2 Shows a flowchart of the steps of the first screening method according to an embodiment of the present invention.
[0080] In one embodiment, in step S1, outlier screening is performed through the first screening method to obtain a first outlier group. Among them, as Figure 2 shown, the first screening method includes steps S111 - S115.
[0081] As Figure 2 shown, in step S111, the first average value of each column of data in the drilling time series data is calculated
[0082] As Figure 2 shown, in step S112, for each measurement value in each column of data, according to the corresponding first average value calculate the residual error G corresponding to each measurement value d .
[0083] As Figure 2 shown, in step S113, the standard deviation S of each column of data is calculated according to Bessel's formula.
[0084] As Figure 2 shown, in step S114, for each measurement value, determine whether the S measurement value is an outlier according to the corresponding residual error and standard deviation.
[0085] As Figure 2 shown, in step S115, the measurement values determined to be outliers are grouped into a first outlier group.
[0086] The present invention uses the first screening method to judge and eliminate outliers containing gross errors. First, the average value of the equally precise independent measurement columns T i (i = 1, 2,..., n) should be calculated and the residual error and the standard deviation S of the measurement column in the drilling record data is calculated according to Bessel's formula. If a certain measurement value t in a certain column of data in the drilling time series datad Residual error (1 ≤ d ≤ g, where d represents the number of measured values of the characteristic parameters included in the column data) satisfies the following formula: |G d | > 3S, then t is considered d to be an outlier containing gross error and, for the first screening method, belongs to the outlier to be excluded.
[0087] Figure 3 FIG. shows a flowchart of steps of a second screening method according to an embodiment of the present invention.
[0088] In one embodiment, in step S1, outlier screening is performed by the second screening method to obtain a second outlier group, as Figure 3 shown, the second screening method includes steps S121 - S124.
[0089] As Figure 3 shown, in step S121, for each column of data, the measured values in the column data that are greater than the corresponding preset measured value are regarded as suspicious values.
[0090] As Figure 3 shown, in step S122, the second average value and the average deviation of the remaining data are calculated, where the remaining data is the data in the column data except the suspicious values.
[0091] As Figure 3 shown, in step S123, it is determined whether the suspicious value is an outlier according to the second average value and the average deviation.
[0092] As Figure 3 shown, in step S124, the suspicious values determined to be outliers are grouped into a second outlier group.
[0093] For some drilling time - series data, the second screening method can be used to judge the selection or rejection of suspicious values. First, find the average value and the average deviation S' of the remaining data except the suspicious values, and compare a certain measured suspicious value t d in a certain column of the drilling data with the average value . If it satisfies: then for the second screening method, discard the suspicious value, otherwise the measured value should be retained.
[0094] Figure 4 FIG. shows a flowchart of steps of a third screening method according to an embodiment of the present invention.
[0095] In one embodiment, in step S1, outlier screening is performed by the third screening method to obtain a third outlier group, as Figure 4 shown, the third screening method includes steps S131 - S139.
[0096] As Figure 4 shown, in step S131, for each column of data, a measurement value is randomly selected from the column data and denoted as the selected measurement value.
[0097] As Figure 4 shown, in step S132, the measurement values in each column of data are sorted according to the magnitude relationship to obtain the sorted data.
[0098] As Figure 4 shown, in step S133, the significance level value of the statistical test for detecting outliers is determined.
[0099] As Figure 4 shown, in step S134, the outlier detection critical value is determined according to the significance level value and the sample size of the measurement values in each column of data.
[0100] As Figure 4 shown, in step S135, for the selected measurement value in each column of data, the first difference and the third difference or the second difference and the third difference are calculated. Among them, the first difference is the difference between the selected measurement value and the previous adjacent measurement value, the second difference is the difference between the second measurement value and the first measurement value in the sorted data, and the third difference is the difference between the selected measurement value and the first measurement value.
[0101] As Figure 4 shown, in step S136, according to the first difference and the third difference or the second difference and the third difference, combined with the outlier detection critical value, it is determined whether the current selected measurement value is an outlier to obtain a judgment result.
[0102] As Figure 4 shown, in step S137, when the judgment result of step S136 is yes, the selected measurement value is removed from the corresponding column data, and the column data after removal is set as the column data, and step S131 is returned until all the measurement values in the column data are traversed.
[0103] As Figure 4 shown, in step S138, when the judgment result of step S136 is no, step S131 is returned.
[0104] As Figure 4 shown, in step S139, all the selected measurement values determined to be outliers are combined into a third outlier group.
[0105] The traditional data acquisition and processing method takes all the data collected by the instrument as valid data. However, in practice, due to environmental influence, some measurement values will be abnormal due to jumps. The present invention processes the measurement values of drilling data based on the third screening method as follows: Step a, for each column of data in the drilling time series data, a measurement value t is randomly selected. d, step b: Sort the measured values in each column from small to large as \(t_{(1)}\lt t_{(2)}\lt\cdots\lt t_{(n)}\), where \(n\) is the number of samples collected for a characteristic parameter; step c: Take the significance level value \(a = 0.05\) for the statistical test of detecting outliers, and determine the critical value \(D(a,n)\) from \((a,n)\); step d: Calculate the differences between the measured value \(t_{(i)}\) and its adjacent values \(t_{(i + 1)}-t_{(i)}\), \(t_{(i + 2)}-t_{(i + 1)}\), etc. Calculate \(D=(t_{(i + 1)}-t_{(i)}) / (t_{(i + 2)}-t_{(i + 1)})\) or \(D=(t_{(i + 2)}-t_{(i + 1)}) / (t_{(i + 3)}-t_{(i + 2)})\); step e: Determine the outlier. When \(D > D(a,n)\), for the third screening method, determine that the measured value \(t_{(i)}\) is the outlier to be removed; step f: After removing the measured value \(t_{(i)}\), select any one of the remaining sample measured values, re-number and sort the remaining sample measured values. The number of measured values in the remaining samples is \(n - 1\). Repeat the process of steps b - f until all outliers are detected. 1 \(<t_{(2)}\) 2 \(<\cdots\lt t_{(n)}\) d \(\cdots\lt t_{(n)}\) g , where \(n\) is the number of samples collected for a characteristic parameter; step c: Take the significance level value \(a = 0.05\) for the statistical test of detecting outliers, and determine the critical value \(D(a,n)\) from \((a,n)\); step d: Calculate the differences between the measured value \(t_{(i)}\) and its adjacent values \(t_{(i + 1)}-t_{(i)}\), \(t_{(i + 2)}-t_{(i + 1)}\), etc. Calculate \(D=(t_{(i + 1)}-t_{(i)}) / (t_{(i + 2)}-t_{(i + 1)})\) or \(D=(t_{(i + 2)}-t_{(i + 1)}) / (t_{(i + 3)}-t_{(i + 2)})\); step e: Determine the outlier. When \(D > D(a,n)\), for the third screening method, determine that the measured value \(t_{(i)}\) is the outlier to be removed; step f: After removing the measured value \(t_{(i)}\), select any one of the remaining sample measured values, re-number and sort the remaining sample measured values. The number of measured values in the remaining samples is \(n - 1\). Repeat the process of steps b - f until all outliers are detected. d and its adjacent value \(t_{(i + 1)}-t_{(i)}\) d \(-t_{(i)}\) d-1 , \(t_{(i + 2)}-t_{(i + 1)}\) 2 \(-t_{(i + 1)}\) 1 , calculate \(D=(t_{(i + 1)}-t_{(i)}) / (t_{(i + 2)}-t_{(i + 1)})\) or \(D=(t_{(i + 2)}-t_{(i + 1)}) / (t_{(i + 3)}-t_{(i + 2)})\); step e: Determine the outlier. When \(D > D(a,n)\), for the third screening method, determine that the measured value \(t_{(i)}\) is the outlier to be removed; step f: After removing the measured value \(t_{(i)}\), select any one of the remaining sample measured values, re-number and sort the remaining sample measured values. The number of measured values in the remaining samples is \(n - 1\). Repeat the process of steps b - f until all outliers are detected. d \(-t_{(i)}\) d-1 \() / (t_{(i + 2)}-t_{(i + 1)})\) d \(-t_{(i + 1)}\) 1 \()\) or \(D=(t_{(i + 2)}-t_{(i + 1)}) / (t_{(i + 3)}-t_{(i + 2)})\); step e: Determine the outlier. When \(D > D(a,n)\), for the third screening method, determine that the measured value \(t_{(i)}\) is the outlier to be removed; step f: After removing the measured value \(t_{(i)}\), select any one of the remaining sample measured values, re-number and sort the remaining sample measured values. The number of measured values in the remaining samples is \(n - 1\). Repeat the process of steps b - f until all outliers are detected. 2 \(-t_{(i + 1)}\) 1 \() / (t_{(i + 3)}-t_{(i + 2)})\) d \(-t_{(i + 2)}\) 1 \()); step e: Determine the outlier. When \(D > D(a,n)\), for the third screening method, determine that the measured value \(t_{(i)}\) is the outlier to be removed; step f: After removing the measured value \(t_{(i)}\), select any one of the remaining sample measured values, re-number and sort the remaining sample measured values. The number of measured values in the remaining samples is \(n - 1\). Repeat the process of steps b - f until all outliers are detected. d as the outlier to be removed; step f: After removing the measured value \(t_{(i)}\), select any one of the remaining sample measured values, re-number and sort the remaining sample measured values. The number of measured values in the remaining samples is \(n - 1\). Repeat the process of steps b - f until all outliers are detected. d ), select any one of the remaining sample measured values, re-number and sort the remaining sample measured values. The number of measured values in the remaining samples is \(n - 1\). Repeat the process of steps b - f until all outliers are detected.
[0106] Figure 5 Figure 42 shows a flowchart of the steps of the method for obtaining smoothed drilling data according to an embodiment of the present invention.
[0107] In one embodiment, in step S3, as Figure 5 shown, the smoothed drilling data is obtained through steps S31 - S36.
[0108] As Figure 5 shown, in step S31, the predicted value \(X(k|k - 1)\) of the drilling system at time \(k\) is calculated using the optimal estimate value \(X(k - 1|k - 1)\) of the drilling system at time \(k - 1\); the optimal estimate value of the drilling system corresponding to the initial time is the identity matrix; the drilling system is the system for collecting drilling data.
[0109] \(X(k|k - 1)=AX(k - 1|k - 1)+BU(k)\)
[0110] In the formula, \(U(k)\) is the control quantity of the current state. If there is none, it can be 0.
[0111] The drilling system is described by a linear differential equation. The drilling system can be understood as a measurement system for obtaining various characteristic parameters.
[0112] X(k) = AX(k - 1) + BU(k) + W(k)
[0113] Z(k) = HX(k) + V(k)
[0114] In the equations, X(k) is the state of the drilling system at time k, U(k) is the control quantity of the system at time k, and A and B are system parameter matrices. Z(k) is the measured value at time k, and H is the measurement parameter matrix. W(k) and V(k) are the noises in the system and the measurement process respectively. It is considered that the noises satisfy the Gaussian white noise model using Kalman filter estimation. Let the covariance matrices of W(k) and V(k) be Q and R respectively.
[0115] As Figure 5 shown, in step S32, the covariance matrix P(k|k - 1) corresponding to the predicted value X(k|k - 1) at time k is calculated using the covariance matrix P(k - 1|k - 1) at time k - 1.
[0116] P(k|k - 1) = APA T + Q
[0117] In the equation, A T is the transpose matrix of A, and Q is the noise of the system.
[0118] As Figure 5 shown, in step S33, the optimal estimated value X(k|k) at time k is calculated based on the predicted value X(k|k - 1) at time k, the measured value Z(k) at time k, and the Kalman gain at time k.
[0119] With the prediction of the system, the estimation is then made by referring to the measured values.
[0120] X(k|k) = X(k|k - 1) + k g (k)(Z(k) - HX(k|k - 1))
[0121] In the equation, k g (k) is the Kalman gain at time k, which can be calculated according to the formula described by the noise.
[0122] To achieve recursion, k g (k) is updated in real time every time.
[0123] k g (k) = P(k|k - 1)H T / (HP(k|k - 1)H T + R)
[0124] As Figure 5As shown, in step S34, the covariance matrix P(k|k) corresponding to the optimal estimated value X(k|k) at time k is calculated based on the Kalman gain and the covariance matrix P(k|k-1) at time k.
[0125] P(k|k) = (1 - k g (k)H)P(k|k-1)
[0126] As Figure 5 shown, in step S35, the measured value at time k is smoothed using the optimal estimated value X(k|k) at time k.
[0127] As Figure 5 shown, in step S36, let k = k + 1, and return to step S31 until all the preprocessed drilling data are smoothed (the measurement data converges, that is, tends to data within a certain range).
[0128] The covariance generated by each estimation is recursively combined with the current measured value to estimate the current best state of the system, realizing the smoothing of the comprehensive logging time series curve.
[0129] Data preprocessing is the most front-end link of the principal component analysis of fault impact parameters. A reasonable and effective data preprocessing model can effectively affect the input of the principal component analysis model by influencing the quality and controlling the quantity of data, making the output result of the extraction of fault important characteristic parameters more accurate.
[0130] The present invention first uses at least three different screening methods to remove outliers in the drilling time series data, and then performs Kalman filter smoothing on the processed data, realizing data fusion processing. The fused data can effectively reduce the risk of data misjudgment and effectively solve the problem of the measurement quality of drilling engineering sensor data.
[0131] Figure 6 Shows a flowchart of the steps of the method for obtaining key parameters related to lost circulation according to an embodiment of the present invention.
[0132] In one embodiment, in step S4, as Figure 6 shown, the key parameters related to lost circulation are obtained through steps S41 - S49.
[0133] As Figure 6 shown, in step S41, a mathematical model for principal component analysis of the lost circulation data set is constructed.
[0134] The expressions of the mathematical model for principal component analysis of the lost circulation data set (including the expression for principal component analysis of the lost circulation data set and the expression for lost circulation risk prediction score) are:
[0135]
[0136] F = b 1 Y 1 + b 2 Y 2 +... + b m Y m
[0137] In the formula, X n is the selected drilling characteristic parameter; Y m is the target principal component; a mn is the factor loading, the value in the component score coefficient matrix; b m is the contribution rate of the m-th target principal component, a value belonging to the sum of squared rotated loadings; F represents the predicted score of lost circulation risk.
[0138] As Figure 6 shown, in step S42, the smoothed drilling data is input into a preset software. Specifically, the smoothed drilling data is input into SPSS software (Statistical Package for the Social Sciences).
[0139] As Figure 6 shown, in step S43, the column data in the smoothed drilling data whose measured values are not all zero is selected to obtain the selected drilling characteristic parameters.
[0140] As Figure 6 shown, in step S44, the correlation test option is selected to confirm whether each selected drilling characteristic parameter is suitable for principal component analysis.
[0141] The existence of correlation among the original variables (among the selected drilling characteristic parameters) is the primary condition for principal component analysis. Otherwise, the original variables cannot be dimensionally reduced. The KMO and Bartlett tests can be used to test whether there is correlation among the variables.
[0142] Introduction to the KMO test: The KMO test is an index for comparing the simple correlation coefficient and partial correlation coefficient among variables. The KMO statistic ranges from 0 to 1. The closer the KMO value is to 1, the stronger the correlation among the variables, and the more suitable the original variables are for principal component analysis; the closer the KMO value is to 0, the weaker the correlation among the variables, and the less suitable the original variables are for principal component analysis.
[0143] Introduction to Bartlett spherical test: The assumptions of the Bartlett spherical test are as follows: H0: The correlation coefficient matrix is an identity matrix (variables are uncorrelated); H1: The correlation coefficient matrix is not an identity matrix (variables are correlated). The statistic of the Bartlett sphericity test is obtained from the determinant of the correlation coefficient matrix. If the significance level is less than the given α, then the null hypothesis should be rejected, that is, there is a correlation between the original variables, and it is suitable for principal component analysis; on the contrary, it is not suitable for principal component analysis.
[0144] When performing a correlation test, a correlation coefficient matrix will be obtained, and based on this matrix, it can be determined whether factor analysis is suitable for each selected drilling characteristic parameter.
[0145] As Figure 6 shown, in step S45, after passing the correlation test, principal component analysis is performed on each selected drilling characteristic parameter to obtain a scree plot curve, the eigenvalues of each principal component, and the cumulative variance contribution rate.
[0146] As Figure 6 shown, in step S46, the target principal components are selected from each principal component based on the scree plot curve, eigenvalues, and cumulative variance contribution rate.
[0147] Generally, the nodes before the node where the curve of the scree plot changes from steep to flat are selected as the principal components. If the eigenvalue is less than 1, it means that the explanatory power of this principal component is very low. Generally, an eigenvalue greater than 1 is used as the standard for screening principal components.
[0148] When the cumulative variance contribution rate of the first m principal components reaches a certain specific value (generally more than 80%), the first m principal components can be retained.
[0149] As Figure 6 shown, in step S47, after extracting each principal component and common factor and performing factor rotation analysis, the total variance explained data and the component score coefficient matrix are output.
[0150] The total variance explained includes information such as the initial eigenvalues of the principal components, the sum of squared loadings extracted, and the sum of squared loadings rotated.
[0151] As Figure 6 shown, in step S48, the values of the sum of squared loadings rotated in the total variance explained data and the values in the component score coefficient matrix are substituted into the principal component analysis mathematical model of the lost circulation dataset.
[0152] As Figure 6 shown, in step S49, the key parameters related to lost circulation are determined according to the principal component analysis mathematical model of the lost circulation dataset after substituting the values.
[0153] In the SPSS software, the initial variables (selecting drilling characteristic parameters) of the lost circulation dataset are standardized and made trend - consistent to eliminate the influence of dimensions. Standardization is carried out through the z - score standardization method in SPSS, and trend - consistency is achieved through index positive - orientation.
[0154] The smoothed drilling data is input into the SPSS software, and corresponding options are selected or corresponding icons are clicked in the software interface according to the operation processes of the principal component analysis method in the software. When using the SPSS software to process data, a rotated component matrix and a component score coefficient matrix will be obtained, and the importance of each characteristic parameter in the lost circulation fault can be observed.
[0155] Figure 7 Shows a schematic diagram of selecting key parameters related to lost circulation based on data fusion processing according to an embodiment of the present invention.
[0156] As Figure 7 shown, for a large amount of drilling data, when analyzing, some deviation points need to be removed. At least three screening methods can be used to perform fusion processing on the data. The principles are respectively making a mathematical comparison between the difference between the data outlier and the average value and the standard deviation to find the outlier; making a mathematical comparison between the difference between the outlier and the average value of the remaining data after removing the outlier and the average deviation to determine whether it is an outlier; judging whether there is an outlier through the statistic of the ratio of the difference between the outlier and the adjacent value to the range.
[0157] In the drilling time - series data, the existence of outliers is inevitable, which will have an adverse impact on the lost circulation analysis. Using at least three outlier detection methods can effectively solve the interference problem of outliers on the lost circulation fault analysis. For example, the 3 - sigma method has higher sensitivity compared to the 4d test method because, compared to the average deviation, the standard deviation can more sensitively reflect the existence of larger - deviation data. However, it may also cause the former to wrongly discard non - abnormal extreme values; the Dixon test method not only has different calculation formulas for the statistic according to the sample size, but also has different critical value tables for one - sided and two - sided tests; therefore, creatively, through the fusion of at least three outlier detection methods, the occurrence of data misjudgment problems is avoided to a greater extent, thereby improving the data quality.
[0158] As Figure 7 shown, the Kalman filter is an algorithm that uses a linear system state equation to perform optimal estimation of the system state through system input - output observation data. Since the observation data includes the influence of noise and interference in the system, the optimal estimation can also be regarded as a filtering process. Taking the weighted average according to the three methods of removing data outliers and combining the Kalman filter to achieve smoothing processing of the drilling time - series data.
[0159] For physical quantities in drilling engineering, various sensors are required for measurement. However, due to factors such as technology or other factors and noises that cannot be predicted or controlled by humans, the prediction of physical quantities by sensors cannot be completely accurate. The present invention obtains the best estimate through Kalman filtering combined with observation and model prediction, realizes iterative update, and effectively solves the problem of the measurement quality of drilling engineering sensor data.
[0160] The core idea of Kalman filtering is: based on the optimal estimate value x at time k - 1 k-1 to predict the predicted value x at time k k|k-1 ; then according to the measurement value Z at time k k and the predicted value x at time k k|k-1 , obtain the best estimate value x at this moment k ; perform iterative loop.
[0161] In the comprehensive mud logging time series analysis task, the greatest advantage of using Kalman filtering in the present invention is that it can use the state space form to represent the unobserved component model. Therefore, according to the Kalman idea, the comprehensive mud logging time series is smoothed, so that the measurement data continuously converges towards the true value.
[0162] As Figure 7 shown, for the selection of key parameter characteristics related to lost circulation, the present invention adopts the principal component analysis method to achieve dimensionality reduction. The technical source of the principal component analysis method (PCA) is the technology of matrix operation, matrix diagonalization and matrix spectral decomposition technology. Its core idea is to explain most of the information of the original variables through the linear combination of the original variables, achieve the purpose of dimensionality reduction, and thus simplify the complexity of the problem. For the multi-dimensional data set with the working condition of lost circulation, the PCA algorithm can be used to extract the key characteristic parameters related to lost circulation.
[0163] The central idea of feature extraction is to reduce the data dimension and simplify the data. The basis of data feature extraction is to map the data according to a certain rule, and map the data from the input space to the feature space through the mapping rule. Generally, this kind of mapping should comply with two criteria: one is that the feature space should be able to retain the main classification information in the original input space to the greatest extent, and the other is that under the rule of data dimensionality reduction, the dimension of the data input space should be greater than the dimension of the feature space. Principal component analysis is the most commonly used feature extraction algorithm in pattern classification. It is a data compression method that meets the above mapping criteria. As one of the classical feature extraction algorithms, principal component analysis can comprehensively describe the internal information contained in the original data, and convert the original data into fewer comprehensive features for representation. Its usage criterion is to make its variance optimal in the sense of statistical mean square.
[0164] The present invention processes the drilling data set by integrating at least three screening methods, weakens the adverse effects of outliers on data analysis, and effectively reduces the risk of data misjudgment.
[0165] The present invention obtains the best estimate of the measurement data of the drilling engineering sensor through Kalman filtering combined with observation and model prediction, realizes iterative update, and effectively solves the problem of the measurement quality of the drilling engineering sensor data.
[0166] The present invention uses the PCA algorithm to extract strongly correlated parameters of well leakage faults, and solves the problems of data redundancy and dimensionality disaster; innovatively uses the SPSS data analysis tool to perform correlation analysis on the well leakage data set after data preprocessing, observes the principal components through the scree plot, analyzes the principal components by eigenvalue and sum of squared loadings, and better explains the relationship between the principal components and each characteristic parameter through factor rotation, obtains the mathematical expression of the well leakage principal component mathematical model, and extracts the key correlated parameters of well leakage.
[0167] The present invention integrates at least three screening methods, uses Kalman filtering to smooth the data sequence, and uses the PCA method to select the key parameter features of well leakage correlation, so as to solve the problems of poor data quality of the well leakage data set in drilling engineering, difficult analysis and processing of massive data, and difficult extraction of key characteristic parameters related to faults.
[0168] A method and system for selecting key parameters related to well leakage based on data fusion processing provided by the present invention can also cooperate with a computer-readable storage medium, on which a computer program is stored, and the computer program is executed to run a method for selecting key parameters related to well leakage based on data fusion processing.
[0169] The computer program can run computer instructions, and the computer instructions include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc.
[0170] The computer-readable storage medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0171] It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.
[0172] According to another aspect of the present invention, there is also provided a system for selecting key parameters related to lost circulation based on data fusion processing, which executes a method for selecting key parameters related to lost circulation based on data fusion processing. The system for selecting key parameters related to lost circulation includes: an outlier determination module, an outlier elimination module, a smoothing processing module, and a key parameter extraction module.
[0173] The outlier determination module is used to screen outliers for the drilling time series data using at least three screening methods to obtain at least three outlier groups. The outlier elimination module is used to eliminate target outliers from the drilling time series data based on the outlier groups to obtain preprocessed drilling data. The smoothing processing module is used to perform smoothing processing on the preprocessed drilling data to obtain smoothed drilling data. The key parameter extraction module is used to perform feature extraction processing on the smoothed drilling data to obtain key parameters related to lost circulation.
[0174] In summary, the present invention provides a method and system for selecting key parameters related to lost circulation based on data fusion processing. Compared with the prior art, it has the following advantages:
[0175] The present invention improves the data quality of the drilling data set by fusing at least three different screening methods to process the drilling data set and simultaneously combining a filtering algorithm to perform smoothing processing on the drilling data set, and then selects key parameters related to lost circulation based on the principal component analysis method, greatly improving the accuracy of selecting key parameters related to lost circulation.
[0176] It should be understood that the embodiments disclosed in the present invention are not limited to the specific structures, processing steps or materials disclosed herein, but should extend to equivalent alternatives of these features understood by those of ordinary skill in the relevant art. It should also be understood that the terms used herein are only for the purpose of describing specific embodiments and do not mean to limit.
[0177] In the description of the present invention, unless otherwise specified, "a plurality of" means two or more; the orientation or positional relationship indicated by the terms "upper", "lower", "left", "right", "inner", "outer", "front end", "rear end", "head", "tail", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", "third", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0178] In the description of the present invention, it should be noted that, unless otherwise clearly specified and defined, the terms "connected" and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0179] Certain terms are used throughout this application to refer to particular system components. As those skilled in the art will recognize, the same components may typically be referred to by different names, and thus this application is not intended to distinguish between components that differ only in name and not in function. In this application, the terms "comprise", "include" and "have" are used in an open-ended fashion, and thus are to be interpreted to mean "including but not limited to...". Additionally, the terms "substantially", "essentially" or "approximately" as may be used herein refer to the industry-accepted tolerances for the corresponding terms. The term "coupled" as may be used herein includes direct coupling and indirect coupling via another component, element, circuit, or module, where for indirect coupling, the intervening component, element, circuit, or module does not change the information of the signal but may adjust its current level, voltage level, and / or power level. Inferred coupling (e.g., where one element is inferred to be coupled to another element) includes direct and indirect coupling between two elements in the same manner as "coupled".
[0180] The "one embodiment" or "embodiment" mentioned in the specification means that the specific features, structures or characteristics described in connection with the embodiment are included in at least one embodiment of the present invention. Therefore, the phrases "one embodiment" or "embodiment" that appear throughout the specification do not necessarily all refer to the same embodiment.
[0181] The embodiments of the present invention are given for purposes of illustration and description, and are not exhaustive or limit the invention to the disclosed form. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are chosen and described in order to better illustrate the principles of the invention and its practical application, and to enable those of ordinary skill in the art to understand the invention and design various embodiments with various modifications suitable for a particular purpose.
[0182] Although the embodiments disclosed in the present invention are as above, the content described is only an embodiment adopted for the convenience of understanding the present invention and is not intended to limit the present invention. Any person skilled in the art within the technical field to which the present invention pertains may make any modifications and changes in the form of implementation and details without departing from the spirit and scope disclosed in the present invention. However, the scope of patent protection of the present invention shall still be subject to the scope defined by the appended claims.
Claims
1. A method for selecting key parameters related to lost circulation based on data fusion processing, characterized in that, the method includes: S1. For the drilling time series data, use at least three screening methods to screen out outliers, and obtain at least three outlier groups; S2. Based on the outlier groups, eliminate the target outliers from the drilling time series data to obtain preprocessed drilling data; S3. Smooth the preprocessed drilling data to obtain smoothed drilling data; S4. Perform feature extraction processing on the smoothed drilling data to obtain key parameters related to lost circulation.
2. The method for selecting key parameters related to lost circulation based on data fusion processing according to claim 1, characterized in that, in step S1, obtain the drilling time series data, wherein the drilling time series data includes multiple sets of column data, and a single column data is a time series data of a drilling characteristic parameter.
3. The method for selecting key parameters related to lost circulation based on data fusion processing according to claim 2, characterized in that, in step S1, perform outlier screening through the first screening method to obtain the first outlier group, wherein the first screening method includes the following steps: S111. Calculate the first average value of each column data in the drilling time series data; S112. For each measured value in each column data, calculate the residual error corresponding to each measured value according to the corresponding first average value; S113. Calculate the standard deviation of each column data according to Bessel's formula; S114. For each measured value, determine whether the measured value is an outlier according to the corresponding residual error and the standard deviation; S115. Form the first outlier group with the measured values determined to be the outliers.
4. The method for selecting key parameters related to lost circulation based on data fusion processing according to claim 2 or 3, characterized in that, in step S1, perform outlier screening through the second screening method to obtain the second outlier group, wherein the second screening method includes the following steps: S121. For each column data, regard the measured values greater than the corresponding preset measured value in the column data as suspicious values; S122. Calculate the second average value and the average deviation of the remaining data, wherein the remaining data is the data in the column data except the suspicious values; S123. Determine whether the suspicious value is the outlier according to the second average value and the average deviation; S124. Form the second outlier group with the suspicious values determined to be the outliers.
5. The method for selecting key parameters related to lost circulation based on data fusion processing according to any one of claims 2-4, characterized in that, in step S1, perform outlier screening through the third screening method to obtain the third outlier group, wherein the third screening method includes the following steps: S131. For each column data, randomly select a measured value from the column data, and record it as the selected measured value; S132, sorting the measurement values in each column of data according to size to obtain sorted data; S133, determining the significance level value of the statistical test for detecting outliers; S134, determining an anomaly detection critical value according to the significance level value and the sample number of the measurement values in each column of data; S135. For each of the selected measurement values in the column data, a first difference and a third difference or a second difference and a third difference are calculated, wherein the first difference is a difference between the selected measurement value and a previous adjacent measurement value, the second difference is a difference between the second measurement value and the first measurement value in the sorted data, and the third difference is a difference between the selected measurement value and the first measurement value; S136, determining whether the currently selected measurement value is the abnormal value according to the first difference and the third difference or the second difference and the third difference in combination with the abnormality detection critical value, and obtaining a judgment result; S137, when the judgment result is yes, the selected measurement value is removed from the corresponding column data, and the column data after the removal is set as the column data, and the process returns to step S131 until all the measurement values in the column data are traversed; S138, when the judgment result is no, return to step S131; S139. All the selected measurement values determined as the abnormal values are grouped into the third abnormal value group.
6. A method for selecting key parameters associated with lost circulation based on data fusion processing according to any one of claims 1 to 5, It is characterized in that In step S2, outliers in different outlier value groups are compared to obtain the target outlier value, wherein the target outlier value is an outlier value that exists in at least any two of the outlier value groups.
7. A method for selecting key parameters associated with lost circulation based on data fusion processing according to any one of claims 1 to 6, It is characterized in that In step S3, the smoothed drilling data is obtained by the following steps: S31, using the optimal estimated value X(k-1|k-1) of the drilling system at time k-1 to calculate the predicted value X(k|k-1) of the drilling system at time k; the optimal estimated value of the drilling system corresponding to the initial time is a unit matrix; the drilling system is a system for collecting the drilling data; S32, using the covariance matrix P(k-1|k-1) at time k-1 to calculate the covariance matrix P(k|k-1) corresponding to the predicted value X(k|k-1) at time k; S33, calculating the optimal estimate value X(k|k) at time k according to the predicted value X(k|k-1) at time k, the measured value Z(k) at time k and the Kalman gain at time k; S34, calculating the covariance matrix P(k|k) corresponding to the optimal estimate X(k|k) at the k moment according to the Kalman gain at the k moment and the covariance matrix P(k|k-1); S35, using the optimal estimated value X(k|k) at the time k to smooth the measured value at the time k; S36. Let \(k = k + 1\), and return to step S31 until all the preprocessed drilling data are smoothed.
8. A method for selecting key parameters related to lost circulation based on data fusion processing according to any one of claims 1 - 7, characterized in that in step S4, the key parameters related to lost circulation are obtained through the following steps: S41. Construct a mathematical model for principal component analysis of the lost circulation data set; S42. Input the smoothed drilling data into a preset software; S43. Select the column data in the smoothed drilling data where the measured values are not all zero to obtain the selected drilling characteristic parameters; S44. Select the correlation test option to confirm whether each of the selected drilling characteristic parameters is suitable for principal component analysis; S45. When passing the correlation test, perform principal component analysis on each of the selected drilling characteristic parameters to obtain a scree plot curve, the eigenvalues of each principal component, and the cumulative variance contribution rate; S46. Select the target principal component from each of the principal components according to the scree plot curve, the eigenvalues, and the cumulative variance contribution rate; S47. Extract each principal component and common factor and perform factor rotation analysis, and then output the total variance explained data and the component score coefficient matrix; S48. Substitute the value of the sum of squared rotation loadings in the total variance explained data and the values in the component score coefficient matrix into the mathematical model for principal component analysis of the lost circulation data set; S49. Determine the key parameters related to lost circulation according to the mathematical model for principal component analysis of the lost circulation data set after substituting the values.
9. A storage medium, characterized in that it contains a series of instructions for executing the method steps according to any one of claims 1 - 8.
10. A system for selecting key parameters related to lost circulation based on data fusion processing, characterized in that it executes the method according to any one of claims 1 - 8, and the system for selecting key parameters related to lost circulation includes: An outlier determination module, which is used to screen outliers for drilling time series data using at least three screening methods to obtain at least three outlier groups; An outlier removal module, which is used to remove target outliers from the drilling time series data based on the outlier groups to obtain preprocessed drilling data; A smoothing processing module, which is used to smooth the preprocessed drilling data to obtain smoothed drilling data; A key parameter extraction module, which is used to perform feature extraction processing on the smoothed drilling data to obtain key parameters related to lost circulation.