Quality evaluation model construction method, quality evaluation method and device for drilling data
By constructing a drilling data quality evaluation model based on Birch algorithm, the inefficiency problem in the existing technology that relies on manual experience is solved, and more efficient and accurate data quality evaluation is achieved.
Patent Information
- Application Number
- CN202311572660.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, drilling data quality evaluation mainly relies on manual experience judgment, which is low in efficiency and inaccurate enough.
By obtaining historical drilling data, performing normal transformation and standardization processing, it is divided into several historical data sets, using the Birch algorithm to determine outliers and default values, calculate quality index values, and build a quality evaluation model for training to achieve automated data quality evaluation.
It improves the efficiency and accuracy of drilling data quality evaluation, and is more automated and reliable than manual judgment.
Smart Images

Figure CN120031433A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of oil exploration, and in particular, to a method for constructing a quality evaluation model for drilling data, a quality evaluation method and a device. Background Art
[0002] In drilling engineering, there are a lot of real-time drilling data, such as ground engineering parameters, downhole engineering parameters, etc. Real-time identification of drilling conditions, real-time evaluation of drilling effects, and real-time optimization of drilling parameters all require calculation and analysis of these real-time drilling data. Due to inherent errors in data acquisition devices, signal fluctuations during data transmission, subjectivity in manual processing, equipment failures, and other reasons, these drilling data may contain some abnormal values or erroneous values, which will affect the reliability and accuracy of subsequent calculation and analysis results. Therefore, it is necessary to pre-process the drilling data and evaluate its data quality. At present, the evaluation of drilling data quality is based on manual experience judgment, which is inefficient.
[0003] Therefore, there is an urgent need for a method to construct a quality evaluation model for drilling data, which can evaluate the treatment of drilling data more efficiently. Summary of the invention
[0004] The purpose of the embodiments of this specification is to provide a method for constructing a quality evaluation model for drilling data, a quality evaluation method and a device, so as to more efficiently evaluate the treatment of drilling data.
[0005] To achieve the above objectives, on the one hand, an embodiment of this specification provides a method for constructing a quality evaluation model for drilling data, comprising:
[0006] Get historical drilling data;
[0007] Performing normal transformation and standardization processing on the historical drilling data to obtain processed historical drilling data;
[0008] Dividing the processed historical drilling data into a plurality of historical data sets, wherein a quality evaluation result of each historical data set is known;
[0009] The Birch algorithm is used to determine the outliers and default values of each historical data set;
[0010] Calculating at least one quality indicator value corresponding to each historical data set according to the abnormal value and the default value;
[0011] A quality evaluation model is constructed, and the quality indicator values corresponding to all historical data sets are used as training sets to train the quality evaluation model to obtain a trained quality evaluation model.
[0012] Preferably, the at least one quality indicator value corresponding to each historical data set includes: missing rate, abnormality rate, data consistency, data timeliness, data credibility and data redundancy.
[0013] Preferably, the missing rate is calculated by the following formula:
[0014]
[0015] Among them, Y 1 is the missing rate of the historical data set, x is the number of default values of the numeric type in the historical data set, a is the total number of data in the historical data set, b is the number of default values of the string type in the historical data set, and c is the number of redundant data in the historical data set.
[0016] Preferably, the abnormality rate is calculated by the following formula:
[0017]
[0018] Among them, Y 2 is the anomaly rate of the historical data set, d is the number of anomalies in the historical data set, a is the total number of data in the historical data set, b is the number of default values of the string type in the historical data set, and c is the number of redundant data in the historical data set.
[0019] Preferably, the data consistency is calculated by the following formula:
[0020]
[0021] Among them, Y 3 is the data consistency of the historical data set, e is the number of anomalies in the reference consistency of different layers in the same well, f is the number of anomalies in the reference consistency of the same layer in different wells, g is the number of repetitions in the anomalies of e and f, a is the total number of data in the historical data set, b is the number of default values of string type in the historical data set, and c is the number of redundant data in the historical data set.
[0022] Preferably, the data timeliness is calculated by the following formula:
[0023]
[0024] Among them, Y 4 is the timeliness of the historical data set, a is the total number of data in the historical data set, b is the number of default values of the string type in the historical data set, and c is the number of redundant data in the historical data set.
[0025] Preferably, the data credibility is calculated by the following formula:
[0026] Y 5 =10×log10 (Ps / Pn);
[0027] Among them, Y 5 is the data credibility of the historical data set, Ps is the average power of the signal, and Pn is the average power of the noise.
[0028] Preferably, the data redundancy is calculated by the following formula:
[0029]
[0030] Among them, Y 6 is the data redundancy of the historical data set, i is the missing rate of the historical data set, β is the balance coefficient, the value range is (0,1), and c is the number of redundant data in the historical data set.
[0031] On the other hand, the embodiment of this specification also provides a method for evaluating the quality of drilling data, including:
[0032] Obtain drilling data to be evaluated;
[0033] Performing normal transformation and standardization processing on the drilling data to be evaluated to obtain processed drilling data to be evaluated;
[0034] Dividing the processed drilling data to be evaluated into a plurality of data sets to be evaluated;
[0035] The Birch algorithm was used to determine the outliers and default values of each data set to be evaluated;
[0036] Calculating at least one quality indicator value to be evaluated corresponding to each data set to be evaluated according to the abnormal value and the default value of each data set to be evaluated;
[0037] The quality index value to be evaluated is input into the trained quality evaluation model to obtain the quality of the drilling data to be evaluated, wherein the quality evaluation model is constructed using any of the above methods.
[0038] On the other hand, an embodiment of the present specification provides a device for constructing a quality evaluation model for drilling data, the device comprising:
[0039] A first acquisition module is used to acquire historical drilling data;
[0040] A first processing module is used to perform normal transformation and standardization processing on the historical drilling data to obtain processed historical drilling data;
[0041] A first partitioning module is used to partition the processed historical drilling data into a plurality of historical data sets, wherein a quality evaluation result of each historical data set is known;
[0042] A first determination module is used to determine the abnormal value and the default value of each historical data set using the Birch algorithm;
[0043] A first calculation module, used for calculating at least one quality indicator value corresponding to each historical data set according to the abnormal value and the default value;
[0044] The construction module is used to construct a quality evaluation model, and the quality indicator values corresponding to all historical data sets are used as training sets to train the quality evaluation model to obtain a trained quality evaluation model.
[0045] In another aspect, the embodiments of this specification further provide a drilling data quality evaluation device, the device comprising:
[0046] A second acquisition module is used to acquire drilling data to be evaluated;
[0047] A second processing module is used to perform normal transformation and standardization processing on the drilling data to be evaluated to obtain processed drilling data to be evaluated;
[0048] A second division module is used to divide the processed drilling data to be evaluated into a plurality of data sets to be evaluated;
[0049] The second determination module is used to determine the abnormal value and the default value of each data set to be evaluated by using the Birch algorithm;
[0050] A second calculation module, used for calculating at least one quality indicator value to be evaluated corresponding to each data set to be evaluated according to the abnormal value and the default value of each data set to be evaluated;
[0051] An evaluation module is used to input the quality index value to be evaluated into a trained quality evaluation model to obtain the quality of the drilling data to be evaluated, wherein the quality evaluation model is constructed using the above-mentioned device.
[0052] On the other hand, an embodiment of the present specification further provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein when the computer program is executed by the processor, the instructions according to any one of the methods described above are executed.
[0053] On the other hand, an embodiment of the present specification further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor of a computer device, the computer program executes instructions according to any one of the methods described above.
[0054] It can be seen from the technical solutions provided in the above embodiments of this specification that, through the method of the embodiments of this specification, after obtaining the historical drilling data, it is subjected to normal transformation and standardization processing, and then divided into several historical data sets, and then the abnormal values and default values of each historical data set are determined, and at least one quality indicator value corresponding to each historical data set is calculated according to the abnormal values and the default values, and then a quality evaluation model is constructed, and the quality indicator values corresponding to all historical data sets are used as training sets to train the quality evaluation model, and the trained quality evaluation model is obtained. The quality evaluation model can be used to evaluate the quality of the drilling data, which is more efficient and more accurate than manual experience judgment.
[0055] In order to make the above and other purposes, features and advantages of the present specification more obvious and easy to understand, the following specifically cites preferred embodiments and describes them in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0057] Figure 1 A flow chart of a method for constructing a quality evaluation model for drilling data provided in an embodiment of this specification is shown;
[0058] Figure 2 A flow chart of a method for evaluating the quality of drilling data provided in an embodiment of this specification is shown;
[0059] Figure 3 A schematic diagram of the module structure of a device for constructing a quality evaluation model for drilling data provided by an embodiment of this specification is shown;
[0060] Figure 4 A schematic diagram of the module structure of a drilling data quality evaluation device provided in an embodiment of this specification is shown;
[0061] Figure 5 A schematic diagram of the structure of a computer device provided in an embodiment of this specification is shown.
[0062] Description of the accompanying symbols:
[0063] 100. A first acquisition module;
[0064] 200, first processing module;
[0065] 300, first division module;
[0066] 400, first determination module;
[0067] 500. A first computing module;
[0068] 600, building blocks;
[0069] 700, second acquisition module;
[0070] 800, second processing module;
[0071] 900, second division module;
[0072] 1000. Second determination module;
[0073] 1100. Second calculation module;
[0074] 1200, evaluation module;
[0075] 502. Computer equipment;
[0076] 504, processor;
[0077] 506. Memory;
[0078] 508, driving mechanism;
[0079] 510, input / output module;
[0080] 512. Input devices;
[0081] 514. Output device;
[0082] 516. Presentation equipment;
[0083] 518. Graphical user interface;
[0084] 520, network interface;
[0085] 522, communication link;
[0086] 524. Communication bus. DETAILED DESCRIPTION
[0087] The following will be combined with the drawings in the embodiments of this specification to clearly and completely describe the technical solutions in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the embodiments of this specification.
[0088] In drilling engineering, there are a lot of real-time drilling data, such as ground engineering parameters, downhole engineering parameters, etc. Real-time identification of drilling conditions, real-time evaluation of drilling effects, and real-time optimization of drilling parameters all require calculation and analysis of these real-time drilling data. Due to inherent errors in data acquisition devices, signal fluctuations during data transmission, subjectivity in manual processing, equipment failures, and other reasons, these drilling data may contain some abnormal values or erroneous values, which will affect the reliability and accuracy of subsequent calculation and analysis results. Therefore, it is necessary to pre-process the drilling data and evaluate its data quality. At present, the evaluation of drilling data quality is based on manual experience judgment, which is inefficient.
[0089] In order to solve the above problems, an embodiment of this specification provides a method for constructing a quality evaluation model for drilling data. Figure 1 This is a flowchart of a method for constructing a quality evaluation model for drilling data provided by an embodiment of this specification. This specification provides method operation steps as described in the embodiment or flowchart, but more or fewer operation steps may be included based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the order of many steps and does not represent the only execution order. When the system or device product is executed in practice, it can be executed in the order of the method shown in the embodiment or the accompanying drawings or in parallel.
[0090] It should be noted that the terms "first", "second", etc. in the description and claims of the embodiments of this specification and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of this specification described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, device, product or equipment that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0091] Reference Figure 1 The embodiment of this specification provides a method for constructing a quality evaluation model for drilling data, including:
[0092] S101: Acquire historical drilling data;
[0093] S102: performing normal transformation and standardization processing on the historical drilling data to obtain processed historical drilling data;
[0094] S103: Divide the processed historical drilling data into a plurality of historical data sets, wherein a quality evaluation result of each historical data set is known;
[0095] S104: using the Birch algorithm to determine the abnormal value and the default value of each historical data set;
[0096] S105: Calculating at least one quality indicator value corresponding to each historical data set according to the abnormal value and the default value;
[0097] S106: Construct a quality evaluation model, and use the quality indicator values corresponding to all historical data sets as training sets to train the quality evaluation model to obtain a trained quality evaluation model.
[0098] Historical drilling data refers to drilling data of known quality. The quality of historical drilling data can be obtained through manual annotation. Historical drilling data is generally obtained one by one according to the drilling time. One historical drilling data includes drilling pressure, drilling speed, hook load, rotary table torque, riser pressure (pump pressure), inlet flow, outlet flow, etc. After obtaining the historical drilling data, it needs to be normally transformed and standardized. The Box-Cox transformation method can be used to perform normal transformation, and the data after normal transformation is further subjected to Z-score standardization to obtain the processed historical drilling data. Since most field data are non-normally distributed, and most quality evaluation models require data to be normally distributed in order to obtain better model convergence speed and prediction effect. The purpose of standardization is to remove dimensionality problems between different drilling data and ensure the accuracy of subsequent analysis.
[0099] The processed historical drilling data is divided to obtain several historical data sets, where the historical data sets are historical data sets formed by the values of certain data at different times within a specified historical time period, such as historical data set A formed by all drilling pressures within a specified historical time period, and historical data set B formed by all rotation speeds within a specified historical time period. Since the historical drilling data are drilling data of known quality, the quality evaluation results of the corresponding historical data sets are also known. Since the quality of the historical drilling data can be obtained through manual annotation, such as 0-0.5 represents poor, 0.5-0.9 represents good, and 0.9-1.0 represents excellent, the scores of different historical drilling data are manually annotated. For the historical data set, the average score of all data in the historical data set is calculated, and the average score is used as the actual quality evaluation result corresponding to the historical data set.
[0100] The Birch algorithm is further used to determine the abnormal values and default values of each historical data set. The Birch algorithm can determine the cluster feature tree (CF Tree). The specific steps are as follows:
[0101] Step 1: Manually set three CF Tree parameters: the maximum number of non-leaf nodes B, the maximum number of CFs L contained in each leaf node, and the maximum radius threshold T for each CF in the leaf node.
[0102] For the historical dataset D, read the p-th data d in the historical dataset D p (0 < p ≤ n) as a sample point and include it in the triple LN p , where each CF is a triple, which can be represented by (N, LS, SS). Here, N represents the number of sample points in the CF, LS represents the sum vector of the feature dimensions of the sample points in the CF, and SS represents the sum of the squares of the feature dimensions of the sample points in the CF.
[0103] Step 2: Read the next data d in the historical dataset D p+1 (0 < p + 1 ≤ n) as a sample point. If the CF node distance of this sample point is still less than the maximum radius threshold T, include it in the triple to which data d p belongs and go to Step 5. Otherwise, go to Step 3.
[0104] Step 3: If the number of CF nodes of this sample point is less than the threshold L, create a new leaf node and put the next data d p+1 into it. This new leaf node is used to store the new CF node, update all the CF triples on the path, generate the new triple LN p+1 , and go to Step 5. Otherwise, go to Step 4.
[0105] Step 4: Divide the current leaf node to obtain two new leaf nodes. Among all the CF triples in the old leaf node, select the two CF triples with the farthest distance within the hypersphere as the first CF nodes of the two new leaf nodes respectively. Then, according to the distance, put the other tuples closest to the corresponding first CF node into the corresponding leaf nodes to perform the splitting of the leaf nodes. Check the parent node upward in turn to determine whether the parent node also needs to be split. If so, its splitting method is the same as that of the leaf node. If not, go to Step 5.
[0106] Step 5: Read the remaining data d one by one t (p + 1 < t ≤ n), and repeat Steps 2 - 5. Complete the construction of the clustering feature tree CF Tree for all data.
[0107] Through the clustering feature tree, dense data is divided into clusters, and sparse data is regarded as outliers. The dense data here can be defined as follows: if the data of any field in the historical data set (limited to the data of one field, such as drilling pressure) is selected, the data has an intersection with the historical data set in the neighborhood without the center, that is dense data. On the contrary, if the data of any field in the historical data set is selected, the data has no intersection with the historical data set in the neighborhood without the center, that is sparse data.
[0108] The number of abnormal values and default values is further counted. For the default value, if the data in the historical data set is of string type, the default value representation is that the string is empty (NULL); if the data in the historical data set is of numeric type, the default value representation is that the numeric value is 0.
[0109] Based on the abnormal values and default values, at least one quality indicator value corresponding to each historical data set can be calculated, and a quality evaluation model can be constructed. The average score is used as the actual quality evaluation result corresponding to the historical data set, and the quality indicator values corresponding to all historical data sets are input into the quality evaluation model to obtain the predicted quality evaluation result. The difference between the actual quality evaluation result and the predicted quality evaluation result is calculated, and the quality evaluation model is further optimized. The above training process is continuously iterated to obtain the trained quality evaluation model.
[0110] Through the method of the embodiment of this specification, after obtaining the historical drilling data, it is subjected to normal transformation and standardization processing, and then divided into several historical data sets, and then the abnormal values and default values of each historical data set are determined, and at least one quality indicator value corresponding to each historical data set is calculated according to the abnormal values and the default values, and then a quality evaluation model is constructed, and the quality indicator values corresponding to all historical data sets are used as training sets to train the quality evaluation model to obtain a trained quality evaluation model. The quality evaluation model can be used to evaluate the quality of the drilling data, which is more efficient and more accurate than manual experience judgment.
[0111] Among them, the at least one quality indicator value corresponding to each historical data set includes: missing rate, abnormality rate, data consistency, data timeliness, data credibility and data redundancy.
[0112] In the examples of this specification, the missing rate is calculated by the following formula:
[0113]
[0114] Among them, Y 1is the missing rate of the historical data set, x is the number of default values of numeric types in the historical data set, a is the total number of data in the historical data set, b is the number of default values of string types in the historical data set, and c is the number of redundant data in the historical data set. It should be noted that redundant data refers to the number of repeated data in each historical data set.
[0115] The abnormality rate is calculated by the following formula:
[0116]
[0117] Among them, Y 2 is the anomaly rate of the historical data set, d is the number of anomalies in the historical data set, a is the total number of data in the historical data set, b is the number of default values of the string type in the historical data set, and c is the number of redundant data in the historical data set.
[0118] The data consistency is calculated by the following formula:
[0119]
[0120] Among them, Y 3 is the data consistency of the historical data set, e is the number of anomalies of reference consistency in different layers of the same well, f is the number of anomalies of logical consistency in the same layer of different wells, g is the number of repetitions in the anomalies of e and f, a is the total number of data in the historical data set, b is the number of default values of string type in the historical data set, and c is the number of redundant data in the historical data set.
[0121] The meanings of e and f are further explained by examples: for the same drilling data, such as drilling pressure data, the data in the current historical data set is the drilling pressure data at a depth of 1000-1200 meters. The drilling pressure data at a depth of 1000-1200 meters in the current well and the drilling pressure data at a depth of 2300-2500 meters (different layers in the same well) are obtained, and the curves obtained by regressing the drilling pressure data at the two locations are compared. The drilling pressure data that is different between the drilling pressure data at a depth of 1000-1200 meters in the current well and the drilling pressure data at a depth of 2300-2500 meters is determined as abnormal data e. Obtain the drilling pressure data at a depth of 1000-1200 meters in the current well, and compare it with the drilling pressure data at a depth of 1000-1200 meters in the adjacent well (different wells and the same layer), and compare the curves obtained by regressing the drilling pressure data of the two places. The drilling pressure data that is different from the drilling pressure data at a depth of 1000-1200 meters in the current well and the drilling pressure data at a depth of 1000-1200 meters in the current well is determined as abnormal data f.
[0122] The data timeliness is calculated by the following formula:
[0123]
[0124] Among them, Y 4 is the timeliness of the historical data set, a is the total number of data in the historical data set, b is the number of default values of the string type in the historical data set, and c is the number of redundant data in the historical data set. It should be noted that the historical drilling data may not be timely when it is acquired. For example, if the drill is pulled out during the acquisition process, there is no historical drilling data in the corresponding time period in the acquired historical drilling data. The corresponding time period in the historical data set corresponds to the default value of the string type.
[0125] The data credibility is calculated by the following formula:
[0126] Y 5 =10×log 10 (Ps / Pn);
[0127] Among them, Y 5 is the data credibility of the historical data set, Ps is the average power of the signal, and Pn is the average power of the noise. It should be noted that the signal refers to the drilling signal, and the noise refers to the interference intensity of various other instruments on site on the downhole mud pulse.
[0128] The data redundancy is calculated by the following formula:
[0129]
[0130] Among them, Y 6 is the data redundancy of the historical data set, i is the missing rate of the historical data set, β is the balance coefficient, the value range is (0,1), and c is the number of redundant data in the historical data set.
[0131] In the embodiment of the present specification, the quality evaluation model is a probabilistic neural network model (PNN), in which there are 6 neurons in the input layer, 9 neurons in the sample layer, 3 neurons in the summation layer, and 1 neuron in the output layer. Each historical data set corresponds to the above six quality indicator values, and the six quality indicator values corresponding to the historical data set are used as training sets to train the quality evaluation model, and the calculated average score is used as the actual quality evaluation result corresponding to the historical data set until the conformity rate of the predicted quality evaluation result of the quality evaluation model reaches more than 80-90%, and the training is terminated.
[0132] Reference Figure 2 Based on the above-mentioned method for constructing a quality evaluation model for drilling data, the embodiment of this specification further provides a method for evaluating the quality of drilling data, the method comprising:
[0133] S201: Acquire drilling data to be evaluated;
[0134] S202: performing normal transformation and standardization processing on the drilling data to be evaluated to obtain processed drilling data to be evaluated;
[0135] S203: Divide the processed drilling data to be evaluated into a plurality of data sets to be evaluated;
[0136] S204: using the Birch algorithm to determine the outliers and default values of each data set to be evaluated;
[0137] S205: Calculating at least one quality indicator value to be evaluated corresponding to each data set to be evaluated according to the abnormal value and the default value of each data set to be evaluated;
[0138] S206: Inputting the quality index value to be evaluated into the trained quality evaluation model to obtain the quality of the drilling data to be evaluated, wherein the quality evaluation model is constructed using the above method.
[0139] Specifically, the drilling data to be evaluated and the quality of the drilling data to be evaluated are stored in a database for subsequent analysis and decision-making by on-site personnel. In the embodiment of this specification, data storage is performed using MySQL database technology.
[0140] Based on the above-mentioned method for constructing a quality evaluation model of drilling data, the embodiment of this specification also provides a device for constructing a quality evaluation model of drilling data. The device may include a system (including a distributed system), software (application), module, component, server, client, etc. using the method described in the embodiment of this specification and a device combined with necessary implementation hardware. Based on the same innovative concept, the device in one or more embodiments provided in the embodiment of this specification is as described in the following embodiments. Since the implementation scheme and method for solving the problem of the device are similar, the implementation of the specific device of the embodiment of this specification can refer to the implementation of the aforementioned method, and the repetitions will not be repeated. As used below, the term "unit" or "module" can implement a combination of software and / or hardware of predetermined functions. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.
[0141] Specifically, Figure 3 This is a schematic diagram of the module structure of an embodiment of a drilling data quality evaluation model construction device provided in the embodiment of this specification, referring to Figure 3 As shown, a drilling data quality evaluation model construction device provided in an embodiment of the present specification includes: a first acquisition module 100, a first processing module 200, a first division module 300, a first determination module 400, a first calculation module 500, and a construction module 600.
[0142] A first acquisition module 100 is used to acquire historical drilling data;
[0143] The first processing module 200 is used to perform normal transformation and standardization processing on the historical drilling data to obtain processed historical drilling data;
[0144] A first partitioning module 300 is used to partition the processed historical drilling data into a plurality of historical data sets, wherein a quality evaluation result of each historical data set is known;
[0145] A first determination module 400 is used to determine abnormal values and default values of each historical data set using a Birch algorithm;
[0146] A first calculation module 500, configured to calculate at least one quality indicator value corresponding to each historical data set according to the abnormal value and the default value;
[0147] The construction module 600 is used to construct a quality evaluation model, and train the quality evaluation model using the quality indicator values corresponding to all historical data sets as training sets to obtain a trained quality evaluation model.
[0148] Reference Figure 4 Based on the above-mentioned method for evaluating the quality of drilling data, the embodiment of this specification also provides a device for evaluating the quality of drilling data. The device includes: a second acquisition module 700, a second processing module 800, a second division module 900, a second determination module 1000, a second calculation module 1100, and an evaluation module 1200.
[0149] The second acquisition module 700 is used to acquire the drilling data to be evaluated;
[0150] The second processing module 800 is used to perform normal transformation and standardization processing on the drilling data to be evaluated to obtain processed drilling data to be evaluated;
[0151] A second division module 900 is used to divide the processed drilling data to be evaluated into a plurality of data sets to be evaluated;
[0152] The second determination module 1000 is used to determine the abnormal value and the default value of each data set to be evaluated by using the Birch algorithm;
[0153] A second calculation module 1100, configured to calculate at least one quality indicator value to be evaluated corresponding to each data set to be evaluated according to the abnormal value and the default value of each data set to be evaluated;
[0154] The evaluation module 1200 is used to input the quality indicator value to be evaluated into the trained quality evaluation model to obtain the quality of the drilling data to be evaluated, wherein the quality evaluation model is constructed using the device described in claim 10.
[0155] Reference Figure 5 As shown, based on the above-mentioned method for constructing a quality assessment model for drilling data and a method for assessing the quality of drilling data, a computer device 502 is also provided in an embodiment of this specification, wherein the above-mentioned method is run on the computer device 502. The computer device 502 may include one or more processors 504, such as one or more central processing units (CPUs) or graphics processing units (GPUs), and each processing unit may implement one or more hardware threads. The computer device 502 may also include any memory 506, which is used to store any kind of information such as code, settings, data, etc. In a specific embodiment, a computer program on the memory 506 and executable on the processor 504, when the computer program is executed by the processor 504, can execute instructions according to the above-mentioned method. Non-limitingly, for example, the memory 506 may include any one or more combinations of the following: any type of RAM, any type of ROM, flash memory device, hard disk, optical disk, etc. More generally, any memory may use any technology to store information. Further, any memory may provide volatile or non-volatile retention of information. Further, any memory may represent a fixed or removable component of the computer device 502. In one embodiment, when the processor 504 executes the associated instructions stored in any memory or combination of memories, the computer device 502 can perform any operation of the associated instructions. The computer device 502 also includes one or more drive mechanisms 508 for interacting with any memory, such as a hard disk drive mechanism, an optical disk drive mechanism, etc.
[0156] The computer device 502 may also include an input / output module 510 (I / O) for receiving various inputs (via input devices 512) and for providing various outputs (via output devices 514). A specific output mechanism may include a presentation device 516 and an associated graphical user interface 518 (GUI). In other embodiments, the input / output module 510 (I / O), input device 512, and output device 514 may not be included, and the computer device 502 may be used as a computer device in a network. The computer device 502 may also include one or more network interfaces 520 for exchanging data with other devices via one or more communication links 522. One or more communication buses 524 couple the components described above together.
[0157] The communication link 522 may be implemented in any manner, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 522 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.
[0158] Corresponds to Figure 1-Figure 2 The method in the embodiment of the present specification also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are executed.
[0159] The embodiment of the present specification also provides a computer-readable instruction, wherein when the processor executes the instruction, the program therein causes the processor to execute the following Figure 1 to Figure 2 The method shown.
[0160] It should be understood that in the various embodiments of this specification, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this specification.
[0161] It should also be understood that in the embodiments of this specification, the term "and / or" is only a description of the association relationship of the associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the embodiments of this specification generally indicates that the associated objects before and after are in an "or" relationship.
[0162] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the embodiments of this specification can be implemented with electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of this specification.
[0163] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0164] In the several embodiments provided in this specification, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or it can be an electrical, mechanical or other form of connection.
[0165] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of this specification.
[0166] In addition, each functional unit in each embodiment of this specification may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0167] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of this specification is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of this specification. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code.
[0168] Specific embodiments are used in this specification to illustrate the principles and implementation methods of the embodiments of this specification. The description of the above embodiments is only used to help understand the methods and core ideas of the embodiments of this specification. At the same time, for those skilled in the art, according to the ideas of the embodiments of this specification, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the embodiments of this specification.
Claims
1. A method for constructing a quality evaluation model for drilling data. It is characterized in that include: Get historical drilling data; Performing normal transformation and standardization processing on the historical drilling data to obtain processed historical drilling data; Dividing the processed historical drilling data into a plurality of historical data sets, wherein a quality evaluation result of each historical data set is known; The Birch algorithm is used to determine the outliers and default values of each historical data set; Calculating at least one quality indicator value corresponding to each historical data set according to the abnormal value and the default value; A quality evaluation model is constructed, and the quality indicator values corresponding to all historical data sets are used as training sets to train the quality evaluation model to obtain a trained quality evaluation model.
2. The method according to claim 1, It is characterized in that The at least one quality indicator value corresponding to each historical data set includes: missing rate, abnormality rate, data consistency, data timeliness, data credibility and data redundancy.
3. The method according to claim 2, It is characterized in that The missing rate is calculated by the following formula: Among them, Y 1 is the missing rate of the historical data set, x is the number of default values of the numeric type in the historical data set, a is the total number of data in the historical data set, b is the number of default values of the string type in the historical data set, and c is the number of redundant data in the historical data set.
4. The method according to claim 2, It is characterized in that The abnormality rate is calculated by the following formula: Among them, Y 2 is the anomaly rate of the historical data set, d is the number of anomalies in the historical data set, a is the total number of data in the historical data set, b is the number of default values of the string type in the historical data set, and c is the number of redundant data in the historical data set.
5. The method according to claim 2, It is characterized in that The data consistency is calculated by the following formula: Among them, Y 3 is the data consistency of the historical data set, e is the number of anomalies in the reference consistency of different layers in the same well, f is the number of anomalies in the reference consistency of the same layer in different wells, g is the number of repetitions in the anomalies of e and f, a is the total number of data in the historical data set, b is the number of default values of string type in the historical data set, and c is the number of redundant data in the historical data set.
6. The method according to claim 2, It is characterized in that The data timeliness is calculated by the following formula: Among them, Y 4 is the timeliness of the historical data set, a is the total number of data in the historical data set, b is the number of default values of the string type in the historical data set, and c is the number of redundant data in the historical data set.
7. The method according to claim 2, It is characterized in that The data credibility is calculated by the following formula: Y 5 =10×log 10 (Ps / Pn); Among them, Y 5 is the data credibility of the historical data set, Ps is the average power of the signal, and Pn is the average power of the noise.
8. The method according to claim 2, It is characterized in that The data redundancy is calculated by the following formula: Among them, Y 6 is the data redundancy of the historical data set, i is the missing rate of the historical data set, β is the balance coefficient, the value range is (0,1), and c is the number of redundant data in the historical data set.
9. A method for evaluating the quality of drilling data. It is characterized in that include: Obtain drilling data to be evaluated; Performing normal transformation and standardization processing on the drilling data to be evaluated to obtain processed drilling data to be evaluated; Dividing the processed drilling data to be evaluated into a plurality of data sets to be evaluated; The Birch algorithm was used to determine the outliers and default values of each data set to be evaluated; Calculating at least one quality indicator value to be evaluated corresponding to each data set to be evaluated according to the abnormal value and the default value of each data set to be evaluated; The quality index value to be evaluated is input into the trained quality evaluation model to obtain the quality of the drilling data to be evaluated, wherein the quality evaluation model is constructed using the method described in any one of claims 1 to 8.
10. A device for constructing a quality evaluation model for drilling data, It is characterized in that The device comprises: A first acquisition module is used to acquire historical drilling data; A first processing module is used to perform normal transformation and standardization processing on the historical drilling data to obtain processed historical drilling data; A first partitioning module is used to partition the processed historical drilling data into a plurality of historical data sets, wherein a quality evaluation result of each historical data set is known; A first determination module is used to determine the abnormal value and the default value of each historical data set using the Birch algorithm; A first calculation module, used for calculating at least one quality indicator value corresponding to each historical data set according to the abnormal value and the default value; The construction module is used to construct a quality evaluation model, and the quality indicator values corresponding to all historical data sets are used as training sets to train the quality evaluation model to obtain a trained quality evaluation model.
11. A drilling data quality evaluation device, It is characterized in that The device comprises: A second acquisition module is used to acquire drilling data to be evaluated; A second processing module is used to perform normal transformation and standardization processing on the drilling data to be evaluated to obtain processed drilling data to be evaluated; A second division module is used to divide the processed drilling data to be evaluated into a plurality of data sets to be evaluated; The second determination module is used to determine the abnormal value and the default value of each data set to be evaluated by using the Birch algorithm; A second calculation module, used for calculating at least one quality indicator value to be evaluated corresponding to each data set to be evaluated according to the abnormal value and the default value of each data set to be evaluated; An evaluation module is used to input the quality indicator value to be evaluated into a trained quality evaluation model to obtain the quality of the drilling data to be evaluated, wherein the quality evaluation model is constructed using the device described in claim 10.
12. A computer device comprising a memory, a processor, and a computer program stored on the memory, It is characterized in that When the computer program is executed by the processor, the computer program executes the instructions of the method according to any one of claims 1 to 9.
13. A computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor of a computer device, the computer program executes the instructions of the method according to any one of claims 1 to 9.