Automatic cleaning method and device for abnormal value of logging data and electronic equipment

Through unsupervised anomaly data detection algorithm and binary tree splitting technology, outliers in the log data are automatically screened and eliminated, solving the problem of time-consuming and error-prone human editing, and improving the quality and processing efficiency of logging data.

CN120047547APending Publication Date: 2025-05-27CHINA PETROLEUM & CHEMICAL CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311595585.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

There are a large number of outliers in the existing logging data. Manual editing and eliminating abnormal data is time-consuming and prone to omissions and errors, which affects the quality of logging data.

Method used

Unsupervised anomaly data detection algorithm is used to randomly split the target log data in the multi-dimensional spatial logging data set by using a binary tree to filter out the outliers based on the formed binary tree structure, and the outliers are eliminated from the target logging curve.

Benefits of technology

Through automated data cleaning methods, the errors of manual editing are reduced, work efficiency is improved, and the quality of logging data processing is improved. It is suitable for pre-processing of large-scale logging data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047547A_ABST
    Figure CN120047547A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic cleaning method and device for abnormal values of well logging data and electronic equipment. The method comprises the steps that a plurality of target well logging curves of a target well form a multi-dimensional space well logging data set according to the actual sampling rate; performing random splitting on target logging data in the multi-dimensional space logging data set by using a binary tree by using an unsupervised abnormal data detection algorithm; screening out abnormal values of the logging data based on the formed binary tree structure; and removing abnormal values from the target logging curve to obtain the target logging curve after data cleaning. According to the invention, errors caused by manual editing and removal of abnormal logging data can be reduced, and the working efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of geophysical logging data preprocessing, and more specifically, relates to a method, device, and electronic device for automatically cleaning outliers in logging data. Background Art

[0002] With the acceleration of the oilfield exploration and development process and the increasing requirement for the accuracy of reservoir description, the role of logging data in reservoir prediction and reservoir description is becoming more prominent, and how to improve the quality of logging data is also becoming increasingly important. To improve the quality of logging data, outliers existing in the existing logging data are often removed through manual interactive editing. In the case of a large amount of logging data, manual interactive editing wastes a lot of time and there are situations of omission and editing errors. Summary of the Invention

[0003] The object of the present invention is to propose a method, device, and electronic device for automatically cleaning outliers in logging data, so as to reduce the error of manually editing and removing logging abnormal data and improve work efficiency.

[0004] To achieve the above object, in a first aspect, the present invention proposes a method for automatically cleaning outliers in logging data, including:

[0005] Composing a multi-dimensional space logging data set from several target logging curves of a target well according to the actual sampling rate;

[0006] Using an unsupervised outlier detection algorithm and randomly splitting the target logging data in the multi-dimensional space logging data set by using a binary tree;

[0007] Screening out outliers in the logging data based on the formed binary tree structure;

[0008] Removing the outliers from the target logging curves to obtain the target logging curves after data cleaning.

[0009] Optionally, the unsupervised outlier detection algorithm includes an isolation forest algorithm.

[0010] Optionally, the randomly splitting the target logging data in the multi-dimensional space logging data set by using a binary tree includes:

[0011] S01: Randomly selecting several data from the multi-dimensional space logging data set as a sample subset and putting the sample subset into the root node of the binary tree;

[0012] S02: Randomly generating a splitting point in the current node data, and the splitting point is generated at any data point between the maximum value and the minimum value in the current node data;

[0013] S03: Generate a hyperplane from the splitting point, dividing the current node data space into two subspaces. The data with values less than the splitting point are placed in the left child node of the current node, and the data with values greater than or equal to the splitting point are placed in the right child node of the current node;

[0014] S04: Recursively loop through steps S02 and S03 to continuously construct new child nodes until there is only one data in the child node or the child node has reached the maximum height defined by the binary tree;

[0015] S05: Loop through S01 to S04 to generate multiple binary trees.

[0016] Optionally, screening out outliers of well logging data based on the formed binary tree structure includes:

[0017] Based on the generated binary tree, calculate the path length of any target well logging data point in the binary tree;

[0018] Calculate the outlier score of the target well logging data point based on the path length;

[0019] Judge whether the target well logging data point belongs to an outlier data point based on the outlier score.

[0020] Optionally, calculating the path length of any target well logging data point in the binary tree is calculated by the following formula:

[0021] h(x) = a + c(n),

[0022]

[0023] where h(x) is the path length of the target data point x, a is the number of edges passed by the target well logging data point x from the root node to the leaf node of the tree, c(n) is a correction value, γ is the Euler constant, and n represents the number of data in the same leaf node as the data point x.

[0024] Optionally, calculating the outlier score of the target well logging data point based on the path length is calculated by the following formula:

[0025]

[0026] where s(x,n) represents the outlier score of the target well logging data point x, and E(h(x)) is the average value of the path lengths of the target well logging data point x on multiple binary trees.

[0027] Optionally, judging whether the target well logging data point belongs to an outlier data point based on the outlier score includes:

[0028] Determine the target well logging data points with outlier scores belonging to the set outlier value range as outliers.

[0029] In a second aspect, the present invention provides an automatic cleaning device for outlier logging data, comprising:

[0030] A data set production module, configured to form a multi-dimensional space logging data set from several target logging curves of a target well according to an actual sampling rate;

[0031] An abnormal data detection module, configured to use an unsupervised abnormal data detection algorithm to randomly split the target logging data in the multi-dimensional space logging data set by using a binary tree, and screen out the outliers of the logging data based on the formed binary tree structure;

[0032] An abnormal data elimination module, configured to eliminate the outliers from the target logging curves to obtain the target logging curves after data cleaning.

[0033] In a third aspect, the present invention provides an electronic device, the electronic device comprising:

[0034] At least one processor; and,

[0035] A memory communicatively connected to the at least one processor; wherein,

[0036] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the automatic cleaning method for outlier logging data according to any one of the first aspect.

[0037] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium, which stores computer instructions for causing a computer to execute the automatic cleaning method for outlier logging data according to any one of the first aspect.

[0038] The beneficial effects of the present invention are as follows:

[0039] The method of the present invention first forms a multi-dimensional space logging data set from several target logging curves of a target well according to an actual sampling rate, then uses an unsupervised abnormal data detection algorithm to randomly split the target logging data in the multi-dimensional space logging data set by using a binary tree, then screens out the outliers of the logging data based on the formed binary tree structure, and then eliminates the outliers from the target logging curves to obtain the target logging curves after data cleaning. This method can greatly improve the efficiency by carrying out relevant data processing based on automation and intelligent technologies, reduce the editing errors caused by human intervention, and improve the quality of logging data processing. The method of the present invention is applicable to the preprocessing of large-scale logging data and the rapid production of intelligent learning sample label data sets, and can lay a good data foundation for later reservoir prediction and reservoir description.

[0040] The system of the present invention has other characteristics and advantages, which will be apparent from the accompanying drawings incorporated herein and the subsequent detailed description, or will be described in detail in the accompanying drawings incorporated herein and the subsequent detailed description. These accompanying drawings and detailed description are used together to explain the specific principles of the present invention. Description of the Drawings

[0041] By describing the exemplary embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present invention will become more apparent. In the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.

[0042] Figure 1 A step diagram showing a method for automatically cleaning outliers in logging data according to the present invention is shown.

[0043] Figure 2 A schematic diagram of original target logging data in a method for automatically cleaning outliers in logging data according to an embodiment of the present invention is shown.

[0044] Figure 3 A schematic diagram of target logging data after automatic cleaning in a method for automatically cleaning outliers in logging data according to an embodiment of the present invention is shown. Detailed Description of the Embodiments

[0045] The present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to fully convey the scope of the present invention to those skilled in the art.

[0046] Embodiment 1

[0047] As Figure 1 shown, this embodiment provides a method for automatically cleaning outliers in logging data, including:

[0048] S1: Composing a multi-dimensional space logging data set from several target logging curves of a target well according to the actual sampling rate;

[0049] Specifically, according to N target logging curves of the target well, an N-curve data set in a multi-dimensional space is composed according to the actual sampling rate.

[0050] S2: Using an unsupervised outlier data detection algorithm, randomly splitting the target logging data in the multi-dimensional space logging data set by using a binary tree;

[0051] Preferably, the unsupervised abnormal data detection algorithm includes the Isolation Forest algorithm. This step specifically includes:

[0052] S201: Randomly select several data from the multi-dimensional space logging data set as a sample subset, and put the sample subset into the root node of the binary tree;

[0053] S202: Randomly generate a splitting point in the current node data, and the splitting point is generated at any data point between the maximum value and the minimum value in the current node data;

[0054] S203: Generate a hyperplane from the splitting point to divide the current node data space into two sub-spaces, where the data with a value less than the splitting point is placed in the left child node of the current node, and the data with a value greater than or equal to the splitting point is placed in the right child node of the current node;

[0055] S204: Recursively loop steps S202 and S203 to continuously construct new child nodes until there is only one data in the child node or the child node has reached the maximum height defined by the binary tree;

[0056] S205: Loop S201 to S204 to generate multiple binary trees.

[0057] Specifically, in this step, a random hyperplane is selected to divide the data space. Each time it is divided, two sub-spaces are generated, and then a random hyperplane is selected to divide each sub-space, and so on, until there is only one data point in each sub-space or the binary tree reaches the specified height.

[0058] S3: Screen out the outliers of the logging data based on the formed binary tree structure;

[0059] This step specifically includes the following steps:

[0060] S301: Based on the generated binary tree, calculate the path length of any target logging data point in the binary tree;

[0061] Specifically, the path length of any target logging data point in the binary tree is calculated by the following formula:

[0062] h(x) = a + c(n),

[0063]

[0064] where h(x) is the path length of the target data point x, a is the number of edges passed by the target logging data point x from the root node to the leaf node of the tree, c(n) is a correction value, γ is the Euler constant, and n represents the number of data in the same leaf node as the data point x.

[0065] S302: Calculate the anomaly score of the target log data point based on the path length;

[0066] Specifically, calculate the anomaly score of the target log data point based on the path length through the following formula:

[0067]

[0068] Where s(x,n) represents the anomaly score of the target log data point x, and E(h(x)) is the average value of the path lengths of the target log data point x on multiple binary trees.

[0069] S303: Determine whether the target log data point belongs to an abnormal data point based on the anomaly score.

[0070] Specifically, it can be seen from the multiple binary trees formed that those normal values with high density are stopped only after being divided many times, but those abnormal points with low density are easily stopped in a subspace very early, and these points with low density are abnormal values. The target log data points whose anomaly scores belong to the set abnormal value range can be determined as abnormal values. For example, when the path length of the data point x is smaller, the anomaly score s(x,n) is closer to 1, and at this time, the probability that the target log data point is an abnormal value is greater, that is, the target log data points with the anomaly score s(x,n) close to 1 are screened out as abnormal value data points.

[0071] S4: Remove the abnormal values from the target log curve to obtain the target log curve after data cleaning.

[0072] Based on the above automatic cleaning method for abnormal values of log data of the present invention, it is a fast and practical preprocessing technology for removing abnormal data. Carrying out relevant data processing based on automation and intelligent technologies can greatly improve efficiency, reduce editing errors caused by human intervention, improve the quality of log data processing, can quickly and cleanly remove abnormal data points in the original log data, and does not lose the effective values of the log data.

[0073] Embodiment 2

[0074] This embodiment provides an automatic cleaning method for abnormal values of log data. The method of this embodiment first forms a data set in a multi-dimensional space according to N target log curves of the target well at the actual sampling rate, then selects a random hyperplane to divide the data space. Each time it is divided, two subspaces are generated, and then a random hyperplane is selected to divide each subspace again, and the cycle continues until there is only one data point in each subspace or the specified height is reached. Furthermore, it can be seen that those normal values with high density are stopped only after being divided many times, but those abnormal points with low density are easily stopped in a subspace very early, and these points with low density are abnormal values.

[0075] The specific steps of this embodiment are as follows:

[0076] (1) Unsupervised training stage:

[0077] ① Randomly select data points from the logging data as a sample subset and put them into the root node of the binary tree.

[0078] ② Randomly generate a splitting point p in the current node data, which is generated between the maximum and minimum values of the current node data.

[0079] ③ Generate a hyperplane with point p as the splitting point, and divide the current node data space into 2 subspaces: the data less than p is placed in the left child node of the current node, and the data greater than or equal to p is placed in the right child node of the current node.

[0080] ④ Recursively loop steps ② and ③, continuously construct new child nodes until there is only one data in the child node or the child node has reached the maximum height defined by the binary tree.

[0081] ⑤ Loop ① to ④ to generate T binary trees.

[0082] (2) Outlier inference and removal processing:

[0083] ① Calculate the path length of the target logging data point x:

[0084] If the number of edges passed by the target logging data point x from the root node to the leaf node of the tree is a, then the path length h(x) of the data point x is:

[0085] h(x) = a + c(n),

[0086]

[0087] where γ is the Euler constant and n represents the number of data in the same leaf node as the data point x.

[0088] ② Calculate the outlier score of the target logging data point x. The specific formula is as follows:

[0089]

[0090] where E(h(x)) is the average value of the path lengths of the data point x on T binary trees.

[0091] ③ Outlier removal:

[0092] From the results obtained according to ②, it can be seen that when the path length of the data point x is smaller, s(x,n) is closer to 1, and at this time, the probability that the logging data point is an outlier is greater. Therefore, the abnormal logging data with an abnormal score close to 1 is removed from the original target logging curve to obtain the target logging curve after automatic cleaning. The original target logging curve in this embodiment is as Figure 2 shown. In the figure, multiple obvious abnormal data points can be seen within the dotted box. The target logging curve after automatic cleaning is as Figure 3 shown, and it can be seen that Figure 3 the abnormal data points in it have all been removed, and the valid values of the logging data have not been lost.

[0093] Embodiment 3

[0094] This embodiment provides an automatic cleaning device for logging data outliers, including:

[0095] A data set production module, configured to form a multi-dimensional space logging data set from several target logging curves of a target well according to the actual sampling rate;

[0096] An abnormal data detection module, configured to use an unsupervised abnormal data detection algorithm, randomly split the target logging data in the multi-dimensional space logging data set by using a binary tree, and screen out the outliers of the logging data based on the formed binary tree structure;

[0097] An abnormal data removal module, configured to remove the outliers from the target logging curve to obtain the target logging curve after data cleaning.

[0098] In this embodiment, the unsupervised abnormal data detection algorithm adopted by the abnormal data detection module includes the isolation forest algorithm.

[0099] The abnormal data detection module is specifically configured to perform the following processing steps:

[0100] S01: Randomly select several data from the multi-dimensional space logging data set as a sample subset, and put the sample subset into the root node of the binary tree;

[0101] S02: Randomly generate a splitting point in the current node data, and the splitting point is generated at any data point between the maximum value and the minimum value of the current node data;

[0102] S03: Generate a hyperplane from the splitting point to divide the current node data space into two sub-spaces, where the data with a value less than the splitting point is placed in the left child node of the current node, and the data with a value greater than or equal to the splitting point is placed in the right child node of the current node;

[0103] S04: Recursively loop through steps S02 and S03 to continuously construct new child nodes until there is only one data in the child nodes or the child nodes have reached the maximum height defined by the binary tree;

[0104] S05: Loop through S01 to S04 to generate multiple binary trees.

[0105] The method for the abnormal data detection module to screen out the abnormal values of well logging data based on the formed binary tree structure includes:

[0106] Based on the generated binary tree, calculate the path length of any target well logging data point in the binary tree;

[0107] Calculate the abnormal score of the target well logging data point based on the path length;

[0108] Based on the abnormal score, determine whether the target well logging data point belongs to an abnormal data point.

[0109] In this embodiment, the calculation of the path length of any target well logging data point in the binary tree is calculated by the following formula:

[0110] h(x) = a + c(n),

[0111]

[0112] where h(x) is the path length of the target data point x, a is the number of edges passed by the target well logging data point x from the root node to the leaf node of the tree, c(n) is a correction value, γ is the Euler constant, and n represents the number of data in the same leaf node as the data point x.

[0113] The calculation of the abnormal score of the target well logging data point based on the path length is calculated by the following formula:

[0114]

[0115] where s(x,n) represents the abnormal score of the target well logging data point x, and E(h(x)) is the average value of the path lengths of the target well logging data point x on multiple binary trees.

[0116] In this embodiment, the method for the abnormal data elimination module to determine whether the target well logging data point belongs to an abnormal data point based on the abnormal score includes:

[0117] Determine the target well logging data points with abnormal scores belonging to the set abnormal value range as abnormal values.

[0118] Embodiment 4

[0119] This embodiment provides an electronic device, and the electronic device includes:

[0120] At least one processor; and,

[0121] A memory communicatively connected to the at least one processor; wherein,

[0122] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the automatic cleaning method for outliers in logging data according to any one of the foregoing embodiments.

[0123] An electronic device according to an embodiment of the present disclosure includes a memory and a processor, and the memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0124] The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In an embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory.

[0125] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain good user experience effects, this embodiment may also include well-known structures such as communication buses, interfaces, etc., and these well-known structures should also be included in the protection scope of the present disclosure.

[0126] For the detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.

[0127] Embodiment 5

[0128] This embodiment provides a non-transitory computer-readable storage medium, which stores computer instructions for causing a computer to execute the automatic cleaning method for outliers in logging data according to any one of the foregoing embodiments.

[0129] A computer-readable storage medium according to an embodiment of the present disclosure stores non-transitory computer-readable instructions thereon. When the non-transitory computer-readable instructions are run by a processor, all or part of the steps of the methods of the foregoing embodiments of the present disclosure are executed.

[0130] The above computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or removable hard disk), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).

[0131] The embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.

Claims

1. An automatic cleaning method for outlier logging data, characterized in that, it includes: Form a multi-dimensional space logging data set by several target logging curves of a target well according to the actual sampling rate; Use an unsupervised outlier data detection algorithm and randomly split the target logging data in the multi-dimensional space logging data set by using a binary tree; Screen out the outliers of the logging data based on the formed binary tree structure; Remove the outliers from the target logging curves to obtain the target logging curves after data cleaning.

2. The automatic cleaning method for outlier logging data according to claim 1, characterized in that, the unsupervised outlier data detection algorithm includes the Isolation Forest algorithm.

3. The automatic cleaning method for outlier logging data according to claim 2, characterized in that, the random splitting of the target logging data in the multi-dimensional space logging data set by using a binary tree includes: S01: Randomly select several data from the multi-dimensional space logging data set as a sample subset, and put the sample subset into the root node of the binary tree; S02: Randomly generate a splitting point in the current node data, and the splitting point is generated at any data point between the maximum value and the minimum value of the current node data; S03: Generate a hyperplane from the splitting point to divide the current node data space into two sub-spaces, where the data with a value less than the splitting point is placed in the left child node of the current node, and the data with a value greater than or equal to the splitting point is placed in the right child node of the current node; S04: Recursively loop steps S02 and S03 to continuously construct new child nodes until there is only one data in the child node or the child node has reached the maximum height defined by the binary tree; S05: Loop S01 to S04 to generate multiple binary trees.

4. The automatic cleaning method for outlier logging data according to claim 3, characterized in that, the screening out of the outliers of the logging data based on the formed binary tree structure includes: Based on the generated binary tree, calculate the path length of any target logging data point in the binary tree; Calculate the outlier score of the target logging data point based on the path length; Judge whether the target logging data point belongs to an outlier data point based on the outlier score.

5. The automatic cleaning method for outlier logging data according to claim 4, characterized in that, the calculation of the path length of any target logging data point in the binary tree is calculated by the following formula: h(x) = a + c(n), where h(x) is the path length of the target data point x, a is the number of edges passed by the target logging data point x from the root node to the leaf node of the tree, c(n) is a correction value, γ is the Euler constant, and n represents the number of data in the same leaf node as the data point x.

6. The automatic cleaning method for outlier logging data according to claim 5, characterized in that, the calculation of the outlier score of the target logging data point based on the path length is calculated by the following formula: where s(x,n) represents the outlier score of the target logging data point x, and E(h(x)) is the average value of the path lengths of the target logging data point x on multiple binary trees.

7. The automatic cleaning method for outliers in logging data according to claim 6, wherein, judging whether the target logging data point belongs to an outlier data point based on the outlier score includes: Determining the target logging data points with outlier scores within a set outlier value range as outliers.

8. An automatic cleaning device for outliers in logging data, wherein, it includes: A data set production module, configured to form a multi-dimensional space logging data set from several target logging curves of a target well according to an actual sampling rate; An outlier data detection module, configured to use an unsupervised outlier data detection algorithm to randomly split the target logging data in the multi-dimensional space logging data set by using a binary tree, and screen out the outliers in the logging data based on the formed binary tree structure; An outlier data elimination module, configured to eliminate the outliers from the target logging curves to obtain the target logging curves after data cleaning.

9. An electronic device, wherein, the electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the automatic cleaning method for outliers in logging data according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium, wherein, the non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the automatic cleaning method for outliers in logging data according to any one of claims 1-7.