Logging data processing method and device, electronic equipment and storage medium
By segmenting, standardizing and screening the quality factor of well recording data, the accuracy of well recording data is solved, and the data quality is improved to support intelligent applications.
Patent Information
- Application Number
- CN202410006670.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
The accuracy of well recording data is affected by the drilling working characteristics and irregular operation, which has led to the failure of the advantages of well recording technology to fully utilize and limit its widespread application.
By obtaining the pending well recording data set, sorting it in time series and segmenting it into multiple data subsets, computing data quality factors, filtering and optimizing data subsets to improve data quality, including data standardization and interpolation processing.
Improve the accuracy and consistency of well recording data and support subsequent intelligent applications.
Smart Images

Figure CN120256414A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of oil engineering data processing, and particularly to a method, device, electronic device and storage medium for well logging data processing. Background Art
[0002] Well logging technology mainly records various data of drilling during the oil extraction process. It is not only the basis for the smooth progress of various projects but also the guarantee for the smooth progress of oil extraction, playing a very important role in oil extraction. In actual situations, the well logging technology can monitor the drilling process in all aspects.
[0003] Although well logging data can play a great role, in actual situations, due to the characteristics of drilling work itself and the non-standard operation of staff, there are certain potential hazards that affect the accuracy of well logging data. Due to the factors of well logging data quality, the advantages of well logging technology have not been fully exerted, which has greatly restricted the development of well logging technology and cannot be widely applied. Therefore, improving the quality of well logging data has a profound impact on the survival and development space of well logging. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method, device, electronic device and storage medium for well logging data processing, which can improve the quality of well logging data for subsequent intelligent applications.
[0005] According to one aspect of the present disclosure, a method for well logging data processing is provided, including:
[0006] Obtaining a well logging data set to be processed, where the well logging data set to be processed is arranged in a time series;
[0007] Dividing the well logging data set to be processed according to a preset time gap value to obtain a plurality of data subsets;
[0008] Calculating a data quality factor for each data subset, where the data quality factor is the ratio of the non-empty data fields in the data subset to the total number of all fields in the data subset;
[0009] Screening data subsets according to the data quality factor, and obtaining target well logging data according to the screened data subsets and the time series.
[0010] In one implementation, the dividing the well logging data set to be processed according to a preset time gap value to obtain a plurality of data subsets includes:
[0011] Dividing two adjacent data in the well logging data set to be processed with a time interval greater than the preset time gap value according to the preset time gap value to obtain a plurality of data subsets.
[0012] In one embodiment, the method further includes:
[0013] Based on the original logging data of multiple transmission channels obtained, an original logging data set is obtained. The original logging data set includes multiple field tags, and each field tag has a corresponding field data set including a field data value, a field unit, and a field timestamp;
[0014] Data standardization is performed on the original logging data set to obtain the logging data set to be processed.
[0015] In one embodiment, the obtaining the original logging data set based on the original logging data of multiple transmission channels includes:
[0016] Based on the transmission channel information of the original logging data through a predefined data dictionary, a field data set including a field data value, a field unit, and a field timestamp corresponding to each field tag in the original logging data set is obtained. The data dictionary includes multiple field tags and is provided with a first mapping relationship between each field tag and the transmission channel.
[0017] In one embodiment, the performing data standardization on the original logging data set includes:
[0018] Performing time format conversion on the field timestamp through a preset time format;
[0019] Performing numerical magnitude conversion on the field data value in the original logging data set through the preset field unit corresponding to the field tag.
[0020] In one embodiment, the performing data standardization on the original logging data set includes:
[0021] Performing numerical standardization processing on the field data value corresponding to each field tag in the original logging data set.
[0022] In one embodiment, the performing numerical standardization processing on the field data value corresponding to each field tag in the original logging data set includes:
[0023] For each field tag, according to a preset error value, the field data with the field data value being the preset error value in the original logging data set is determined as incorrect field data, or, for each field tag, according to a preset field threshold range, the field data with the field data value exceeding the preset field threshold range in the original logging data set is determined as incorrect field data, and the incorrect field data is replaced with a null value identifier.
[0024] In one embodiment, the numerical standardization process for the field data values corresponding to each field label in the original logging dataset includes:
[0025] For the field dataset corresponding to each field label, divide the field dataset into multiple first subsequences according to a first preset time window;
[0026] For the field data values within each first subsequence, determine the numerical distribution of the field data values in the first subsequence. When the numerical distribution satisfies the normal distribution, determine the field data values that exceed three standard deviations from the mean as abnormal field data, or when the numerical distribution does not satisfy the normal distribution, determine the field data values that are less than the 5th percentile or greater than the 95th percentile as abnormal field data, and replace the abnormal field data with null value identifiers.
[0027] In one embodiment, the numerical standardization process for the field data values corresponding to each field label in the original logging dataset includes:
[0028] For the field dataset corresponding to each field label, divide the field dataset into multiple second subsequences according to a second preset time window;
[0029] Determine the median value of the data in each second subsequence according to the field data values in the second subsequence;
[0030] Determine the abnormal second subsequences in the field dataset according to the median values of the data in each second subsequence;
[0031] Replace the field data values in the abnormal second subsequences with null value identifiers.
[0032] In one embodiment, the screening of the data subset according to the data quality factor includes:
[0033] Delete the data subsets whose ratios are less than a preset ratio according to the data quality factor.
[0034] In one embodiment, the obtaining of the target logging data according to the screened data subset and the time series includes:
[0035] Perform linear interpolation processing on the null value identifiers in the screened data subset to obtain a target data subset;
[0036] Resample the target data subset according to the target data subset and the time series to obtain the target logging data.
[0037] According to another aspect of the present disclosure, there is provided a logging data processing device, including:
[0038] A data acquisition module for acquiring a well logging data set to be processed, where the well logging data set to be processed is arranged in a time series;
[0039] A data segmentation module for segmenting the well logging data set to be processed according to a preset time gap value to obtain a plurality of data subsets;
[0040] A quality analysis module for calculating a data quality factor for each data subset, where the data quality factor is the ratio of the non-empty data fields in the data subset to the total number of all fields in the data subset;
[0041] A data optimization module for screening data subsets according to the data quality factor, and obtaining target well logging data based on the screened data subsets and the time series.
[0042] In one implementation, the device further includes:
[0043] A data standardization module for obtaining an original well logging data set based on the original well logging data of multiple transmission channels, and performing data standardization on the original well logging data set to obtain the well logging data set to be processed, where the original well logging data set includes multiple field tags, and each field tag has a corresponding field data set including a field data value, a field unit, and a field timestamp.
[0044] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0045] At least one processor; and
[0046] At least one memory storing a computer program,
[0047] The processor calls the computer program to cause the processor to execute the well logging data processing method according to one aspect of the present disclosure as described above.
[0048] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing a computer program, where the computer program is used to cause a computer to execute the well logging data processing method according to one aspect of the present disclosure as described above.
[0049] The above technical features can be combined in various suitable ways or replaced by equivalent technical features as long as the object of the present invention can be achieved.
[0050] One or more technical solutions provided in the embodiments of the present disclosure can segment a to-be-processed logging data set according to a preset time interval value to obtain multiple data subsets, then calculate the data quality factor of each data subset, screen the data subsets based on the data quality factor, and then obtain the final target logging data according to the screened data subsets and time series, which can improve the quality of logging data and facilitate subsequent intelligent applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features and advantages of the present disclosure are disclosed. In the drawings:
[0052] Figure 1 A flowchart showing a method for processing logging data according to an exemplary embodiment of the present disclosure is shown;
[0053] Figure 2 A flowchart showing another method for processing logging data according to an exemplary embodiment of the present disclosure is shown;
[0054] Figure 3 A schematic block diagram showing a device for processing logging data according to an exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0056] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0057] As used herein, the term "including" and its variants are open-ended, that is, "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0058] It should be noted that the modification of "one" and "multiple" mentioned in this disclosure is illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".
[0059] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0060] The solutions of the embodiments of this disclosure will be described below with reference to the accompanying drawings.
[0061] The embodiments of this disclosure provide a method for processing logging data. Refer to Figure 1 , Figure 1 which shows a schematic flowchart of a method for processing logging data according to an exemplary embodiment of this disclosure. The method includes:
[0062] S101, obtaining a logging data set to be processed, where the logging data set to be processed is arranged in a time series;
[0063] S102, dividing the logging data set to be processed according to a preset time gap value to obtain a plurality of data subsets;
[0064] S103, calculating a data quality factor for each data subset, where the data quality factor is the ratio of the non-empty data fields in the data subset to the total number of all fields in the data subset;
[0065] S104, screening the data subsets according to the data quality factor, and obtaining target logging data according to the screened data subsets and the time series.
[0066] In the embodiments of this disclosure, the logging data set to be processed is divided according to a preset time gap value to obtain a plurality of data subsets, then the data quality factor of each data subset is calculated, the data subsets are screened based on the data quality factor, and finally the target logging data is obtained according to the screened data subsets and the time series, which can improve the quality of logging data and facilitate subsequent intelligent applications.
[0067] Drilling operations are usually discontinuous in time, so the logging data collected by sensors also has a certain time interval.
[0068] By setting a preset time gap value, the logging data set to be processed can be divided to obtain a plurality of data subsets. To prevent excessive data interpolation from being required subsequently, which may affect data accuracy, the preset time gap value cannot be set too large. In continuous exploration, a feasible way is to set the preset time gap value to 60s to prevent excessive interpolation of long data sequences and thus reduce deviations.
[0069] In step S102, the to-be-processed logging dataset is segmented according to a preset time interval value. Specifically, two adjacent data in the to-be-processed logging dataset with a time interval greater than the preset time interval value can be segmented to obtain multiple data subsets. Exemplarily, for the to-be-processed logging dataset, when the time interval between two pieces of data is greater than 60s, these two pieces of data are segmented, and thus the to-be-processed logging dataset is divided into multiple data subsets.
[0070] In some embodiments of the present disclosure, screening the data subsets according to the data quality factor may include: deleting the data subsets with a ratio less than a preset ratio according to the data quality factor. Exemplarily, for each data subset, calculate the ratio of the non-empty data fields in the data subset to the total number of all fields. If the ratio is less than 80%, then discard the data subset. By setting the data quality factor to 80%, the data quality can be ensured.
[0071] In some embodiments of the present disclosure, as Figure 2 shown, Figure 2 FIG. shows a schematic flowchart of another logging data processing method according to an exemplary embodiment of the present disclosure. The method includes:
[0072] S201, based on the original logging data of multiple transmission channels obtained, obtain an original logging dataset. The original logging dataset includes multiple field tags, and each field tag has a corresponding field dataset. The field dataset includes a field data value, a field unit, and a field timestamp corresponding to the field tag;
[0073] S202, perform data standardization on the original logging dataset to obtain a to-be-processed logging dataset;
[0074] S203, segment the to-be-processed logging dataset according to a preset time interval value to obtain multiple data subsets;
[0075] S204, calculate the data quality factor of each data subset. The data quality factor is the ratio of the non-empty data fields in the data subset to the total number of all fields in the data subset;
[0076] S205, screen the data subsets according to the data quality factor, and obtain the target logging data according to the screened data subsets and the time series.
[0077] In the embodiments of the present disclosure, by obtaining the original logging data of multiple transmission channels to obtain an original logging dataset, performing data standardization on the original logging dataset, and combining the calculation of the data quality factors of the multiple segmented data subsets, the problems caused by different data sources having different properties or different orders of magnitude can be avoided, and the data quality can be optimized.
[0078] The original logging data at the wellsite, including various sensor monitoring data, is usually obtained through sensor detection. The sensor monitoring data of different field transmission channels can be fused into a single data stream and transmitted to a WITSML (Wellsite Information Transfer Standard Markup Language) server for further processing and analysis. WITSML is a platform- and language-independent data format for data transmission and exchange, which is built on XML technology. It contains metadata associated with timestamp tags (such as transmission channel tags, channel tags, mnemonics, units, and service company names, etc.) and sensor measurement values or data values.
[0079] The original logging data of multiple transmission channels detected at the wellsite can be further processed after being received by the receiving-end server of the logging data. The logging data processing method of the exemplary embodiments of the present disclosure can be deployed on the receiving-end server of the logging data. Through this method, the original logging data of multiple transmission channels obtained is processed to improve the quality of the logging data, and the target logging data is output. Furthermore, the target logging data is pushed to the database of the intelligent application system for subsequent further applications.
[0080] In some embodiments of the present disclosure, based on the original logging data of multiple transmission channels obtained, an original logging data set can be obtained, which may include: through a predefined data dictionary, based on the transmission channel information of the original logging data, a field data set corresponding to each field tag in the original logging data set is obtained, including field data values, field units, and field timestamps, where the data dictionary includes multiple field tags and a first mapping relationship between each field tag and the transmission channel is set.
[0081] The files transmitted in WITSML format may contain multiple redundant channels for representing the same field. For example, time, depth, and pressure are each a transmission channel, but different logging instruments may require different channels for the same field. For example, for some logging instruments, there is only one channel for the depth field, while for some logging instruments, there are multiple channels or channels for the depth field. Therefore, the exemplary embodiments of the present disclosure define a data dictionary, which includes multiple field tags and a mapping relationship between each field tag and the transmission channel is set to ensure that on the receiving-end server side of the logging data, based on the original logging data of multiple transmission channels obtained, that is, the actual values are obtained, and then the field data set corresponding to the custom field tags on the receiving-end server side is obtained. The field tags are used to indicate different sensor data types, such as pressure field tags, temperature field tags, and more specifically, formation pressure field tags, riser pressure field tags, drill pressure field tags, etc.
[0082] After obtaining the original logging dataset, in some embodiments of the present disclosure, data standardization is performed on the original logging dataset. Exemplarily, it may include: converting the time format of the field timestamp through a preset time format; converting the numerical magnitude of the field data values in the original logging dataset through the preset field unit corresponding to the field label.
[0083] Since the received original logging data is the monitoring data of various sensors, considering that when different sensors collect data, there are differences in the format of the acquisition time of the sensor detection data or the numerical unit of the sensor monitoring data. For example, time is usually stored as a string and can appear in multiple formats, such as yyyy - mm - dd hh:mm:ss or yyyy:mm:dd hh:mm:ss, while the numerical unit can be manually defined during the data acquisition process. Therefore, different units or unit labels are usually found in the collected data. For example, the unit or unit label of the hook load can be lbf (pound - force), klbf (kilo - pound - force). Therefore, by defining a preset time format, that is, a unified time format, the time format of the field timestamp is converted to convert the time of various sensor monitoring data into a unified time format, which is convenient for subsequent further analysis or processing. If the time cannot be accurate to the second, this piece of data is discarded. By defining the preset field unit corresponding to each field label and performing conversion, the numerical magnitude of the fields is unified.
[0084] In some embodiments of the present disclosure, performing data standardization on the original logging dataset may further include: performing numerical standardization processing on the field data values corresponding to each field label in the original logging dataset.
[0085] Specifically, performing numerical standardization processing on the field data values corresponding to each field label in the original logging dataset may include: for each field label, determining the field data with the field data value being the preset error value in the original logging dataset as the error field data according to the preset error value, or, for each field label, determining the field data with the field data value exceeding the preset field threshold range in the original logging dataset as the error field data according to the preset field threshold range, and replacing the error field data with a null value identifier.
[0086] Downhole instruments assign predefined error values in the case of failure to read and store measurement values during the sensor data acquisition process. Different downhole instruments have different defined error values. For example, typical error values can be "-999", "-999900", "-9.99". Exemplarily, when the read and storage of measurement data fails, the data is directly assigned "-999". During subsequent analysis, it can be known that this measurement data is an error value. Therefore, for each field label, according to the preset error value, the field data with the field data value of the preset error value in the original mud logging dataset is determined as the error field data, and the error field data is replaced with a null value identifier.
[0087] In addition, there may be unstable measurement values that are physically meaningless, perhaps because they are negative or too large to have a true value, such as a pressure value cannot be less than zero. In this case, define the preset field threshold range corresponding to each field label, such as the minimum value and the maximum value, to identify obviously untrue measurement values. Once confirmed as error field data, the error field data is replaced with a control identifier such as a Null identifier.
[0088] In some embodiments of the present disclosure, the numerical normalization process for the field data values corresponding to each field label in the original mud logging dataset may further include: for the field dataset corresponding to each field label, dividing the field dataset into multiple first subsequences according to the first preset time window; for the field data values within each first subsequence, determining the numerical distribution of the field data values in the first subsequence. In the case where the numerical distribution satisfies the normal distribution, the field data values that exceed three standard deviations from the average value are determined as abnormal field data, or in the case where the numerical distribution does not satisfy the normal distribution, the field data values that are less than the 5th percentile or greater than the 95th percentile are determined as abnormal field data, and the abnormal field data is replaced with a null value identifier.
[0089] Considering that the time span of drilling operations is generally very long, the obtained original logging dataset may be data lasting for several months. Therefore, in the exemplary embodiments of the present disclosure, the field dataset is divided into multiple first subsequences according to a first preset time window. The first preset time window, such as a 24-hour time window, and the length of the first preset time window can be adjusted according to data characteristics or determined according to the change rate of the time series trend, and can be between one day and one month. By evaluating the statistical distribution of the field data values on the first preset time window, abnormal field data is then determined. If the numerical distribution satisfies the normal distribution, the field data values that exceed three standard deviations from the average value are determined as abnormal field data, and values that exceed 3 standard deviations from the average value can be eliminated by this method. Or when the numerical distribution does not satisfy the normal distribution, that is, when there is a significant deviation from the normal distribution, the field data values less than the 5th percentile or greater than the 95th percentile are determined as abnormal field data. Once abnormal field data is detected, a control identifier such as a Null identifier is used to replace the abnormal field data.
[0090] In some embodiments of the present disclosure, performing numerical standardization processing on the field data values corresponding to each field label in the original logging dataset may further include: dividing the field dataset into multiple second subsequences according to a second preset time window for each field dataset corresponding to a field label; determining the data median of each second subsequence according to the field data values in the second subsequence; determining the abnormal second subsequences in the field dataset according to the data medians of the respective second subsequences; and replacing the field data values in the abnormal second subsequences with null identifiers.
[0091] Here, the second preset time window may be different from the length of the first preset time window. For example, a 6-hour time window and a 1-hour step size are used to divide the field dataset. Of course, they may also be the same to improve the efficiency of subsequent analysis.
[0092] According to the field data values in the second subsequences, determine the data median of each second subsequence. The data median may be the median of normal data. After determining the data medians of each second subsequence, the data median distribution of the entire field dataset can be statistically analyzed, and then combined with the data median distribution of the entire field dataset to determine the abnormal second subsequences in the field dataset. For example, if the data median of a certain second subsequence is significantly abnormal in the data median distribution of the entire field dataset, then the second subsequence is determined as an abnormal second subsequence, and all field data values in the abnormal second subsequence are replaced with a control identifier such as a Null identifier.
[0093] After the above numerical standardization process, there will be Null identifiers for the field data values corresponding to each field tag, that is, there will be Null identifiers in the to-be-processed logging dataset obtained. Therefore, further, in some embodiments of the present disclosure, for the to-be-processed logging dataset obtained after data standardization processing, when splitting the to-be-processed logging dataset according to a preset time gap value to obtain multiple data subsets, calculating the data quality factor of each data subset, and deleting the data subsets with a ratio less than 80% according to the ratio, the target logging data can be obtained based on the filtered data subsets and the time series, including: performing linear interpolation processing on the null value identifiers in the filtered data subsets to obtain target data subsets; resampling the target data subsets according to the target data subsets and the time series to obtain target logging data.
[0094] Linear interpolation is mainly for the incorrect field data, abnormal field data, and field data in the abnormal second subsequence replaced by null value identifiers during the previous data standardization process. By combining linear interpolation with the setting of a preset time gap value, it is possible to recover as many values in each data subset as possible without introducing biases caused by inaccurate data interpolation. Then, combined with the time series, the time series is resampled to a unified sampling frequency. The resampling method can use interpolation techniques such as linear interpolation to estimate the missing numerical values between the original time series by linear interpolation. Finally, the target data subsets after performing interpolation processing and resampling are concatenated to obtain the target logging data.
[0095] The embodiments of the present disclosure also provide a logging data processing device, as Figure 3 shown, Figure 3 The schematic block diagram of the logging data processing device according to the exemplary embodiments of the present disclosure is shown. The device includes:
[0096] A data acquisition module 301, configured to acquire a to-be-processed logging dataset, and the to-be-processed logging dataset is arranged in a time series;
[0097] A data splitting module 302, configured to split the to-be-processed logging dataset according to a preset time gap value to obtain multiple data subsets;
[0098] A quality analysis module 303, configured to calculate the data quality factor of each data subset, and the data quality factor is the ratio of the non-null data fields in the data subset to the total number of all fields in the data subset;
[0099] A data optimization module 304, configured to filter data subsets according to the data quality factor, and obtain target logging data based on the filtered data subsets and the time series.
[0100] In the embodiments of the present disclosure, the to-be-processed logging dataset is segmented according to a preset time gap value to obtain multiple data subsets, then the data quality factor of each data subset is calculated, the data subsets are screened based on the data quality factor, and finally the target logging data is obtained according to the screened data subsets and the time series, which can improve the quality of logging data and facilitate subsequent intelligent applications.
[0101] In some embodiments of the present disclosure, the data segmentation module 302 is further configured to segment two adjacent data in the to-be-processed logging dataset with a time interval greater than the preset time gap value according to the preset time gap value to obtain multiple data subsets.
[0102] In some embodiments of the present disclosure, the logging data processing device further includes: a data standardization module, configured to obtain an original logging dataset based on the original logging data of multiple transmission channels, and perform data standardization on the original logging dataset to obtain the to-be-processed logging dataset, where the original logging dataset includes multiple field tags, and each field tag has a corresponding field dataset including a field data value, a field unit, and a field timestamp.
[0103] In some embodiments of the present disclosure, the data standardization module obtains, through a predefined data dictionary, a field dataset corresponding to each field tag in the original logging dataset, including a field data value, a field unit, and a field timestamp, based on the transmission channel information of the original logging data, where the data dictionary includes multiple field tags and is provided with a first mapping relationship between each field tag and the transmission channel.
[0104] In some embodiments of the present disclosure, the data standardization module performs time format conversion on the field timestamp through a preset time format; and performs numerical magnitude conversion on the field data value in the original logging dataset through the preset field unit corresponding to the field tag.
[0105] In some embodiments of the present disclosure, the data standardization module is configured to perform numerical standardization processing on the field data values corresponding to each field tag in the original logging dataset.
[0106] In some embodiments of the present disclosure, when the data standardization module performs numerical standardization processing on the field data values corresponding to each field tag in the original logging dataset, for each field tag, according to a preset error value, the field data value in the original logging dataset that is the preset error value is determined as incorrect field data, or, for each field tag, according to a preset field threshold range, the field data value in the original logging dataset that exceeds the preset field threshold range is determined as incorrect field data, and the incorrect field data is replaced with a null value identifier.
[0107] In some embodiments of the present disclosure, when the data normalization module performs numerical normalization processing on the field data values corresponding to each field label in the original logging dataset, for the field dataset corresponding to each field label, the field dataset is divided into multiple first subsequences according to a first preset time window; for the field data values within each first subsequence, the numerical distribution of the field data values in the first subsequence is determined. When the numerical distribution satisfies the normal distribution, the field data values that exceed three standard deviations from the average value are determined as abnormal field data, or when the numerical distribution does not satisfy the normal distribution, the field data values that are less than the 5th percentile or greater than the 95th percentile are determined as abnormal field data, and the abnormal field data is replaced with a null value identifier.
[0108] In some embodiments of the present disclosure, when the data normalization module performs numerical normalization processing on the field data values corresponding to each field label in the original logging dataset, for the field dataset corresponding to each field label, the field dataset is divided into multiple second subsequences according to a second preset time window; according to the field data values in the second subsequences, the median value of the data in each second subsequence is determined; according to the median values of the data in each second subsequence, the abnormal second subsequences in the field dataset are determined; and the field data values in the abnormal second subsequences are replaced with a null value identifier.
[0109] In some embodiments of the present disclosure, the data optimization module 304 is specifically configured to delete the data subsets with a ratio less than a preset ratio according to the data quality factor.
[0110] In some embodiments of the present disclosure, the data optimization module 304 is further configured to perform linear interpolation processing on the null value identifiers in the filtered data subsets to obtain a target data subset; and resample the target data subset according to the target data subset and the time series to obtain the target logging data.
[0111] The relevant content of the logging data processing device provided by the embodiments of the present disclosure corresponds to the above-mentioned logging data processing method. For the matters not covered, reference may be specifically made to the relevant description of the foregoing logging data processing method, which will not be elaborated herein.
[0112] An exemplary embodiment of the present disclosure further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program that can be executed by the at least one processor, and when the computer program is executed by the at least one processor, it is configured to cause the electronic device to execute the method according to the embodiments of the present disclosure.
[0113] The exemplary embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is configured to cause the computer to execute the method according to the embodiments of the present disclosure.
[0114] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not depart from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the different dependent claims and the features described herein can be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a separate embodiment can be used in other described embodiments.
Claims
1. A method for processing mud logging data, characterized in that Including: Obtain a well logging data set to be processed, where the well logging data set to be processed is arranged in a time series; Divide the well logging data set to be processed according to a preset time gap value to obtain a plurality of data subsets; Calculate the data quality factor of each data subset, where the data quality factor is the ratio of the non-empty data fields to the total number of all fields in the data subset; According to the data quality factor, filter the data subsets, and obtain the target well logging data according to the filtered data subsets and the time series.
2. The method according to claim 1, characterized in that The dividing the well logging data set to be processed according to a preset time gap value to obtain a plurality of data subsets includes: According to the preset time gap value, divide the two adjacent data in the well logging data set to be processed with a time interval greater than the preset time gap value to obtain a plurality of data subsets.
3. The method according to claim 1, wherein The method further includes: Based on the original well logging data of multiple transmission channels, obtain an original well logging data set, where the original well logging data set includes multiple field tags, and each field tag has a corresponding field data set including a field data value, a field unit, and a field timestamp; Perform data standardization on the original well logging data set to obtain the well logging data set to be processed.
4. The method according to claim 3, wherein The obtaining the original well logging data set based on the original well logging data of multiple transmission channels includes: Through a predefined data dictionary, based on the transmission channel information of the original well logging data, obtain a field data set corresponding to each field tag in the original well logging data set, where the data dictionary includes multiple field tags and is provided with a first mapping relationship between each field tag and the transmission channel.
5. The method according to claim 4, characterized in that The performing data standardization on the original well logging data set includes: Perform time format conversion on the field timestamp through a preset time format; Perform numerical magnitude conversion on the field data value in the original well logging data set through the preset field unit corresponding to the field tag.
6. The method according to claim 3, wherein The performing data standardization on the original well logging data set includes: Perform numerical standardization processing on the field data values corresponding to each field tag in the original well logging data set.
7. The method according to claim 6, characterized in that, The performing numerical standardization processing on the field data values corresponding to each field tag in the original well logging data set includes: For each field tag, according to a preset error value, determine the field data with the field data value of the preset error value in the original well logging data set as error field data, or, for each field tag, according to a preset field threshold range, determine the field data with the field data value exceeding the preset field threshold range in the original well logging data set as error field data, and replace the error field data with a null value identifier.
8. The method according to claim 6, wherein The performing numerical standardization processing on the field data values corresponding to each field tag in the original well logging data set includes: For the field data set corresponding to each field tag, divide the field data set into multiple first subsequences according to a first preset time window; For each field data value within each first subsequence, determine the numerical distribution of the field data values in the first subsequence. When the numerical distribution satisfies the normal distribution, determine the field data values that exceed three standard deviations from the mean as abnormal field data. Or when the numerical distribution does not satisfy the normal distribution, determine the field data values that are less than the 5th percentile or greater than the 95th percentile as abnormal field data, and replace the abnormal field data with null value identifiers.
9. The method according to claim 6, wherein The numerical standardization process for the field data values corresponding to each field label in the original logging dataset includes: For each field dataset corresponding to a field label, divide the field dataset into multiple second subsequences according to a second preset time window; Determine the median value of the data in each second subsequence based on the field data values in the second subsequence; Determine the abnormal second subsequences in the field dataset based on the median values of the data in each second subsequence; Replace the field data values in the abnormal second subsequences with null value identifiers.
10. The method according to claim 1, characterized in that, The screening of the data subsets according to the data quality factor includes: Delete the data subsets with a ratio less than a preset ratio according to the data quality factor.
11. The method according to claim 1, wherein The obtaining of the target logging data according to the screened data subsets and the time series includes: Perform linear interpolation on the null value identifiers in the screened data subsets to obtain a target data subset; Resample the target data subset according to the target data subset and the time series to obtain the target logging data.
12. A mud logging data processing device, characterized in that, It includes: A data acquisition module, configured to acquire a logging dataset to be processed, where the logging dataset to be processed is arranged in a time series; A data segmentation module, configured to segment the logging dataset to be processed according to a preset time gap value to obtain multiple data subsets; A quality analysis module, configured to calculate the data quality factor of each data subset, where the data quality factor is the ratio of the non-null data fields to the total number of all fields in the data subset; A data optimization module, configured to screen the data subsets according to the data quality factor, and obtain the target logging data according to the screened data subsets and the time series.
13. The device according to claim 12, characterized in that, The device further includes: A data standardization module, configured to obtain an original logging dataset based on the original logging data of multiple transmission channels, and perform data standardization on the original logging dataset to obtain the logging dataset to be processed, where the original logging dataset includes multiple field labels, and each field label has a corresponding field dataset including field data values, field units, and field timestamps.
14. An electronic device, characterized in that, It includes: At least one processor; And At least one memory storing a computer program, The processor calls the computer program to cause the processor to execute the method according to any one of claims 1-11.
15. A non-transitory computer-readable storage medium storing a computer program, characterized in that, The computer program is used to cause the computer to execute the method according to any one of claims 1-11.