Comparative predictive coding-based logging data representation method and system, terminal and medium
By comparing the prediction coding network, the logging data is expressed and interpreted, and the problems of inefficient manual logging interpretation are solved, and the efficient logging interpretation effect is achieved under small sample conditions.
Patent Information
- Application Number
- CN202311617137.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, manual logging interpretation has low efficiency, large results differences, and high error rate. Especially under small sample conditions, the classification accuracy of deep learning methods is reduced, making it difficult to meet the requirements of logging interpretation.
The well log data representation method based on contrast prediction encoding is adopted, and the well log data is preprocessed and standardized, and the comparison prediction encoding network structure is built, and the model is trained using noise comparison estimation loss function to achieve effective representation of the well log data.
Under the conditions of small sample logging data, the accuracy and efficiency of logging interpretation are improved, the errors caused by manual judgment are reduced, the work efficiency of staff is improved, and the cost loss of the production process is reduced.
Smart Images

Figure CN120067525A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of oil and gas exploration and development, and particularly to a method, system, terminal and medium for representing logging data based on contrastive predictive coding. Background Art
[0002] With the rapid development of society and economy, the consumption of oil by various countries is increasing day by day. However, as a non-renewable resource, the reserves of oil will only become less and less, and even dry up. Compared with developing new oil fields to alleviate the problem of tight oil resources, it is more cost-effective and feasible to deeply develop the remaining oil and gas resources in existing developed oil fields and promote the high and stable production of existing oil fields. For a long time, logging interpretation tasks have mainly adopted manual processing methods. Specifically, logging interpreters mainly make subjective judgments based on the morphological transformation characteristics of logging curves and their own experience. Manual logging interpretation not only wastes manpower and material resources and has low efficiency, but also the interpretation results are affected by subjective factors of people and are limited by the technical level of logging interpreters. Especially for interpreters who are not very familiar with the geological conditions and lack experience, the division results vary greatly and the error rate is relatively high.
[0003] In the actual oil and gas exploration process, at the block determination stage, there is only the data of one labeled oil testing well and the data of 10 surrounding testing wells with expert interpretation conclusions. Under a large number of logging data samples, many deep learning methods have achieved good results in the geological layer division task. However, as the training data decreases, the classification accuracy of various deep learning methods also decreases. A small amount of logging data in a new block cannot support a deep convolutional neural network. The accuracy of machine learning methods under few-sample conditions cannot meet the requirements of logging interpretation, and it is difficult for feature engineering to contain all the information in logging data. Therefore, it is necessary to design a method for representing logging data under few-sample conditions, which can obtain better logging interpretation results under few-sample logging data conditions. Summary of the Invention
[0004] In order to overcome the defects existing in the above-mentioned prior art, the purpose of the present invention is to provide a method, system, terminal and medium for representing logging data based on contrastive predictive coding, so as to solve the technical problems of low manpower and material resources, low efficiency, large differences in division results and high error rate in manual logging interpretation in the prior art.
[0005] The present invention is realized through the following technical solutions:
[0006] A method for representing logging data based on contrastive predictive coding includes the following steps:
[0007] Step 1, obtain the original logging data, preprocess the original logging data, and perform standardization processing on the preprocessed original logging data to obtain several columns of logging data curves;
[0008] Step 2: Combine several columns of logging data curves to form a two-dimensional matrix, and obtain a two-dimensional representation of the logging data based on the two-dimensional matrix.
[0009] Step 3: Use the obtained two-dimensional representation of the logging data to divide the dataset into a training set and a test set for the learning environment. The training set is used to train the contrast prediction coding network model, and the test set is used to verify and evaluate the model performance.
[0010] Step 4: Build a contrast prediction coding network structure based on the two-dimensional representation of the logging data, rely on the noise contrast estimation loss function, and train the contrast prediction coding model according to the loss function.
[0011] Step 5: Verify the feature expression ability of the logging data after encoding by the contrast prediction coding network, and complete the work of representing the logging data by contrast prediction coding.
[0012] Preferably, in Step 1, the preprocessing of the original logging data includes removing outliers and performing simple filtering on the measurement curves to remove sharp values generated by machine noise.
[0013] Preferably, in Step 1, the preprocessed original logging data is normalized. The normalization formula is as follows:
[0014]
[0015] where, x i represents the normalized logging data, x i represents the original logging data, X max and X min represent the maximum and minimum values of a certain curve in a well, respectively.
[0016] Preferably, in Step 2, the two-dimensional representation of the logging data is obtained according to the two-dimensional matrix. The specific process is as follows:
[0017] Set the numerical range of the two-dimensional matrix to linearly vary from 0 to 255 to obtain a single-channel grayscale image. The depth of each logging data has an interval distance. There are total data points in a well. Use a window with a length set to L meters to slide on the single-channel gray image, and set the sliding step size to s. Thus, the single-channel gray image can be segmented into a two-dimensional representation of the logging data with a size of L×6. The label determination formula for the two-dimensional representation of the logging data is as follows:
[0018] label pic = argmax(bincount(label[j:j + L]))
[0019] Among them, j is the sampling point of the j-th row, and j is used as a variable during transformation; label pic is the geological layer label of the two-dimensional characterization of well logging data; argmax is the index of the geological layer label that appears the most times; bincount is the number of times each geological layer label appears; label is the label corresponding to the one-dimensional sampling point of the well logging curve.
[0020] Preferably, in step 4, a contrastive prediction coding network structure is built according to the two-dimensional characterization of well logging data. The contrastive prediction coding network structure includes an encoder and an autoregressive network; the two-dimensional characterization of well logging data is transformed into an embedding space, and prediction coding is performed through an autoregressive model in the space, and the error between positive and negative samples is calculated using mutual information. In the actual calculation process, the mutual information between multiple samples is calculated through a logarithmic model, and the calculation formula is:
[0021]
[0022] where x t+k represents the positive sample, represents the feature representation of the positive sample, c t represents the current cumulative feature;
[0023] Relying on the noise contrastive estimation loss function, the loss function calculation formula of the contrastive prediction coding network is:
[0024]
[0025] where, x t+k represents the positive sample, and the others are negative samples, c t represents the current cumulative feature, and E represents the average value of N randomly input samples.
[0026] Preferably, in step 4, the specific process of training the contrastive prediction coding model according to the loss function is as follows:
[0027] First, perform well logging sample selection, process the training set, and construct it in the form of anchor data, positive samples, and negative sample pairs; contrastive prediction coding learns by comparing the mutual information between the current cumulative feature and positive and negative samples. Appropriate anchor data is selected based on the characteristics of well logging data and the temporal information of the two-dimensional characterization of well logging data. The samples after selecting the anchor data are used as positive samples, and the samples with labels different from the positive samples are used as negative samples;
[0028] Train the network model, train the contrast prediction coding network using the training set, and plot the loss curve and accuracy curve during the training process; according to the contrast prediction coding network model and the input training set, use the well logging sample selection method to construct training data, input it into the network model, train the network model, set the hyperparameters of the number of training epochs and the model update frequency, and plot the loss curve and accuracy change curve during the training process.
[0029] Furthermore, the specific process of well logging sample selection is as follows:
[0030] Select 5 consecutive two-dimensional representations of well logging data as the anchor data along the depth sequence;
[0031] After selecting the anchor data, select the two-dimensional representation of the third well logging data in the depth order as the positive sample. If the label category of this sample is different from that of the anchor data, then select the interface as the positive sample;
[0032] For the data with different categories from the positive sample, use the stratified sampling method to select 100 two-dimensional representations of well logging data for each category;
[0033] In each category, use the Kmeans clustering algorithm to cluster these 100 two-dimensional representations of well logging data into three categories, select the three cluster centers of each category as the negative samples of the current category, and a total of 27 negative samples are selected for each positive sample.
[0034] A well logging data representation system based on contrast prediction coding includes:
[0035] The first data processing module is used to obtain the original well logging data, preprocess the original well logging data, and perform standardization processing on the preprocessed original well logging data to obtain several columns of well logging data curves;
[0036] The second data processing module is used to combine several columns of well logging data curves to form a two-dimensional matrix, and obtain the two-dimensional representation of well logging data according to the two-dimensional matrix;
[0037] The data division module is used to divide the data set into the training set and test set used in the learning environment by using the obtained two-dimensional representation of well logging data, where the training set part is used to train the contrast prediction coding network model, and the test set is used to verify and evaluate the model performance;
[0038] The data modeling module is used to build a contrast prediction coding network structure according to the two-dimensional representation of well logging data, rely on the noise contrast estimation loss function, and train the contrast prediction coding model according to the loss function;
[0039] The data verification module is used to verify the feature expression ability of the well logging data after encoding by the contrast prediction coding network, and complete the well logging data representation work of contrast prediction coding.
[0040] A mobile terminal includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of a method for representing logging data based on contrast prediction coding as described above are implemented.
[0041] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of a method for representing logging data based on contrast prediction coding as described above are implemented.
[0042] Compared with the prior art, the present invention has the following beneficial technical effects:
[0043] The present invention provides a method for representing logging data based on contrast prediction coding. By preprocessing and normalizing the logging data, several columns of logging data curves are obtained; and the two-dimensional characterization of the logging data is transformed. The obtained two-dimensional characterization of the logging data is randomly divided into a training set and a test set used for the learning environment, and a contrast prediction coding network structure is built. The selection and training of logging samples are carried out in the network model, and the loss curve and accuracy change curve of the training process can be effectively drawn; the present invention reconstructs the logging feature representation for the original six-dimensional logging curve data. Using the reconstructed logging feature representation can obtain good logging interpretation effects under the condition of small-sample logging data. It reduces the errors that may be brought by manual judgment, improves the work efficiency of the staff, and reduces the cost loss in the production process. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 It is a flow chart of the method for representing logging data based on contrast prediction coding in the present invention;
[0045] Figure 2 It is a flow chart of the application of the contrast prediction coding network in the geological layer division task in the present invention;
[0046] Figure 3 It is a schematic diagram of the two-dimensional characterization conversion step of the logging data in the present invention;
[0047] Figure 4 It is a schematic diagram of the contrast prediction coding network structure in the present invention;
[0048] Figure 5 It is the loss change curve and accuracy change curve during the training process of the contrast prediction coding network in the present invention;
[0049] Figure 6 It is the loss change curve and accuracy change curve during the training process of the geological layer division task in the present invention;
[0050] Figure 7This is a comparison result graph between the geological layers identified by the logging data representation method in the present invention and the actual geological layers. Detailed implementation manners
[0051] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0052] The present invention will be further described in detail below with reference to the accompanying drawings:
[0053] The purpose of the present invention is to provide a logging data representation method, system, terminal and medium based on contrast prediction coding, so as to solve the technical problems of low manpower, material resources, efficiency, large differences in division results and high error rates in manual logging interpretation in the prior art.
[0054] See Figure 1 and Figure 2 In an embodiment of the present invention, a logging data representation method based on contrast prediction coding is provided, including the following steps:
[0055] Step 1, obtain the original logging data, preprocess the original logging data, and perform standardization processing on the preprocessed original logging data to obtain several columns of logging data curves;
[0056] Specifically, the preprocessing of the original logging data includes removing outliers and performing simple filtering on the measurement curves to remove sharp values generated by machine noise.
[0057] Specifically, the standardization processing of the preprocessed original logging data is carried out, and the standardization processing formula is as follows:
[0058]
[0059] where x i represents the normalized logging data, x i represents the original logging data, X max and X min respectively represent the maximum and minimum values of a certain curve in a well.
[0060] Step 2, combine several columns of logging data curves to form a two-dimensional matrix, and obtain a two-dimensional representation of the logging data according to the two-dimensional matrix;
[0061] Specifically, the process of obtaining the two-dimensional representation of the logging data according to the two-dimensional matrix is as follows:
[0062] Set the numerical range of the two-dimensional matrix to linearly vary from 0 to 255 to obtain a single-channel grayscale image. The depth of each well logging data has an interval distance. One well has a total of total data points. Use a window with a length of L meters to slide on the single-channel gray image, and set the sliding step size to s. Thus, the single-channel gray image can be segmented into a two-dimensional representation of well logging data with a size of L×6. The formula for determining the label of the two-dimensional representation of well logging data is as follows:
[0063] label pic = argmax(bincount(label[j:j + L]))
[0064] where j is the sampling point of the j-th row, and j is used as a variable during transformation; label pic is the geological layer label of the two-dimensional representation of well logging data; argmax is the index of the geological layer label with the most occurrences; bincount is the number of occurrences of each geological layer label; label is the label corresponding to the one-dimensional sampling point of the well logging curve.
[0065] Step 3: Use the obtained two-dimensional representation of well logging data to divide the training set and test set used in the learning environment of the dataset. The training set part is used to train the contrast prediction coding network model, and the test set is used to verify and evaluate the model performance; the setting of the reward and punishment rules includes but is not limited to the above description, and the parameter values of rewards and punishments can be flexibly set according to the input data.
[0066] Step 4: Build a contrast prediction coding network structure based on the two-dimensional representation of well logging data, rely on the noise contrast estimation loss function, and train the contrast prediction coding model according to the loss function;
[0067] Specifically, build a contrast prediction coding network structure based on the two-dimensional representation of well logging data. The contrast prediction coding network structure includes an encoder and an autoregressive network; transform the two-dimensional representation of well logging data into the embedding space, and predict and code in the space through the autoregressive model. Use mutual information to calculate the error between positive and negative samples. In the actual calculation process, calculate the mutual information between multiple samples through the logarithmic model. The calculation formula is:
[0068]
[0069] where x t+k represents the positive sample, represents the feature representation of the positive sample, c t represents the current cumulative feature;
[0070] Rely on the noise contrast estimation loss function. The calculation formula for the loss function of the contrast prediction coding network is:
[0071]
[0072] Among them, x t+k represents the positive sample, and the others are negative samples. c t represents the current cumulative feature, and E represents the average value of N randomly input samples.
[0073] Specifically, the specific process of training the contrast prediction coding model according to the loss function is as follows:
[0074] First, perform well logging sample selection, process the training set, and construct it in the form of anchor data, positive samples, and negative sample pairs; contrast prediction coding learns by comparing the mutual information between the current cumulative feature and positive and negative samples. Appropriate anchor data is selected based on the characteristics of well logging data and the time series information of the two-dimensional representation of well logging data. The samples after selecting the anchor data are used as positive samples, and the samples with labels different from the positive samples are used as negative samples;
[0075] Train the network model, use the training set to train the contrast prediction coding network, and draw the loss curve and accuracy curve during the training process; according to the contrast prediction coding network model, read the training set, use the well logging sample selection method to construct the training data, input it into the network model, train the network model, set the hyperparameters of the number of training epochs and the model update frequency, and draw the loss curve and accuracy change curve during the training process.
[0076] Among them, the specific process of well logging sample selection is as follows:
[0077] Select 5 consecutive two-dimensional representations of well logging data as anchor data along the depth sequence;
[0078] After selecting the anchor data, use the third two-dimensional representation of well logging data in the depth order as the positive sample. If the label category of this sample is different from the anchor data category, then select the interface as the positive sample;
[0079] For the data with categories different from the positive sample, use the stratified sampling method to select 100 two-dimensional representations of well logging data for each category;
[0080] In each category, use the Kmeans clustering algorithm to cluster these 100 two-dimensional representations of well logging data into three categories, and select the three cluster centers of each category as the negative samples of the current category. A total of 27 negative samples are selected for each positive sample.
[0081] Among them, the well logging sample selection algorithm in the present invention is as follows:
[0082]
[0083] Step 5: Verify the feature expression ability of the well logging data encoded by the contrast prediction coding network, and complete the representation of the well logging data by contrast prediction coding.
[0084] Embodiment
[0085] This embodiment provides a method for representing well logging data based on contrast prediction coding. The specific process is as follows:
[0086] (1) Remove outliers from the provided original data, such as values outside the reasonable range like 0 and -9999, and perform simple filtering on the measurement curves to remove sharp values caused by machine noise, etc. To ensure the features used for unified training, the data needs to be standardized.
[0087] (2) After the well logging data undergoes outlier processing and normalization, it becomes a curve that varies according to depth for each column. Combine each column to form a two-dimensional matrix, and linearly transform the numerical range of this two-dimensional matrix to 0 to 255, from which a single-channel grayscale image can be obtained. The depth interval of each well logging data is 0.125 meters, and a well has total data points. Use a window with a length of L meters to slide on the single-channel grayscale image with a sliding step of s. Thus, the single-channel grayscale image can be segmented into a two-dimensional representation of well logging data with a size of L×6. The conversion process of the two-dimensional representation of well logging data is as Figure 3 shown.
[0088] (3) Use the method in the second step to convert the well logging data into a two-dimensional representation of well logging data. Randomly divide the data set into a training set and a test set used for the learning environment, with a division ratio of 0.8:0.2. The training set part is used to train the contrast prediction coding network model, and the test set is used to verify and evaluate the model performance.
[0089] (4) Transform the two-dimensional representation of well logging data into a more compact latent embedding space to make conditional prediction easier to model. Secondly, a powerful autoregressive model is used in this latent space to predict many future steps. Finally, use a loss function that depends on noise contrast estimation to train the entire model. The schematic diagram of the network model is as Figure 4 shown.
[0090] (5) Process the training set and construct it in the form of anchor data, positive samples, and negative sample pairs. Select the samples after the anchor data as positive samples, and select the samples with labels different from the positive samples as negative samples.
[0091] (6) Train the contrast prediction coding network using the training set, and plot the loss curve and accuracy curve during the training process. According to the network model built in step (4), read in the training set in step (3), construct training data using the well logging sample selection method in the step, input it into the network model, train the network model, set hyperparameters such as the number of training epochs and the model update frequency, and plot the loss curve and accuracy change curve during the training process. The results are as Figure 5 shown.
[0092] (7) Add a linear classification layer after the encoder, fix the structure and parameters of the contrast prediction coding network, verify the feature expression ability of the well logging data after encoding by the contrast prediction coding network, and use the data in the test set to train only the linear classifier. The accuracy and loss change curves during the training process are as Figure 6 shown. Take the accuracy of the well logging features encoded by the test set in the geological layer division task as the judgment criterion. The comparison result of a certain well is as Figure 7 shown. The actual layer is the result of expert interpretation, and the intelligent layer is the result identified by this method. The accuracy comparison of different methods is shown in Table 1.
[0093] Table 1: Accuracy comparison of different methods (%)
[0094]
[0095]
[0096] Judging from the above results, the obvious advantages of the method of the present invention are as follows: The method can achieve an accuracy of 85% in geological layer division under the condition of a training set with only 10 well logging data, exceeding the traditional method and the convolutional neural network.
[0097] The present invention also provides a contrast prediction coding based well logging data representation system, including a first data processing module, a second data processing module, a data division module, a data modeling module, and a data verification module;
[0098] The first data processing module is used to obtain the original well logging data, preprocess the original well logging data, and perform standardization processing on the preprocessed original well logging data to obtain several columns of well logging data curves;
[0099] The second data processing module is used to combine several columns of well logging data curves to form a two-dimensional matrix, and obtain a two-dimensional representation of the well logging data according to the two-dimensional matrix;
[0100] The data division module is used to divide the data set into a training set and a test set used in the learning environment by using the two-dimensional representation of the well logging data obtained, where the training set part is used to train the contrast prediction coding network model, and the test set is used to verify and evaluate the model performance;
[0101] A data modeling module, configured to build a contrast prediction coding network structure based on the two-dimensional characterization of logging data, rely on a noise contrast estimation loss function, and train a contrast prediction coding model according to the loss function;
[0102] A data verification module, configured to verify the feature expression ability of the logging data after being encoded by the contrast prediction coding network, and complete the work of representing the logging data by contrast prediction coding.
[0103] The present invention also provides a mobile terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor, such as a program for representing logging data based on contrast prediction coding.
[0104] When the processor executes the computer program, the steps of the above method for representing logging data based on contrast prediction coding are implemented, including the following steps:
[0105] Step 1, obtain original logging data, preprocess the original logging data, and perform standardization processing on the preprocessed original logging data to obtain several columns of logging data curves;
[0106] Step 2, combine several columns of logging data curves to form a two-dimensional matrix, and obtain a two-dimensional characterization of the logging data according to the two-dimensional matrix;
[0107] Step 3, use the obtained two-dimensional characterization of the logging data to divide the dataset into a training set and a test set used for the learning environment, where the training set is used to train the contrast prediction coding network model, and the test set is used to verify and evaluate the model performance;
[0108] Step 4, build a contrast prediction coding network structure based on the two-dimensional characterization of the logging data, rely on a noise contrast estimation loss function, and train a contrast prediction coding model according to the loss function;
[0109] Step 5, verify the feature expression ability of the logging data after being encoded by the contrast prediction coding network, and complete the work of representing the logging data by contrast prediction coding.
[0110] Alternatively, when the processor executes the computer program, the functions of each module in the above system are implemented, for example:
[0111] A first data processing module, configured to obtain original logging data, preprocess the original logging data, and perform standardization processing on the preprocessed original logging data to obtain several columns of logging data curves;
[0112] A second data processing module, configured to combine several columns of logging data curves to form a two-dimensional matrix, and obtain a two-dimensional characterization of the logging data according to the two-dimensional matrix;
[0113] A data partitioning module, configured to partition the data set into a training set and a test set used by the learning environment by using the two-dimensional characterization of the logging data, where the training set is used to train the contrast prediction coding network model, and the test set is used to verify and evaluate the model performance;
[0114] A data modeling module, configured to build a contrast prediction coding network structure according to the two-dimensional characterization of the logging data, rely on the noise contrast estimation loss function, and train the contrast prediction coding model according to the loss function;
[0115] A data verification module, configured to verify the feature expression ability of the logging data after being encoded by the contrast prediction coding network, and complete the logging data representation work of the contrast prediction coding.
[0116] Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the mobile terminal. For example, the computer program may be divided into a first data processing module, a second data processing module, a data partitioning module, a data modeling module, and a data verification module;
[0117] The specific functions of each module are as follows:
[0118] A first data processing module, configured to obtain the original logging data, preprocess the original logging data, and perform standardization processing on the preprocessed original logging data to obtain several columns of logging data curves;
[0119] A second data processing module, configured to combine several columns of logging data curves to form a two-dimensional matrix, and obtain the two-dimensional characterization of the logging data according to the two-dimensional matrix;
[0120] A data partitioning module, configured to partition the data set into a training set and a test set used by the learning environment by using the two-dimensional characterization of the logging data, where the training set is used to train the contrast prediction coding network model, and the test set is used to verify and evaluate the model performance;
[0121] A data modeling module, configured to build a contrast prediction coding network structure according to the two-dimensional characterization of the logging data, rely on the noise contrast estimation loss function, and train the contrast prediction coding model according to the loss function;
[0122] A data verification module, configured to verify the feature expression ability of the logging data after being encoded by the contrast prediction coding network, and complete the logging data representation work of the contrast prediction coding.
[0123] The mobile terminal may be a computing device such as a desktop computer, a notebook, a handheld computer, and a cloud server. The mobile terminal may include, but is not limited to, a processor and a memory.
[0124] The processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the mobile terminal, and connects various parts of the entire mobile terminal through various interfaces and lines.
[0125] The memory may be used to store the computer program and / or module. The processor realizes various functions of the mobile terminal by running or executing the computer program and / or module stored in the memory, and by calling the data stored in the memory.
[0126] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0127] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of a method for representing logging data based on contrast prediction coding are realized.
[0128] If the modules / units integrated in the mobile terminal are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium.
[0129] Based on such understanding, all or part of the processes in the above methods of the present invention can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method for representing logging data based on contrast prediction coding can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc.
[0130] The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0131] It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent substitutions can still be made to the specific embodiments of the present invention, and any modification or equivalent substitution that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A method for representing logging data based on contrastive predictive coding, characterized in that, it includes the following steps: Step 1, obtain the original logging data, preprocess the original logging data, and perform standardization processing on the preprocessed original logging data to obtain several columns of logging data curves; Step 2, combine several columns of logging data curves to form a two-dimensional matrix, and obtain a two-dimensional representation of the logging data according to the two-dimensional matrix; Step 3, use the obtained two-dimensional representation of the logging data to divide the training set and the test set used in the learning environment of the dataset, where the training set part is used to train the contrastive predictive coding network model, and the test set is used to verify and evaluate the model performance; Step 4, build a contrastive predictive coding network structure according to the two-dimensional representation of the logging data, rely on the noise contrast estimation loss function, and train the contrastive predictive coding model according to the loss function; Step 5, verify the feature expression ability of the logging data after encoding by the contrastive predictive coding network, and complete the work of representing the logging data by contrastive predictive coding.
2. A method for representing logging data based on contrastive predictive coding according to claim 1, characterized in that, in Step 1, the preprocessing of the original logging data includes removing outliers and performing simple filtering on the measurement curve to remove sharp values generated by machine noise.
3. A method for representing logging data based on contrastive predictive coding according to claim 1, characterized in that, in Step 1, perform standardization processing on the preprocessed original logging data, and the formula for the standardization processing is as follows: Among them, x i represents the normalized logging data, and x i represents the original logging data. X max and X min respectively represent the maximum and minimum values of a certain curve in a well.
4. A method for representing logging data based on contrastive predictive coding according to claim 1, characterized in that, in Step 2, obtain a two-dimensional representation of the logging data according to the two-dimensional matrix, and the specific process is as follows: Set the numerical range of the two-dimensional matrix to linearly vary from 0 to 255 to obtain a single-channel grayscale image. The depth of each logging data has an interval distance. One well has total data points. Use a window with a length set to L meters to slide on the single-channel gray image, and set the sliding step size to s. Thus, the single-channel gray image can be segmented into a two-dimensional representation of the logging data with a size of L×6. The formula for determining the label of the two-dimensional representation of the logging data is as follows: label pic = argmax(bincount(label[j:j+L])) Among them, j is the sampling point of the j-th row, and j is used as a variable during transformation; label pic is the geological layer label of the two-dimensional characterization of well logging data; argmax is the index of the geological layer label with the most occurrences; bincount is the number of occurrences of each geological layer label; label is the label corresponding to the one-dimensional sampling point of the well logging curve.
5. A method for representing logging data based on contrastive predictive coding according to claim 1, characterized in that, Step 4, build a contrastive predictive coding network structure according to the two-dimensional representation of the logging data. The contrastive predictive coding network structure includes an encoder and an autoregressive network; transform the two-dimensional representation of the logging data into the embedding space, and perform predictive coding in the space through the autoregressive model. Use mutual information to calculate the error between positive and negative samples. In the actual calculation process, calculate the mutual information between multiple samples through the logarithmic model. The calculation formula is: where x t+k represents a positive sample, represents the feature representation of the positive sample, c t represents the current cumulative feature; Rely on the noise contrast estimation loss function, and the calculation formula for the loss function of the contrastive predictive coding network is: Among them, x t+k represents the positive sample, and the others are all negative samples. c t represents the current cumulative feature, and E represents the average value of the N random samples input.
6. A method for representing logging data based on contrastive predictive coding according to claim 1, characterized in that, in Step 4, the specific process of training the contrastive predictive coding model according to the loss function is as follows: First, perform well logging sample selection, process the training set, and construct it in the form of anchor data, positive samples, and negative sample pairs. The contrast prediction coding learns by comparing the mutual information between the current cumulative features and the positive and negative samples. Appropriate anchor data is selected based on the characteristics of well logging data and the temporal information of the two-dimensional representation of well logging data. The samples after selecting the anchor data are used as positive samples, and the samples with different labels from the positive samples are used as negative samples. Train the network model, use the training set to train the contrast prediction coding network, and draw the loss curve and accuracy curve during the training process. According to the contrast prediction coding network model, the input training set, use the well logging sample selection method to construct training data, input it into the network model, train the network model, set the hyperparameters of the number of training epochs and the model update frequency, and draw the loss curve and accuracy change curve during the training process.
7. A method for representing well logging data based on contrast prediction coding according to claim 6, characterized in that, the specific process of the well logging sample selection is as follows: Select 5 consecutive two-dimensional representations of well logging data as anchor data along the depth sequence; Select the third two-dimensional representation of well logging data in the depth order after selecting the anchor data as the positive sample. If the label category of this sample is different from the anchor data category, then select the interface as the positive sample; For the data with different categories from the positive sample, use the stratified sampling method to select 100 two-dimensional representations of well logging data for each category; In each category, use the Kmeans clustering algorithm to cluster these 100 two-dimensional representations of well logging data into three categories, and select the three cluster centers of each category as the negative samples of the current category. A total of 27 negative samples are selected for each positive sample.
8. A system for representing well logging data based on contrast prediction coding, characterized in that, comprising: The first data processing module is used to obtain the original well logging data, preprocess the original well logging data, and perform standardization processing on the preprocessed original well logging data to obtain several columns of well logging data curves; The second data processing module is used to combine several columns of well logging data curves to form a two-dimensional matrix, and obtain the two-dimensional representation of well logging data according to the two-dimensional matrix; The data division module is used to divide the data set into the training set and the test set used in the learning environment by using the obtained two-dimensional representation of well logging data, where the training set part is used to train the contrast prediction coding network model, and the test set is used to verify and evaluate the model performance; The data modeling module is used to build a contrast prediction coding network structure according to the two-dimensional representation of well logging data, rely on the noise contrast estimation loss function, and train the contrast prediction coding model according to the loss function; The data verification module is used to verify the feature expression ability of the well logging data after the contrast prediction coding network encodes, and complete the work of representing the well logging data by contrast prediction coding.
9. A mobile terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the steps of a method for representing well logging data based on contrast prediction coding according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, when the computer program is executed by a processor, the steps of a method for representing logging data based on contrastive predictive coding according to any one of claims 1-7 are implemented.