Well logging data anomaly detection method

By introducing a multi-layer feature extraction layer and a risk weighting matrix, the difficulty of identifying outliers caused by noise interference in well logging data is solved, thereby improving the accuracy and robustness of well logging data anomaly detection.

CN116522125BActive Publication Date: 2025-12-09CHINA PETROLEUM & CHEMICAL CORP +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210071283.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-21
Publication Date
2025-12-09
Estimated Expiration
2042-01-21

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively identify and process outliers caused by noise interference in well logging data, particularly in lithology identification, and supervised learning methods lack robustness.

Method used

By employing a multi-layer feature extraction layer and an intra-class scatter matrix to construct abstract features in the feature extraction layer, and introducing a risk weighting matrix in the output classification layer, the robustness of the training results to noise is improved.

Benefits of technology

By introducing a multi-layer feature extraction layer and a risk weighting matrix, the resolution and robustness to noise in well logging data anomaly detection are improved, thereby enhancing the accuracy of outlier identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116522125B_ABST
    Figure CN116522125B_ABST
Patent Text Reader

Abstract

The application provides a well logging data anomaly detection method, which comprises the following steps: step 1, constructing a well logging data sample set; step 2, performing system initialization; step 3, performing feature extraction layer training; step 4, outputting classification layer training; and step 5, performing anomaly value identification. The well logging data anomaly detection method introduces a plurality of feature extraction layers, and introduces an intra-class scatter matrix when the feature extraction layers are constructed, so that abstract features with resolution can be obtained. When the output classification layer is trained, a risk weighting matrix is introduced, so that the robustness of the training result to noise is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of tectonic geology and petroleum geology, and particularly relates to a well logging data anomaly detection method. BACKGROUND

[0002] In data mining, anomaly detection is the identification of items, events or observations that do not match the expected pattern or other items in the dataset. Often, the anomalous items will translate to some kind of problem such as bank frauds, structural issues, medical problems, text errors, etc. Anomalies are also called outliers, novelties, noise or exceptions. Common anomaly detection methods include: density-based methods, subspace and correlation-based outlier detection for high-dimensional data, one-class support vector machines, etc. When collecting well logging data, due to the influence of well logging fluid and rock debris, the well logging values at certain depths may be abnormal, which requires us to study the anomaly detection of well logging data.

[0003] In the Chinese patent application with the application number CN201910308560.6, a time series anomaly value detection method and device are involved. The method includes: obtaining historical data belonging to the same time series within a preset time length before the detection time; performing standardization processing on the historical data; and determining whether the data at the detection time is an anomaly value according to the historical data after standardization processing and a pre-trained calculation model. This method essentially utilizes time correlation and has strong hypothesis conditions.

[0004] In the Chinese patent application with the application number CN201610970187.7, an anomaly value detection method and system in an LTE network are involved. The measured data is divided into a training set and a test set, clusters and parameters are defined in the training set, the clustering algorithm is used to find the cluster to which each data point belongs, the likelihood value of each data point is calculated according to the parameter value and the clustering result, and the likelihood value is divided into an abnormal area, an intermediate area and a normal area according to the set early warning threshold and alarm threshold; the model that has been calculated is applied to the test set, the likelihood value of each data point in the test set is calculated, and each data point is divided into the areas, so that the anomaly value in the test set is found. Although this method uses machine learning, it is essentially an unsupervised method and has great uncertainty. In addition, there are few reports on anomaly detection of well logging data. Well logging data has obvious physical significance and can be used for lithology identification, so a supervised learning method can be used. At the same time, well logging data often has certain noise, which brings great difficulty to anomaly detection.

[0005] In the Chinese patent application with the application number CN201611123548.0, a micro-resistivity scanning imaging logging data anomaly correction method and device are related to the oil engineering logging technical field. The method first collects logging data, judges whether the logging data is abnormal and the abnormal type, then calculates the average value of the adjacent normal data of the abnormal data according to the abnormal type, then calculates the reliability of the abnormal data by using the similarity of the abnormal data and the adjacent normal data, and judges whether the abnormal data is reliable according to the reliability, finally, when the abnormal data is not reliable, the average value is used to replace the abnormal data as the correction value, and when the abnormal data is reliable, the change characteristics of the abnormal data are extracted, and the change characteristics and the corresponding average value are added as the correction value of the abnormal data. The correction method can make the corrected abnormal data consistent with the adjacent normal data in size as a whole, and at the same time, the change characteristics of the abnormal data are retained in the local, and after visual imaging, a good imaging effect is obtained.

[0006] The above prior art is quite different from the present application, and cannot solve the technical problems we want to solve. Therefore, we have invented a new logging data anomaly detection method. SUMMARY

[0007] The purpose of the present application is to provide a logging data anomaly detection method which can obtain abstract features with resolution and improve the robustness of the training results to noise.

[0008] The purpose of the present application can be achieved by the following technical measures: a logging data anomaly detection method, which comprises

[0009] The purpose of the present application can also be achieved by the following technical measures:

[0010] The logging data anomaly detection method comprises:

[0011] Step 1, constructing a logging data sample set;

[0012] Step 2, system initialization;

[0013] Step 3, feature extraction layer training;

[0014] Step 4, output classification layer training;

[0015] Step 5, abnormal value identification.

[0016] In step 1, logging data of one or more wells is obtained, all logging values at each depth form a feature vector, i.e. a sample, and a sample is expressed as wherein represents a real number field, and d is the sample dimension; the sample set composed of logging data is sample matrix where n is the total number of samples in the set; is the lithology-labeled well logging data.

[0017] In step 1, the well logging values include natural gamma, natural potential, caliper, acoustic time, density, compensated neutron, deep lateral resistivity and shallow lateral resistivity.

[0018] In step 2, let the experience loss coefficient τ>0, the learning rate coefficients a, b>0, the stop coefficient ∈>0, the maximum number of feature extraction layers m be a positive integer, the number of high-dimensional features N h be a positive integer; let the Gaussian kernel width σ>0, the sample rejection coefficient α∈(0, 100); the feature extraction layer weight matrix be a zero matrix, the momentum vector be a zero matrix.

[0019] In step 2, τ∈(10,10 3 ), a∈(0.5,1), b∈(0.01,0.1), ∈=10 -3 , m>5, a∈(0.1,1), α=10, N h >200.

[0020] Step 3 includes:

[0021] Step 301, initialization: let i=1, Y0=X, is the kth row of Y0;

[0022] Step 302, randomly generate the weight matrix Randomly generate the bias vector Then construct the high-dimensional feature matrix The kth column of H is φ is an activation function;

[0023] Step 303, calculate the weight matrix of the ith feature extraction layer

[0024] Step 304, let i increase by 1, if i>m, jump to step 4, otherwise jump to step 302.

[0025] In step 302, the activation function uses Sigmoid, Tanh or Softmax.

[0026] In step 302, the probability distribution used by the random generation method is Gaussian distribution or uniform distribution.

[0027] Step 303 includes:

[0028] Step 30301, according to Solve the within-class or overall scatter matrix S;

[0029] Step 30302, calculate the gradient where If jump to step 30303, if assign the value of current θ to Θ i , let Y i = φ(Θ i Y i-1 ), then jump to step 304;

[0030] Step 30303, calculate the updated momentum vector:

[0031]

[0032] Then calculate the updated weight matrix θ * = θ - bv;

[0033] Step 30304, assign the value of θ * to θ, and the value of v * to v, then jump to step 30302.

[0034] Step 4 includes:

[0035] Step 401, for each sample , calculate its encoded feature representation as follows: where

[0036] Step 402, for each sample , calculate the error where is an all-one vector, and I is an identity matrix;

[0037] Step 403, calculate the risk-weighted matrix M is a diagonal matrix, and the kth diagonal element is

[0038] Step 404, calculate the output weight matrix as follows:

[0039]

[0040] In step 5, for a new sample set p is the number of samples in the new sample set, and calculate its encoded feature representation as follows: where Then calculate the prediction error right Sort the samples from largest to smallest, and the samples corresponding to the top α% of prediction errors in the sorted set are the outliers.

[0041] The well logging data anomaly detection method of this invention includes the following steps: constructing a well logging data sample set, system initialization, training a feature extraction layer, training an output classification layer, and outlier identification. Compared with existing technologies, this invention introduces multiple feature extraction layers and incorporates an intra-class scatter matrix during the construction of the feature extraction layers, enabling the acquisition of abstract features with discriminative power; a risk weighting matrix is ​​introduced during the training of the output classification layer, improving the robustness of the training results to noise. Attached Figure Description

[0042] Figure 1 This is a flowchart of a specific embodiment of the well logging data anomaly detection method of the present invention. Detailed Implementation

[0043] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0044] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the present invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, and / or combinations thereof.

[0045] The well logging data anomaly detection method of the present invention includes the following steps: constructing a well logging data sample set, system initialization, feature extraction layer training, output classification layer training, and outlier identification.

[0046] The following are several embodiments of applying the present invention.

[0047] Example 1

[0048] In a specific embodiment 1 of the present invention, such as Figure 1 As shown, Figure 1 This is a flowchart of the well logging data anomaly detection method of the present invention.

[0049] Step 1, constructing a log data sample set

[0050] Obtain the log data of one or more wells, all log values at each depth form a feature vector, that is, a sample, the log values include but are not limited to natural gamma, natural potential, caliper, acoustic time difference, density, compensated neutron, deep lateral resistivity and shallow lateral resistivity; a sample is expressed as Wherein represents a real number field, d is the sample dimension; the sample set composed of log data is Sample matrix Wherein n is the total number of samples in the set; the log data is labeled with lithology

[0051] Step 2, system initialization

[0052] Let the experience loss coefficient τ>0, the learning rate coefficient a,b>0, the stop coefficient ∈>0, the maximum feature extraction layer m is a positive integer, and the high-dimensional feature quantity N h Is a positive integer; let the Gaussian kernel width σ>0, the sample rejection coefficient α∈(0,100); the feature extraction layer weight matrix Is a zero matrix, and the momentum vector Is a zero matrix;

[0053] Wherein, τ∈(10,10 3 ), a∈(0.5,1), b∈(0.01,0.1), ∈=10 -3 , m>5, a∈(0.1,1), α=10, N h >200.

[0054] Step 3, feature extraction layer training

[0055] Step 301, initialization: let i=1, Y0=X, Is the kth row of Y0;

[0056] Step 302, randomly generate weight matrix Randomly generate bias vector Then construct the high-dimensional feature matrix The kth column of H is φ is the activation function;

[0057] The activation function adopts Sigmoid.

[0058] Step 303, calculate the weight matrix of the i-th feature extraction layer As follows:

[0059] Step 30301, according to Solve the within-class or total scatter matrix S;

[0060] Step 30302, calculate gradient where If go to step 30303, if assign the value of current θ to Θ i , let Y i = φ(Θ i Y i-1 ), then go to step 304

[0061] Step 30303, calculate updated momentum vector Then calculate updated weight matrix θ * = θ - bv;

[0062] Step 30304, assign the value of θ * to θ, and the value of v * to v, then go to step 30302;

[0063] Step 304, let i increase by 1, if i > m, go to step 4, otherwise go to step 302;

[0064] Step 4, output classification layer training

[0065] Step 401, for each sample , calculate its encoded feature representation as follows: where

[0066] Step 402, for each sample , calculate error where is an all-1 vector, and I is an identity matrix;

[0067] Step 403, calculate risk weighting matrix M is a diagonal matrix, whose kth diagonal element is

[0068] Step 404, calculate output weight matrix as follows:

[0069]

[0070] Step 5, outlier identification

[0071] For a new sample set p is the number of samples in the new sample set, calculate its encoded feature representation as follows: where Then the prediction error is calculated The prediction error is calculated The samples corresponding to the top α% of the prediction errors in the sorted set are the abnormal samples.

[0072] Embodiment 2

[0073] In a specific embodiment 2 of the application, the well logging data anomaly detection method of the application comprises the following steps:

[0074] Step 1, constructing a well logging data sample set

[0075] Obtain the well logging data of one or more wells, and all well logging values at each depth form a feature vector, i.e., a sample, wherein the well logging values include but are not limited to natural gamma, natural potential, caliper, acoustic time difference, density, compensated neutron, deep lateral resistivity, and shallow lateral resistivity; one sample is expressed as wherein represents a real number field, and d is the sample dimension; the sample set composed of well logging data is Sample matrix wherein n is the total number of samples in the set; and the well logging data is labeled with lithology.

[0076] Step 2, system initialization

[0077] Let the experience loss coefficient τ>0, the learning rate coefficient a,b>0, the stop coefficient ∈>0, the maximum feature extraction layer number m be a positive integer, and the high-dimensional feature number N h be a positive integer; let the Gaussian kernel width σ>0, the sample rejection coefficient α∈(0,100); the feature extraction layer weight matrix be a zero matrix, and the momentum vector be a zero matrix.

[0078] wherein τ∈(10,10 3 ), a∈(0.5,1), b∈(0.01,0.1), ∈=10 -3 , m>5, a∈(0.1,1), α=10, N h >200.

[0079] Step 3, feature extraction layer training

[0080] Step 301, initialization: let i=1, Y0=X, be the kth row of Y0;

[0081] Step 302, randomly generating a weight matrix Randomly generate a bias vector Then construct a high-dimensional feature matrix The kth column of H is φ is an activation function;

[0082] The activation function is Softmax, and the probability distribution used in the random generation method is Gaussian distribution.

[0083] Step 303, calculate the weight matrix of the ith feature extraction layer As follows:

[0084] Step 30301, according to Solve the within-class or total scatter matrix S;

[0085] Step 30302, calculate the gradient Wherein If Jump to step 30303, if Assign the value of the current θ to Θ i , let Y i = φ(Θ i Y i-1 ), and then jump to step 304

[0086] Step 30303, calculate the updated momentum vector Then calculate the updated weight matrix θ * = θ-bv;

[0087] Step 30304, assign the value of θ * to θ, and assign the value of v * to v, and then jump to step 30302;

[0088] Step 304, let i increase by 1, if i > m, jump to step 4, otherwise jump to step 302;

[0089] Step 4, output the classification layer training

[0090] Step 401, for each sample Calculate its encoded feature representation As follows: Wherein

[0091] Step 402, for each sample Calculate the error Wherein Is an element-wise vector of 1, and I is an identity matrix;

[0092] Step 403, calculate the risk-weighted matrix M is a diagonal matrix, and the kth diagonal element is

[0093] Step 404: Calculate the output weight matrix, as follows:

[0094]

[0095] Step 5: Outlier Identification

[0096] For the new sample set p is the number of samples in the new sample set, and we calculate its encoded feature representation. as follows: in Then calculate the prediction error. right Sort the samples from largest to smallest, and the samples corresponding to the top α% of prediction errors in the sorted set are the outliers.

[0097] Example 3

[0098] In a specific embodiment 3 of the present invention, the well logging data anomaly detection method of the present invention includes the following steps:

[0099] Step 1: Construct a well logging data sample set

[0100] Obtain logging data from one or more wells. All logging values ​​at each depth form a feature vector, i.e., a sample. These logging values ​​include, but are not limited to, natural gamma, spontaneous potential, wellbore, sonic transit time, density, compensated neutrons, deep lateral resistivity, and shallow lateral resistivity. A sample is expressed as... in The sample set consists of real number field d and sample dimension d; the sample set of well logging data is d = d + d + d. Sample matrix Where n is the total number of samples in the set; lithological labels are added to the well logging data;

[0101] Step 2: System Initialization

[0102] Let the empirical loss coefficient τ > 0, the learning rate coefficients a, b > 0, the stopping coefficient ∈ > 0, the maximum feature extraction layer number m be a positive integer, and the number of high-dimensional features N. h Let σ be a positive integer; let the Gaussian kernel width σ > 0, and the sample rejection coefficient α ∈ (0, 100); the feature extraction layer weight matrix. Zero matrix, momentum vector It is a zero matrix;

[0103] Where, τ∈(10,10 3 ), a∈(0.5,1), b∈(0.01,0.1), ∈=10 -3 , m>5, a∈(0.1,1), α=10, N h >200.

[0104] Step 3, feature extraction layer training

[0105] Step 301, initialization: let i = 1, Y0= X, The kth row of Y0;

[0106] Step 302, randomly generate weight matrix Randomly generate bias vector Then construct high-dimensional feature matrix The kth column of H is φ is an activation function;

[0107] The activation function uses Tanh, and the probability distribution used in the random generation method is uniform distribution.

[0108] Step 303, calculate the weight matrix of the ith feature extraction layer As follows:

[0109] Step 30301, according to Solve the within-class or total scatter matrix S;

[0110] Step 30302, calculate the gradient Where If Then jump to step 30303, if Assign the value of the current θ to Θ i , let Y i = φ(Θ i Y i-1 ), and then jump to step 304

[0111] Step 30303, calculate the updated momentum vector Then calculate the updated weight matrix θ * = θ - bv;

[0112] Step 30304, assign the value of θ * to θ, and the value of v * to v, and then jump to step 30302;

[0113] Step 304, let i increase by 1, if i > m, jump to step 4, otherwise jump to step 302;

[0114] Step 4, output classification layer training

[0115] Step 401, for each sample Calculate its encoded feature representation respectively As follows: Where

[0116] Step 402, sample Calculate error Wherein Is a 1 vector, I is a unit matrix;

[0117] Step 403, calculate risk weighted matrix M is a diagonal matrix, the k diagonal element is

[0118] Step 404, calculate output weight matrix, as follows:

[0119]

[0120] Step 5, outlier identification

[0121] For new sample set P is the number of samples in the new sample set, and the encoded feature representation is obtained As follows: Wherein Then calculate the prediction error For Sort from large to small, and the samples corresponding to the top α% of the prediction error in the sorted set are the abnormal samples.

[0122] Finally, it should be pointed out that: the above only for the preferred embodiments of the present application, and not for limiting the present application, although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, it still can be modified, or part of the technical features of the equivalent replacement. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application, should be included in the protection scope of the present application.

[0123] In addition to the technical features described in the specification, they are known to those skilled in the art.

Claims

1. A method of detecting anomalies in well logging data, characterized in that, The well logging data anomaly detection method comprises: Step 1, constructing a well logging data sample set; Step 2, system initialization; Step 3, feature extraction layer training; Step 4, output classification layer training; Step 5, anomaly value identification; In step 1, the logging data of one or more wells are acquired, and all logging values at each depth form a feature vector, i.e. a sample, and a sample is expressed as wherein represents a real number field, d is the sample dimension; and a sample set composed of the logging data is The sample matrix is wherein n is the total number of samples in the set; and the logging data is labeled with lithology. In step 1, the well logging values include natural gamma, natural potential, caliper, acoustic time difference, density, compensated neutron, deep lateral resistivity and shallow lateral resistivity; In step 2, let the experience loss coefficient τ> 0, the learning rate coefficient a, b> 0, the stop coefficient ∈> 0, the maximum feature extraction layer number m is a positive integer, and the high-dimensional feature number N h is a positive integer; let the Gaussian kernel width σ> 0, the sample rejection coefficient α∈(0, 100); the feature extraction layer weight matrix is a zero matrix, and the momentum vector is a zero matrix; Step 3 comprises: Step 301, initialization: let i = 1, Y0 = X, k = 1,..., n; is the kth row of Y0. Step 302, randomly generating a weight matrix randomly generating a bias vector Then construct a high-dimensional feature matrix The kth column of H is φ is an activation function; Step 303, calculating the weight matrix of the i-th feature extraction layer Step 304, let i increase by 1, if i>m, jump to step 4, otherwise jump to step 302; Step 303 comprises: Step 30301, according to Solving the within-class or overall scatter matrix S; Step 30302, compute gradient where If then go to step 30303, else if then assign the value of current θ to Θ i , let Y i = φ(Θ i Y i-1 ), then go to step 304; Step 30303, calculate the updated momentum vector: Then the updated weight matrix θ * = θ - bv; Step 30304, put θ * Assign the value of v to θ, and put v * The value is assigned to v, and then the process jumps to step 30302; Step 4 comprises: Step 401, obtaining a sample Respectively obtain the encoded feature representation As follows: Where i = 1,..., m, φ is an activation function, and θ is a feature extraction layer weight matrix. Step 402, sample Computing error wherein τ is an empirical loss coefficient, is an all-ones vector, and I is an identity matrix. Step 403, calculating a risk-weighted matrix M is a diagonal matrix, whose kth diagonal element is Step 404, calculate the output weight matrix as follows:

2. The method of claim 1, wherein, In step 2, τ∈(10, 10 3 ), a∈(0.5, 1), b∈(0.01, 0.1), ∈=10 -3 , m>5, a∈(0.1, 1), α=10, N h >200.

3. The method of claim 1, wherein, In step 302, the activation function adopts Sigmoid, Tanh or Softmax.

4. The method of claim 1, wherein, In step 302, the probability distribution adopted by the random generation method is Gaussian distribution or uniform distribution.

5. The method of claim 1, wherein, In step 5, for the new sample set p is the number of samples in the new sample set, and we calculate its encoded feature representation. as follows: in i = 1, ..., m; then calculate the prediction error. right Sort the samples from largest to smallest, and the samples corresponding to the top α% of prediction errors in the sorted set are the percentage of outlier samples.

Citation Information

Patent Citations

  • Outlier detection method and system in lte network

    CN106572493B

  • Method and device for correcting abnormal micro-resistivity scanning imaging logging data

    CN106646634A

  • Time series abnormal value detection method and device

    CN110334083A

  • Power equipment fault monitoring method based on mutual reconstruction single-class auto-encoder

    CN112381180A