Dimensionality reduction method, device and storage medium for post-casing electromagnetic logging curve

Through dynamic time regularization algorithm and hierarchical clustering algorithm, the dimensionality reduction of the electromagnetic well logging curve after sleeve is solved, and the problem of low accuracy and comprehensiveness of dimensionality reduction in the existing technology is achieved, a more accurate and comprehensive understanding of the stratigraphic structure and reservoir characteristics is achieved, and data processing efficiency is improved.

CN118734094BActive Publication Date: 2025-05-13SHENZHEN FENGHE SHUZHI TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410707833.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-03
Publication Date
2025-05-13
Estimated Expiration
2044-06-03

AI Technical Summary

Technical Problem

The accuracy and comprehensiveness of the dimensionality reduction method of the post-shut electromagnetic well logging curve in the prior art is low, which leads to difficulty in understanding the formation structure and reservoir characteristics, and reduces the accuracy and comprehensiveness of subsequent analysis.

Method used

The dynamic time regularization algorithm is used to calculate the similarity distance between the electromagnetic well logging curves after the sleeve, and the curves are clustered through the hierarchical clustering algorithm to obtain the curve after dimensionality reduction to retain key information about formation changes and reservoir characteristics.

Benefits of technology

By combining dynamic time regularization algorithm and hierarchical clustering algorithm, effective dimensionality reduction of the post-sleeved electromagnetic well logging curve is achieved, data interpretability and visualization effect are improved, the understanding of stratigraphic structure and reservoir characteristics is enhanced, and the efficiency and accuracy of data processing are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118734094B_ABST
    Figure CN118734094B_ABST
Patent Text Reader

Abstract

The present application provides a method, device, storage medium and computer equipment for reducing the dimension of a post-casing electromagnetic logging curve, which realizes the dimensionality reduction of the post-casing electromagnetic logging curve by combining a dynamic warping algorithm and a hierarchical clustering algorithm. The reduced-dimensional curve can retain key information for reflecting formation changes and reservoir characteristics, thereby improving the accuracy, comprehensiveness and sensitivity of subsequent data analysis. In addition, since the dynamic time warping algorithm and the hierarchical clustering algorithm do not rely on linear assumptions or Gaussian distribution, they have better adaptability to nonlinear and non-Gaussian distributed geological data, so that complex geological data structures can be processed more flexibly, thereby improving the generalization ability of the dimensionality reduction effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of oil and gas exploration and development, and in particular to a dimensionality reduction method, device, storage medium and computer equipment for a post-casing electromagnetic logging curve. Background Art

[0002] In the modern petroleum industry, the post-casing electromagnetic logging technology is widely used in the fields of formation parameter identification and reservoir evaluation. Among them, the post-casing electromagnetic logging technology refers to the technology of exciting the electromagnetic field in a wellbore with casing installed (i.e., cased well), and receiving the electromagnetic response of the cased well to the electromagnetic field to obtain the post-casing electromagnetic logging curve. As an important geophysical data, the post-casing electromagnetic logging curve can provide electrical information about the underground reservoir, such as the resistivity parameters and conductivity parameters of the underground reservoir.

[0003] However, the complexity of field data acquisition and the high-dimensional characteristics of the post-casing electromagnetic logging curves have led to huge challenges in data processing and data interpretation. Therefore, the dimensionality reduction research of the post-casing electromagnetic logging curve is not only of great significance in theory, but also of great value in practical applications. By reducing the dimensionality of the post-casing electromagnetic logging curve, the redundancy of the data can be reduced, the influence of random noise can be reduced, the interpretability and visualization of the data can be improved, and the characteristic structure of the data can be effectively exposed, so that formation engineers and petroleum engineers can better understand the formation structure and reservoir characteristics based on the reduced dimensionality data. Furthermore, the dimensionality reduction of the post-casing electromagnetic logging curve can improve the efficiency and speed of data processing, accelerate the process of oilfield exploration and development, and is of great significance to improving the efficiency of oil and gas resource exploration and development. In addition, the dimensionality reduction of the post-casing electromagnetic logging curve can also help researchers discover the laws and characteristics hidden behind the data, providing more accurate and reliable data support for oilfield exploration and production.

[0004] It can be seen that the research on dimensionality reduction of post-casing electromagnetic logging curves has important theoretical significance and practical application value. It can provide a data basis for promoting the development of the oil industry and improving the efficiency of exploration and development, and is of great significance to promoting the innovation and development of geophysical exploration technology.

[0005] In recent years, various dimensionality reduction methods have been proposed in the prior art, including principal component analysis (PCA) and independent component analysis (ICA), aiming to extract the most critical and effective information from complex data. However, the inventors have found that if the existing dimensionality reduction methods are used to reduce the dimension of the post-casing electromagnetic logging curve, the data after dimensionality reduction will lose key information used to reflect formation changes and reservoir characteristics, making it difficult for engineers or researchers to accurately understand the formation structure and reservoir characteristics based on the data after dimensionality reduction, reducing the accuracy and comprehensiveness of subsequent analysis. Summary of the invention

[0006] The purpose of this application is to solve at least one of the above-mentioned technical defects, especially the technical defects of low dimensionality reduction accuracy and low comprehensiveness in the prior art.

[0007] In a first aspect, an embodiment of the present application provides a method for reducing the dimension of a post-casing electromagnetic logging curve, comprising:

[0008] Obtaining an N-dimensional electromagnetic logging curve after casing; wherein the N-dimensional electromagnetic logging curve after casing is an electromagnetic response curve obtained by performing electromagnetic logging on the same casing well;

[0009] Using a dynamic time warping algorithm, respectively calculating the first similarity distance between each two-dimensional set of electromagnetic logging curves in the N-dimensional set of electromagnetic logging curves;

[0010] When the preset clustering rules are met, the N-dimensional nested electromagnetic logging curves are clustered using a hierarchical clustering algorithm according to each of the first similarity distances, and an M-dimensional dimension reduction curve is obtained based on the clustering results; wherein M is less than N, and both M and N are positive integers.

[0011] In one embodiment, the dynamic time warping algorithm is used to calculate the first similarity distance between each two-dimensional nested electromagnetic logging curves in the N-dimensional nested electromagnetic logging curves, including:

[0012] respectively calculating the first-order derivative data and the second-order derivative data of the post-casing electromagnetic logging curve in each dimension;

[0013] According to the dynamic time warping algorithm and the N-dimensional post-casing electromagnetic logging curve, respectively calculating the first conventional similarity distance between each two-dimensional post-casing electromagnetic logging curve;

[0014] According to the dynamic time warping algorithm and the first derivative data of the post-casing electromagnetic logging curves in each dimension, respectively calculating the first first derivative similarity distance between the post-casing electromagnetic logging curves in each two dimensions;

[0015] According to the dynamic time warping algorithm and the second-order derivative data of the post-casing electromagnetic logging curves in each dimension, respectively calculating the first and second-order derivative similarity distances between the post-casing electromagnetic logging curves in each two dimensions;

[0016] Based on each of the first conventional similarity distances, each of the first first-order derivative similarity distances and each of the first second-order derivative similarity distances, the first similarity distance between each two-dimensional post-casing electromagnetic logging curve is obtained.

[0017] In one embodiment, clustering the N-dimensional nested electromagnetic logging curves using a hierarchical clustering algorithm according to each of the first similarity distances, and obtaining an M-dimensional dimension reduction curve based on the clustering result includes:

[0018] Determine each initial cluster of the current clustering round and each distance between initial clusters of the current clustering round; wherein, if the current clustering round is the first clustering round, each initial cluster of the current clustering round is the N-dimensional nested electromagnetic logging curve, and each distance between initial clusters of the current clustering round is each of the first similarity distances;

[0019] According to a predetermined density statistical threshold and each initial inter-cluster distance of the current clustering round, respectively counting the local density of each initial cluster in the current clustering round; wherein the local density of each initial cluster is the number of initial clusters whose initial inter-cluster distance with the initial cluster is less than the density statistical threshold;

[0020] Based on a preset screening threshold, determining a core cluster of the current clustering round from each initial cluster of the current clustering round;

[0021] Taking each core cluster of the current clustering round as the clustering center, clustering each initial cluster of the current clustering round using the hierarchical clustering algorithm until a preset clustering end condition is met and a clustering result of the current clustering round is obtained;

[0022] Determine whether to execute the next clustering round according to the clustering result of the current clustering round;

[0023] If necessary, the initial clusters of the next clustering round and the distances between the initial clusters of the next clustering round are updated according to the clustering results of the current clustering round, and the next clustering round is entered;

[0024] If not necessary, the M-dimensional dimensionality reduction curve is obtained.

[0025] In one of the embodiments, the clustering end condition is to merge the K initial clusters of the current clustering round into K / I cluster clusters, where I is a preset positive integer;

[0026] The step of judging whether to execute the next clustering round according to the clustering result of the current clustering round includes:

[0027] For each cluster obtained in the current clustering round, if the cluster includes at least two-dimensional post-casing electromagnetic logging curves, then the post-casing electromagnetic logging curves of each dimension included in the cluster are merged into a one-dimensional post-casing electromagnetic logging curve;

[0028] The dynamic time warping algorithm is used to calculate the second similarity distance between each two-dimensional electromagnetic logging curves in the K / I-dimensional electromagnetic logging curves;

[0029] If at least one of the second similarity distances is less than a preset similarity threshold, it is determined that the next clustering round needs to be executed; otherwise, it is determined that the next clustering round does not need to be executed.

[0030] In one embodiment, the updating of each initial cluster of the next clustering round and each initial inter-cluster distance of the next clustering round according to the clustering result of the current clustering round includes:

[0031] The K / I-dimensional electromagnetic logging curves are used as initial clusters for the next clustering round, and the second similarity distances are used as initial inter-cluster distances for the next clustering round.

[0032] In one embodiment, the step of merging the various dimensions of the post-casing electromagnetic logging curves included in the cluster into a one-dimensional post-casing electromagnetic logging curve comprises:

[0033] The mean value of each dimension of the post-casing electromagnetic logging curve included in the cluster is calculated to obtain a one-dimensional post-casing electromagnetic logging curve.

[0034] In one of the embodiments, the preset clustering rule is that at least one of the first similarity distances is smaller than a preset similarity threshold.

[0035] In a second aspect, an embodiment of the present application provides a dimension reduction device for a post-casing electromagnetic logging curve, comprising:

[0036] A curve acquisition module is used to acquire an N-dimensional electromagnetic logging curve after casing; wherein the N-dimensional electromagnetic logging curve after casing is an electromagnetic response curve obtained by performing electromagnetic logging on the same casing well;

[0037] A first similarity distance calculation module is used to calculate the first similarity distance between each two-dimensional set of electromagnetic logging curves in the N-dimensional set of electromagnetic logging curves by using a dynamic time warping algorithm;

[0038] A hierarchical clustering module is used to cluster the N-dimensional nested electromagnetic logging curves according to each of the first similarity distances using a hierarchical clustering algorithm when a preset clustering rule is met, and obtain an M-dimensional dimension reduction curve based on the clustering result; wherein M is less than N, and both M and N are positive integers.

[0039] In a third aspect, an embodiment of the present application provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the dimensionality reduction method of the post-casing electromagnetic logging curve described in any of the above embodiments.

[0040] In a fourth aspect, an embodiment of the present application provides a computer device, the computer device comprising: one or more processors, and a memory;

[0041] The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the dimension reduction method of the post-casing electromagnetic logging curve described in any of the above embodiments are executed.

[0042] In the dimension reduction method, device, storage medium and computer equipment of the post-casing electromagnetic logging curve provided in some embodiments of the present application, by adopting a dynamic regularization algorithm suitable for time series, the spatiotemporal information can be fully utilized and retained in the dimension reduction process, and the geological information contained in the post-casing electromagnetic logging curve can be better captured, so that the dimension reduction curve can better reflect the changes in the stratigraphic structure and the characteristics of the reservoir. By using a hierarchical clustering algorithm suitable for nonlinear clustering, the present application can accurately aggregate post-casing electromagnetic logging curves with similar characteristics into the same category and perform dimension reduction accordingly, so that the existence of redundant information can be reduced and noise data can be effectively filtered, thereby achieving de-redundancy and noise reduction, and improving the quality and credibility of the data.

[0043] It can be seen that the dimension reduction of the post-casing electromagnetic logging curve can be achieved by combining the dynamic time warping algorithm and the hierarchical clustering algorithm. The dimension reduction curve can retain the key information used to reflect the formation changes and reservoir characteristics, thereby improving the accuracy, comprehensiveness and sensitivity of subsequent data analysis. In addition, since the dynamic time warping algorithm and the hierarchical clustering algorithm do not rely on linear assumptions or Gaussian distribution, they have better adaptability to nonlinear and non-Gaussian distributed geological data, so that complex geological data structures can be processed more flexibly, thereby improving the generalization ability of the dimension reduction effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0045] Figure 1 One of the flow diagrams of a method for reducing the dimension of a post-casing electromagnetic logging curve in an embodiment;

[0046] Figure 2 A schematic diagram of the principle of using a dynamic time warping algorithm to calculate the first similarity distance between two-dimensional set electromagnetic logging curves in an embodiment;

[0047] Figure 3 A schematic flow chart of a step of respectively calculating the first similarity distance between each two-dimensional set of electromagnetic logging curves in an N-dimensional set of electromagnetic logging curves using a dynamic time warping algorithm in an embodiment;

[0048] Figure 4 In one embodiment, a schematic flow chart of the steps of clustering N-dimensional nested electromagnetic logging curves using a hierarchical clustering algorithm according to each first similarity distance, and obtaining an M-dimensional dimension reduction curve based on the clustering result is provided;

[0049] Figure 5 A schematic diagram of the first similarity distance between electromagnetic logging curves after each two-dimensional set in one embodiment;

[0050] Figure 6 For Figure 5 The schematic diagram of clustering results obtained by clustering the post-casing electromagnetic logging curves is shown;

[0051] Figure 7 FIG2 is a second flow chart of a method for reducing the dimension of a post-casing electromagnetic logging curve in an embodiment;

[0052] Figure 8 Schematic diagram of the structure of a dimension reduction device for a post-casing electromagnetic logging curve in one embodiment;

[0053] Fig. 9 FIG. 1 is a diagram of the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0054] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0055] Herein, a curve may be understood as a data set including multiple acquisition data, and a data set may include a limited number of acquisition data. For example, a post-casing electromagnetic logging curve may be a data set including multiple electromagnetic response data, and each electromagnetic response data may be a response time and an electromagnetic response value corresponding to the response time.

[0056] The inventors have found that due to the complexity and irregularity of the stratigraphic structure, the time series information of the post-casing electromagnetic logging curve often contains rich geological information, which can reflect the dynamic change characteristics of the geological structure. There is a spatiotemporal correlation between the time series information and the stratigraphic structure. The reason why the existing technology causes the data after dimensionality reduction to lose the key information used to reflect stratigraphic changes and reservoir characteristics is that the existing dimensionality reduction methods such as PCA and ICA, although they can achieve data dimensionality reduction to a certain extent, cannot effectively capture the dynamic change characteristics of the time series, and ignore the spatiotemporal correlation between the time series information and the stratigraphic structure, resulting in obvious deficiencies in data interpretation and stratigraphic feature extraction under complex geological conditions, causing the data after dimensionality reduction to lose the key information used to reflect stratigraphic changes and reservoir characteristics.

[0057] In addition to the above shortcomings, when the existing dimensionality reduction method is used to reduce the dimension of the post-casing electromagnetic logging curve, there are still the following shortcomings:

[0058] (1) For the post-casing electromagnetic logging curve, a single dimensionality reduction method is difficult to consider the redundant information phenomenon and noise influence between logging curves of different dimensions, and therefore cannot effectively deal with the noise influence and redundancy phenomenon of the post-casing electromagnetic logging curve.

[0059] (2) Existing dimensionality reduction methods have relatively limited adaptability to geological data with nonlinear and non-Gaussian distributions, and their generalization ability is insufficient. However, the post-casing electromagnetic logging curves often have complex nonlinear structures and non-Gaussian distribution characteristics, which results in the existing dimensionality reduction methods based on linear assumptions or Gaussian distribution having limited dimensionality reduction effects on the post-casing electromagnetic logging curves.

[0060] Based on this, the present application provides a method, device, storage medium and computer equipment for reducing the dimension of post-casing electromagnetic logging curves, and realizes the dimensionality reduction of post-casing electromagnetic logging curves by combining dynamic time warping algorithm and hierarchical clustering algorithm. The reduced-dimensional curve can retain key information for reflecting formation changes and reservoir characteristics, thereby improving the accuracy, comprehensiveness and sensitivity of subsequent data analysis. In addition, since the dynamic time warping algorithm and the hierarchical clustering algorithm do not rely on linear assumptions or Gaussian distribution, they have better adaptability to nonlinear and non-Gaussian distributed geological data, so that complex geological data structures can be processed more flexibly, thereby improving the generalization ability of the dimensionality reduction effect.

[0061] In one embodiment, the present application provides a dimensionality reduction method for post-casing electromagnetic logging curves. The following embodiments are described by taking the method applied to a computer device as an example. It can be understood that the computer device herein can be any device with data processing functions, which can be but is not limited to a desktop computer, a laptop computer, a server, etc.

[0062] like Figure 1As shown, the dimension reduction method of the post-casing electromagnetic logging curve of the present application includes the following steps:

[0063] S102: Obtaining N-dimensional electromagnetic logging curves.

[0064] Among them, the post-casing electromagnetic logging curve can be a response curve obtained by measuring the cased well using electromagnetic logging technology. In the electromagnetic logging process, the electromagnetic logging equipment can be used to excite the electromagnetic field at different depths of the cased well to obtain the electromagnetic response of the cased well. At each depth, the electromagnetic logging equipment can excite multiple electromagnetic fields with different voltage intensities / current intensities in the order of intensity from small to large or from large to small, so as to obtain the response of the same casing well to the electromagnetic field with different voltage intensities / current intensities at the same depth. In this article, the N-dimensional post-casing electromagnetic logging curve can be an electromagnetic response curve obtained by electromagnetic logging of the same casing well, and the dimension of each dimension post-casing electromagnetic logging curve is determined according to the order of voltage intensity / current intensity from large to small or from small to large. Each dimension post-casing electromagnetic logging curve is used to reflect the change of the electromagnetic response of the same casing well with depth under the same voltage intensity / current intensity.

[0065] For example, if during the electromagnetic logging process, the electromagnetic logging equipment excites 400 electromagnetic fields of different voltage intensities at each depth of the cased well in the order of intensity from large to small or from small to large, and receives the electromagnetic response of the cased well, then a 400-dimensional post-casing electromagnetic logging curve can be obtained after the electromagnetic logging, and the 400-dimensional post-casing electromagnetic logging curve corresponds to the electromagnetic fields of 400 voltage intensities one by one. Each dimension of the post-casing electromagnetic logging curve is the change of the electromagnetic response of the same cased well with depth under the same voltage intensity. At the same time, the dimension of the 400-dimensional post-casing electromagnetic logging curve is determined in the order of voltage intensity from large to small or from small to large. For example, the first dimension of the post-casing electromagnetic logging curve corresponds to the electromagnetic field with the maximum voltage intensity, the second dimension of the post-casing electromagnetic logging curve corresponds to the electromagnetic field with the second largest voltage intensity, and so on, the 400th dimension of the post-casing electromagnetic logging curve corresponds to the electromagnetic field with the minimum voltage intensity.

[0066] S104: using a dynamic time warping algorithm, respectively calculating the first similarity distance between each two-dimensional nested electromagnetic logging curves in the N-dimensional nested electromagnetic logging curves.

[0067] Among them, the Dynamic Time Warping (DTW) algorithm can be used to compare and align time series data. The DTW algorithm detects similar shapes with different stages by allowing "elastic" transformation of the time series, thereby minimizing the impact of time offset and distortion. Compared with the Euclidean distance, the DTW algorithm can calculate the distance between two different lengths of post-casing electromagnetic logging curves based on nonlinear distortion, thereby more accurately determining the similarity between two-dimensional post-casing electromagnetic logging curves.

[0068] Combination Figure 2 , assuming that the two-dimensional post-casing electromagnetic logging curves are P and Q:

[0069]

[0070]

[0071] Among them, p1 is the electromagnetic response of the one-dimensional post-casing electromagnetic logging curve at the first depth, p i is the electromagnetic response of one-dimensional post-casing electromagnetic logging curve at the i-th depth, is one of the one-dimensional electromagnetic logging curves at the L P The electromagnetic response at a depth of P is the depth length of P. q1 is the electromagnetic response of the electromagnetic logging curve after another dimension at the first depth, q j is the electromagnetic response of another dimensional set of electromagnetic logging curves at the jth depth, The electromagnetic logging curve after another dimension is Q The electromagnetic response at a depth of Q is the depth length of Q.

[0072] In this case, the cost matrix C of P and Q can be calculated as follows:

[0073] C i,j =d(p i ,q j )=|p i -q j |

[0074] Among them, the matrix size of C is L P ×L Q . C i,j is the element in the i-th row and j-th column of C, representing the Euclidean distance between the first i elements of P and the first j elements of Q.

[0075] The cumulative distance W between P and Q PQ It can be calculated by the following formula:

[0076]

[0077] in, is P at depth α l The electromagnetic response under is Q at depth β l The electromagnetic response under the condition of L P and L Q The maximum value in .

[0078] The goal of the DTW algorithm is to find the optimal matching path so that the cumulative distance W PQ Based on the above formula, the DTW algorithm can find the optimal matching path between two-dimensional electromagnetic logging curves according to three constraints: boundary constraint, continuity constraint and monotonicity constraint.

[0079] Among them, the boundary constraint condition restricts the matching path to start at (1,1) and end at (L P ,L Q ), that is, the matching path starts at d(p1,q1) and ends at d(L P ,L Q ). The continuity constraint restricts the depth-varying curved paths when aligning sequences. Under the continuity constraint, the matching path must move forward one step at a time. That is, α l+1 -α l ≤1 and β l+1 -β l ≤ 1. The monotonicity constraint is used to preserve the temporal order of the points, restricting the path to move forward and not decrease.

[0080] In this step, the similarity between each two-dimensional set of electromagnetic logging curves can be calculated by the DTW algorithm to obtain each first similarity distance. The smaller the first similarity distance, the more similar the data sequences are. Conversely, the larger the first similarity distance, the more dissimilar the data sequences are. In one example, after obtaining each first similarity distance, a similarity matrix including each first similarity distance can be constructed, and the redundancy of each feature sequence can be analyzed based on this.

[0081] For example, when N is 3, this step can respectively calculate the first similarity distance between the first-dimensional electromagnetic logging curve and the second-dimensional electromagnetic logging curve, the first similarity distance between the first-dimensional electromagnetic logging curve and the third-dimensional electromagnetic logging curve, and the first similarity distance between the second-dimensional electromagnetic logging curve and the third-dimensional electromagnetic logging curve.

[0082] S106: When the preset clustering rules are met, a hierarchical clustering algorithm is used to cluster the N-dimensional nested electromagnetic logging curves according to each first similarity distance, and an M-dimensional dimension reduction curve is obtained based on the clustering result.

[0083] Wherein, M is less than N, and both M and N are positive integers.

[0084] In this step, the preset clustering rule may be a rule for determining whether clustering needs to be performed, and the specific content of the rule may be determined according to the actual situation. In one example, the preset clustering rule may be that at least one first similarity distance is less than a preset similarity threshold. It can be understood that if there is at least one first similarity distance less than the preset similarity threshold, it indicates that among the N-dimensional electromagnetic logging curves, at least two-dimensional electromagnetic logging curves have a high data similarity, so the N-dimensional electromagnetic logging curves can be clustered. Furthermore, if each first similarity distance is greater than or equal to the preset similarity threshold, it indicates that the data similarity between the two N-dimensional electromagnetic logging curves is low, and the execution process can be directly exited.

[0085] In the process of clustering, this step can cluster the N-dimensional electromagnetic logging curves based on the hierarchical clustering algorithm on the basis of the DTW algorithm, so as to obtain clusters of each similar group, and then obtain the M-dimensional dimension reduction curve. By reducing the data dimension through the DTW algorithm and the hierarchical clustering algorithm, which are strong generalization feature engineering methods, redundancy removal and noise reduction can be achieved.

[0086] In this application, the dimension reduction of the post-casing electromagnetic logging curve is achieved by combining the dynamic time warping algorithm and the hierarchical clustering algorithm. The dimension reduction curve can retain key information for reflecting the formation changes and reservoir characteristics, thereby improving the accuracy, comprehensiveness and sensitivity of subsequent data analysis. In addition, since the dynamic time warping algorithm and the hierarchical clustering algorithm do not rely on linear assumptions or Gaussian distribution, they have better adaptability to nonlinear and non-Gaussian distributed geological data, so that complex geological data structures can be processed more flexibly, thereby improving the generalization ability of the dimension reduction effect.

[0087] In one embodiment, Figure 3 As shown, the dynamic time warping algorithm is used to calculate the first similarity distance between each two-dimensional set of electromagnetic logging curves in the N-dimensional set of electromagnetic logging curves, including:

[0088] S302: Calculate the first-order derivative data and the second-order derivative data of each dimension of the electromagnetic logging curve after casing respectively;

[0089] S304: calculating the first conventional similarity distance between each two-dimensional nested electromagnetic logging curves according to the dynamic time warping algorithm and the N-dimensional nested electromagnetic logging curves;

[0090] S306: Calculate the first derivative similarity distance between each two-dimensional electromagnetic logging curves according to the dynamic time warping algorithm and the first derivative data of each dimensional electromagnetic logging curve;

[0091] S308: Calculate the first and second derivative similarity distances between each two-dimensional set of electromagnetic logging curves according to the dynamic time warping algorithm and the second derivative data of each dimensional set of electromagnetic logging curves;

[0092] S310: Based on each first conventional similarity distance, each first first-order derivative similarity distance and each first second-order derivative similarity distance, obtain a first similarity distance between each two-dimensional set of electromagnetic logging curves.

[0093] Considering the similarity of the changing rules of electromagnetic response data in the casing electromagnetic logging curve, this embodiment can use the casing electromagnetic logging curve and its first-order derivative data and second-order derivative data to simultaneously perform sequence similarity analysis to improve the accuracy of the similarity analysis of the casing electromagnetic logging curve of a single well.

[0094] Specifically, in the process of determining each first similarity distance, the present application can perform first-order derivative and second-order derivative on each dimensional post-set electromagnetic logging curve respectively to obtain first-order derivative data and second-order derivative data corresponding to each dimensional post-set electromagnetic logging curve. Exemplarily, the present application can perform first-order derivative and second-order derivative with time as the independent variable.

[0095] Further, in one example, before calculating the first conventional similarity distance, the first first-order derivative similarity distance, and the first second-order derivative similarity distance, the present application may normalize the N-dimensional set electromagnetic logging curve, normalize the first-order derivative data corresponding to the N-dimensional set electromagnetic logging curve, and normalize the second-order derivative data corresponding to the N-dimensional set electromagnetic logging curve, and calculate each first conventional similarity distance, each first first-order derivative similarity distance, and each first second-order derivative similarity distance based on the normalized data. In this way, data regularization can be achieved, thereby eliminating the dimensional differences between features and reducing the scale differences between features.

[0096] For N-dimensional post-casing electromagnetic logging curves, the present application may use the DTW algorithm to calculate the similarity distance between each two-dimensional post-casing electromagnetic logging curves, and obtain each first conventional similarity distance. The present application may use the DTW algorithm to calculate the similarity distance between the first-order derivative data corresponding to each two-dimensional post-casing electromagnetic logging curves, and obtain each first first-order derivative similarity distance. The present application uses the DTW algorithm to calculate the similarity distance between the second-order derivative data corresponding to each two-bit post-casing electromagnetic logging curves, and obtain each first second-order derivative similarity distance.

[0097] It can be understood that in this article, the first conventional similarity distance is the similarity distance between the two-dimensional post-embedded electromagnetic logging curves, the first first-order derivative similarity distance is the similarity distance between the first-order derivative data corresponding to the two-dimensional post-embedded electromagnetic logging curves, and the first second-order derivative similarity distance is the similarity distance between the second-order derivative data corresponding to the two-dimensional post-embedded electromagnetic logging curves.

[0098] After obtaining the first conventional similarity distance, the first first-order derivative similarity distance, and the first second-order derivative similarity distance, the present application can determine the first similarity distance between each two-dimensional set electromagnetic logging curve based on the above three data. In one example, for each two-dimensional set electromagnetic logging curve, the present application can perform weighted summation on the first conventional similarity distance, the first first-order derivative similarity distance, and the first second-order derivative similarity distance of the two-dimensional set electromagnetic logging curve to obtain the first similarity distance corresponding to the two-dimensional set electromagnetic logging curve. It should be noted that the weight corresponding to the first conventional similarity distance, the weight corresponding to the first first-order derivative similarity distance, and the weight corresponding to the first second-order derivative similarity distance may be the same or different, and this document does not make specific restrictions on this. In some implementations, the three weights may be the same, that is, the first similarity distance is obtained by mean calculation.

[0099] In one embodiment, Figure 4 As shown, according to each first similarity distance, a hierarchical clustering algorithm is used to cluster the N-dimensional electromagnetic logging curves, and an M-dimensional dimension reduction curve is obtained based on the clustering results, including:

[0100] S402: Determine each initial cluster of the current clustering round and each distance between initial clusters of the current clustering round; wherein, if the current clustering round is the first clustering round, each initial cluster of the current clustering round is an N-dimensional nested electromagnetic logging curve, and each distance between initial clusters of the current clustering round is each first similarity distance;

[0101] S404: according to a predetermined density statistical threshold and each initial inter-cluster distance of the current clustering round, respectively counting the local density of each initial cluster in the current clustering round; wherein the local density of each initial cluster is the number of initial clusters whose initial inter-cluster distance with the initial cluster is less than the density statistical threshold;

[0102] S406: Determine a core cluster of the current clustering round from each initial cluster of the current clustering round based on a preset screening threshold;

[0103] S408: taking each core cluster of the current clustering round as the clustering center, clustering each initial cluster of the current clustering round using a hierarchical clustering algorithm until a preset clustering end condition is met and a clustering result of the current clustering round is obtained;

[0104] S410: judging whether it is necessary to execute the next clustering round according to the clustering result of the current clustering round;

[0105] S412: if necessary, updating each initial cluster of the next clustering round and each initial inter-cluster distance of the next clustering round respectively according to the clustering result of the current clustering round, and entering the next clustering round;

[0106] S414: If not necessary, obtain an M-dimensional dimensionality reduction curve.

[0107] In this embodiment, by adopting a density-based hierarchical clustering algorithm, the order relationship and spatial relationship of the post-casing electromagnetic logging curves can be considered in the clustering process, which can not only generate a tree structure to clearly show the similarity relationship between the post-casing electromagnetic logging curves, but also use the dimensional order of the post-casing electromagnetic logging curves to consider the distance similarity of the electromagnetic response. In this way, the accuracy of clustering can be improved to further improve the accuracy and comprehensiveness of the cooling curve. For example, according to Figure 5 After clustering with the first similarity matrix and the density-based hierarchical clustering algorithm, we can get Figure 6 The clustering results are shown.

[0108] Among them, the density-based hierarchical clustering algorithm can combine the density clustering algorithm with the hierarchical clustering algorithm, and can gradually merge or split the N-dimensional set of electromagnetic logging curves through relative distance to build a hierarchical structure, which has advantages in processing complex data sets and discovering multi-level structures. Specifically, the present application can complete clustering through multiple rounds of iterations, and steps S402 to S414 can be executed in each clustering round to complete a round of clustering.

[0109] In the current clustering round, each initial cluster of the current clustering round and the inter-cluster distance between each initial cluster (i.e., each initial inter-cluster distance) can be determined. It can be understood that if the current clustering round is the first clustering round, the N-dimensional nested electromagnetic logging curves can be used as N initial clusters, and each first similarity distance is used as each initial inter-cluster distance.

[0110] When determining the initial clusters and the distances between the initial clusters in the current clustering round, the predetermined density statistics threshold can be used as a given radius, and the number of neighboring clusters corresponding to each initial cluster falling within the given radius can be counted to obtain the local density of each initial cluster. For example, when the density statistics threshold is 10, for initial cluster A, the number of initial clusters whose first similarity distance to initial cluster A is less than 10 can be counted, and this can be used as the local density of initial cluster A. Similarly, the local densities of the remaining initial clusters can be determined according to the above steps.

[0111] In the current clustering round, after determining the local density corresponding to each initial cluster, the density reachability can be judged based on this. That is, the local density corresponding to each initial cluster is compared with the preset screening threshold, and based on the size comparison result, the core cluster of the current clustering round is determined in each initial cluster. It can be understood that the screening rules of the core cluster can be determined according to the actual situation, and the specific value of the preset threshold can also be determined according to the actual situation. This article does not make specific restrictions on this. In one example, the initial cluster whose local density is greater than the preset screening threshold can be used as the core cluster. In another example, the preset screening threshold may be 10.

[0112] In this article, the core cluster can be an initial cluster with a large local density. When the local density of the initial cluster is large, it indicates that the initial cluster is relatively similar to the remaining initial clusters with a large number. Therefore, each core cluster of the current clustering round can be used as the clustering center, and the hierarchical clustering algorithm can be used to merge the initial clusters of the current clustering round until the preset clustering end condition is met. In the clustering process of the current clustering round, based on each first similarity distance, a single link, a complete link or an average link can be used to merge the two clusters with the closest inter-cluster distance into one cluster, and the inter-cluster distances of the current clustering round are updated based on the merged result.

[0113] In one example, the present application may merge two clusters with the smallest average link distance, so that the merged cluster can maintain the average similarity between internal samples as much as possible. Further, the average link distance D(A, B) between clusters A and B may be calculated based on the following expression:

[0114]

[0115] Where w(i,j) is the initial inter-cluster distance between the i-th dimension post-casing electromagnetic logging curve of cluster A and the j-th dimension post-casing electromagnetic logging curve of cluster B, |A| is the number of post-casing electromagnetic logging curves included in cluster A, and |B| is the number of post-casing electromagnetic logging curves included in cluster B.

[0116] When the clustering end condition is met, it can be determined that the clustering of the current clustering round is completed. In this case, it can be determined whether the next clustering round needs to be executed based on the clustering result of the current clustering round. If the next clustering round needs to be executed, the initial clusters of the next clustering round and the distances between the initial clusters of the next clustering round can be updated, and the next clustering round can be entered. Otherwise, an M-dimensional dimensionality reduction curve can be obtained.

[0117] In one embodiment, the clustering end condition is to merge the K initial clusters of the current clustering round into K / I cluster clusters, where K is the number of initial clusters of the current clustering round, and I is a preset positive integer. For example, when I=2 and N=400, the first clustering round can cluster the 400-dimensional nested electromagnetic curves into 200 clusters, and obtain 200-dimensional nested electromagnetic curves based on the clustering results. The second clustering round can cluster the 200-dimensional nested electromagnetic curves into 100 clusters, and obtain 100-dimensional nested electromagnetic curves based on the clustering results, and so on, until the next clustering round does not need to be executed.

[0118] In this embodiment, judging whether to execute the next clustering round according to the clustering result of the current clustering round may include the following steps:

[0119] Step A1: for each cluster obtained in the current clustering round, if the cluster includes at least two-dimensional nested electromagnetic logging curves, then the nested electromagnetic logging curves of each dimension included in the cluster are merged into a one-dimensional nested electromagnetic logging curve;

[0120] Step A3: using a dynamic time warping algorithm, respectively calculating the second similarity distance between every two-dimensional electromagnetic logging curves in the K / I-dimensional electromagnetic logging curves;

[0121] Step A5: If at least one second similarity distance is smaller than a preset similarity threshold, it is determined that the next clustering round needs to be executed; otherwise, it is determined that the next clustering round does not need to be executed.

[0122] For ease of description, this embodiment uses the clustering round of clustering 400-dimensional electromagnetic logging curves into 200 clusters as an example for explanation. It can be understood that the other clustering rounds can also be understood in this way, and the specific values ​​of 400, 200, 100, etc. are only for explaining the judgment process and should not constitute a limitation to this application.

[0123] Specifically, when clustering 400-dimensional nested electromagnetic logging curves into 200 clusters, at least some of the clusters include multi-dimensional nested electromagnetic logging curves. For clusters including multi-dimensional nested electromagnetic logging curves, the present application can merge the multi-dimensional nested electromagnetic logging curves in the same cluster into a one-dimensional nested electromagnetic logging curve. In this way, the 400-dimensional nested electromagnetic logging curve can be reduced in dimension to a 200-dimensional nested electromagnetic logging curve based on the clustering results.

[0124] In the case of obtaining a 200-dimensional set of electromagnetic logging curves, the present application may use a dynamic time warping algorithm to calculate the second similarity distance between each two-dimensional set of electromagnetic logging curves. Among them, the relevant description and acquisition method of the second similarity distance can refer to the description of the first similarity distance, which will not be repeated here. For example, the present application may first calculate the second conventional similarity distance between the 200-dimensional set of electromagnetic logging curves, and calculate each second first-order derivative similarity distance and each second second-order derivative similarity distance based on the first-order derivative data and second-order derivative data of the 200-dimensional set of electromagnetic logging curves, and obtain each second similarity distance based on each second conventional similarity distance, each second first-order derivative similarity distance and each second second-order derivative similarity distance.

[0125] If at least one of the second similarity distances corresponding to the current clustering round is less than or equal to the preset similarity threshold, it can be determined that the next clustering round is needed to cluster the 200-dimensional electromagnetic logging curves into 100 clusters, and so on.

[0126] If each second similarity distance corresponding to the current clustering round is greater than or equal to the preset similarity threshold, it can be determined that the clustering is finished and there is no need to execute the next clustering round.

[0127] In this embodiment, after each clustering round, dimensionality reduction can be performed based on the clustering results of the current clustering round, and the similarity distance between the dimensionality reduction curves can be updated to determine whether to execute the next clustering round. In this way, the accuracy of the clustering results can be further improved, so as to further improve the dimensionality reduction accuracy of the post-casing electromagnetic logging curve.

[0128] In one embodiment, merging the dimensional post-embedded electromagnetic logging curves included in the cluster into a one-dimensional post-embedded electromagnetic logging curve includes: performing mean calculation on the dimensional post-embedded electromagnetic logging curves included in the cluster to obtain the one-dimensional post-embedded electromagnetic logging curve. In this way, the time series noise in the post-embedded electromagnetic logging curve can be smoothed and feature dimensionality reduction can be achieved.

[0129] In this embodiment, each dimension of the post-casing electromagnetic logging curve may include the electromagnetic response of the cased well at different depths. Therefore, when performing the mean calculation, the present application may average the electromagnetic response at each depth of each dimension of the post-casing electromagnetic logging curve to merge the multi-dimensional post-casing electromagnetic logging curve into a one-dimensional curve.

[0130] For example, a cluster includes a first set of post-electromagnetic logging curves and a second set of post-electromagnetic logging curves. The first set of post-electromagnetic logging curves includes (D1, R1)(D2, R2), ... (Dn, Rn). The second set of post-electromagnetic logging curves includes (D1, E1)(D2, E2), ... (Dn, En). Among them, D1, D2, Dn are different depths, and R1, R2, Rn, E1, E2, En are corresponding electromagnetic responses. In this case, the average value A1 of R1 and E1, the average value A2 of R2 and E2, and so on can be calculated until the average value An of Rn and En is calculated. In this way, a merged one-dimensional set of post-electromagnetic logging curves can be obtained, which includes (D1, A1)(D2, A2), ... (Dn, An).

[0131] In one embodiment, each initial cluster of the next clustering round and each initial inter-cluster distance of the next clustering round are updated respectively according to the clustering results of the current clustering round, including: using the electromagnetic logging curve after K / I dimension as each initial cluster of the next clustering round, and using each second similarity distance as each initial inter-cluster distance of the next clustering round.

[0132] For example, after merging 400-dimensional electromagnetic logging curves into 200 clusters and reducing the dimension to 200-dimensional electromagnetic logging curves, if it is determined that the next clustering round needs to be performed, the 200-dimensional electromagnetic logging curves obtained by dimensionality reduction can be used as the 200 initial clusters for the next clustering round, and the second similarity distances between the 200-dimensional electromagnetic logging curves are used as the distances between the initial clusters for the next clustering round, and the next clustering round is entered to reduce the dimension of the 200-dimensional electromagnetic logging curves to 100 dimensions.

[0133] To facilitate understanding of this application, a specific example is provided below. Figure 7 As shown, the dimension reduction method of the post-casing electromagnetic logging curve provided in this application may include the following steps:

[0134] S502: Obtaining an N-dimensional post-casing electromagnetic logging curve of the same casing well;

[0135] S504: normalizing the N-dimensional electromagnetic logging curves, and using the normalized results as N independent time series;

[0136] S506: respectively obtain the first-order derivative data of each dimension of the electromagnetic logging curve after being covered, and perform normalization processing on the first-order derivative data corresponding to the N-dimensional electromagnetic logging curve after being covered;

[0137] S508: respectively obtain the second-order derivative data of each dimension of the electromagnetic logging curve after being covered, and perform normalization processing on the second-order derivative data corresponding to the N-dimensional electromagnetic logging curve after being covered;

[0138] S510: Calculate the similarity distances between the N-dimensional electromagnetic logging curves by using a DWT algorithm to obtain a first similarity matrix;

[0139] S512: Calculate the similarity distance between the first-order derivative data of each two-dimensional set of electromagnetic logging curves by using a DWT algorithm to obtain a second similarity matrix;

[0140] S514: calculating the similarity distance between the second-order derivative data of each two-dimensional set of electromagnetic logging curves by using a DWT algorithm to obtain a third similarity matrix;

[0141] S516: averaging the first similarity matrix, the second similarity matrix, and the third similarity matrix to obtain a mean similarity matrix;

[0142] S518: Determine whether there is a value less than a preset similarity threshold in the mean similarity matrix, if so, proceed to S520, otherwise, proceed to S528;

[0143] S520: taking the post-casing electromagnetic logging curve of each dimension as an initial cluster, and calculating the local density of each initial cluster;

[0144] S522: determining a core cluster according to the local density of each initial cluster;

[0145] S524: Based on the hierarchical clustering algorithm, each core cluster is used as the clustering center to merge similar clusters; specifically, the present application can merge similar clusters according to the average link formula, and use 1 / 2 of the initial number of clusters as the number of clusters in each clustering round;

[0146] S526: averaging within the cluster, averaging the values ​​of the multi-dimensional electromagnetic logging curves at each depth in the same cluster, achieving dimensionality reduction, and entering S512; in this way, random noise and other noises can be smoothed, and feature dimensionality reduction is completed, and the number of features is reduced to the number of clusters of the same type;

[0147] S528: Output an M-dimensional dimensionality reduction curve.

[0148] It should be understood that although Figure 1-Figure 7 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1-Figure 7At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0149] The following describes a dimension reduction device for a post-casing electromagnetic logging curve provided in an embodiment of the present application. The dimension reduction device for a post-casing electromagnetic logging curve described below and the dimension reduction method for a post-casing electromagnetic logging curve described above can refer to each other.

[0150] In one embodiment, Figure 8 As shown, the present application provides a dimension reduction device 800 for a post-casing electromagnetic logging curve, specifically comprising:

[0151] The curve acquisition module 810 is used to acquire an N-dimensional electromagnetic logging curve after casing; wherein the N-dimensional electromagnetic logging curve after casing is an electromagnetic response curve obtained by performing electromagnetic logging on the same casing well;

[0152] A first similarity distance calculation module 820 is used to calculate the first similarity distance between each two-dimensional nested electromagnetic logging curves in the N-dimensional nested electromagnetic logging curves by using a dynamic time warping algorithm;

[0153] The hierarchical clustering module 830 is used to cluster the N-dimensional nested electromagnetic logging curves according to each of the first similarity distances using a hierarchical clustering algorithm when the preset clustering rules are met, and obtain an M-dimensional dimension reduction curve based on the clustering results; wherein M is less than N, and both M and N are positive integers.

[0154] In one embodiment, the first similarity distance calculation module 820 of the present application includes:

[0155] A derivative unit, used for respectively calculating the first-order derivative data and the second-order derivative data of the post-casing electromagnetic logging curve in each dimension;

[0156] A first conventional similarity calculation unit is used to calculate the first conventional similarity distance between each two-dimensional post-casing electromagnetic logging curves according to the dynamic time warping algorithm and the N-dimensional post-casing electromagnetic logging curves;

[0157] A first first-order derivative similarity calculation unit is used to calculate the first first-order derivative similarity distance between each two-dimensional post-casing electromagnetic logging curves according to the dynamic time warping algorithm and the first-order derivative data of each dimension of the post-casing electromagnetic logging curve;

[0158] A first second-order derivative similarity calculation unit is used to calculate the first second-order derivative similarity distance between each two-dimensional post-casing electromagnetic logging curves according to the dynamic time warping algorithm and the second-order derivative data of each dimension of the post-casing electromagnetic logging curve;

[0159] The similarity distance calculation unit is used to obtain the first similarity distance between each two-dimensional post-casing electromagnetic logging curve based on each of the first conventional similarity distances, each of the first first-order derivative similarity distances and each of the first second-order derivative similarity distances.

[0160] In one embodiment, the hierarchical clustering module 830 of the present application includes:

[0161] An initial cluster and inter-cluster distance determination unit, used to determine each initial cluster of the current clustering round and each initial inter-cluster distance of the current clustering round; wherein, if the current clustering round is the first clustering round, each initial cluster of the current clustering round is the N-dimensional nested electromagnetic logging curve, and each initial inter-cluster distance of the current clustering round is each of the first similarity distances;

[0162] A local density statistics unit is used to count the local density of each initial cluster in the current clustering round according to a predetermined density statistics threshold and each initial inter-cluster distance of the current clustering round; wherein the local density of each initial cluster is the number of initial clusters whose initial inter-cluster distance with the initial cluster is less than the density statistics threshold;

[0163] A core cluster determination unit, configured to determine a core cluster of the current clustering round from among the initial clusters of the current clustering round based on a preset screening threshold;

[0164] A clustering unit, used to cluster each initial cluster of the current clustering round using the hierarchical clustering algorithm, taking each core cluster of the current clustering round as a clustering center, until a preset clustering end condition is met and a clustering result of the current clustering round is obtained;

[0165] A judging unit, used to judge whether the next clustering round needs to be executed according to the clustering result of the current clustering round;

[0166] An updating unit, used for updating each initial cluster of the next clustering round and each initial inter-cluster distance of the next clustering round according to the clustering result of the current clustering round when the next clustering round needs to be executed, and entering the next clustering round;

[0167] The dimension reduction curve acquisition unit is used to obtain the M-dimensional dimension reduction curve when there is no need to execute the next clustering round.

[0168] In one embodiment, the clustering end condition is to merge the K initial clusters of the current clustering round into K / I cluster clusters, where I is a preset positive integer. The judgment unit of the present application includes:

[0169] A merging unit, for merging the post-casing electromagnetic logging curves of each dimension included in each cluster cluster obtained in the current clustering round into a one-dimensional post-casing electromagnetic logging curve if the cluster cluster includes at least two-dimensional post-casing electromagnetic logging curves;

[0170] A second similarity distance calculation unit is used to calculate the second similarity distance between every two-dimensional electromagnetic logging curves in the K / I-dimensional electromagnetic logging curves by using the dynamic time warping algorithm;

[0171] The similarity comparison unit is used to determine that the next clustering round needs to be executed if at least one of the second similarity distances is less than a preset similarity threshold, and otherwise, determine that the next clustering round does not need to be executed.

[0172] In one embodiment, the updating unit of the present application includes:

[0173] The initial cluster and inter-cluster distance updating unit is used to use the K / I dimension-fitted electromagnetic logging curve as each initial cluster of the next clustering round, and use each of the second similarity distances as each initial inter-cluster distance of the next clustering round.

[0174] In one embodiment, the merging unit of the present application includes:

[0175] The mean calculation unit is used to perform mean calculation on the post-casing electromagnetic logging curves of each dimension included in the cluster to obtain a one-dimensional post-casing electromagnetic logging curve.

[0176] In one embodiment, the preset clustering rule is that at least one of the first similarity distances is smaller than a preset similarity threshold.

[0177] In one embodiment, the present application further provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the dimension reduction method of the post-casing electromagnetic logging curve in any embodiment.

[0178] In one embodiment, the present application also provides a computer device having computer-readable instructions stored therein. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the dimensionality reduction method of the post-casing electromagnetic logging curve in any embodiment.

[0179] Indicatively, Fig. 9The internal structure diagram of a computer device provided in an embodiment of the present application is shown in FIG. 1 . In one example, the computer device may be a server. Fig. 9 The computer device 900 includes a processing component 902, which further includes one or more processors, and a memory resource represented by a memory 901, for storing instructions executable by the processing component 902, such as an application. The application stored in the memory 901 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 902 is configured to execute instructions to perform the steps of the dimensionality reduction method of the post-casing electromagnetic logging curve described in any of the above embodiments.

[0180] The computer device 900 may further include a power supply component 903 configured to perform power management of the computer device 900, a wired or wireless network interface 904 configured to connect the computer device 900 to a network, and an input / output (I / O) interface 905. The computer device 900 may operate based on an operating system stored in the memory 901, such as Windows Server TM, Mac OS X TM, Unix TM, Linux TM, Free BSD TM or the like.

[0181] Those skilled in the art will understand that the internal structure of the computer device shown in the present application is merely a block diagram of a partial structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0182] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the term "include", "comprise" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not clearly listed, or also includes elements inherent to such process, method, article or equipment. In the absence of more restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or equipment including the elements. Herein, "one", "one", "said", "the" and "it" may also include plural forms, unless the context clearly indicates another way. A plurality refers to at least two cases, such as 2, 3, 5 or 8, etc. "And / or" includes any and all combinations of the relevant listed items.

[0183] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.

[0184] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A dimensionality reduction method for post-casing electromagnetic logging curves, characterized in that: include: Obtaining an N-dimensional electromagnetic logging curve after casing; wherein the N-dimensional electromagnetic logging curve after casing is an electromagnetic response curve obtained by performing electromagnetic logging on the same casing well; Using a dynamic time warping algorithm, respectively calculating the first similarity distance between each two-dimensional set of electromagnetic logging curves in the N-dimensional set of electromagnetic logging curves; In the case where the preset clustering rules are met, the N-dimensional nested electromagnetic logging curves are clustered using a hierarchical clustering algorithm according to each of the first similarity distances, and an M-dimensional dimension reduction curve is obtained based on the clustering result; wherein M is less than N, and both M and N are positive integers; The method of using a dynamic time warping algorithm to respectively calculate the first similarity distance between each two-dimensional set of electromagnetic logging curves in the N-dimensional set of electromagnetic logging curves includes: respectively calculating the first-order derivative data and the second-order derivative data of the post-casing electromagnetic logging curve in each dimension; According to the dynamic time warping algorithm and the N-dimensional post-casing electromagnetic logging curve, respectively calculating the first conventional similarity distance between each two-dimensional post-casing electromagnetic logging curve; According to the dynamic time warping algorithm and the first derivative data of the post-casing electromagnetic logging curves in each dimension, respectively calculating the first first derivative similarity distance between the post-casing electromagnetic logging curves in each two dimensions; According to the dynamic time warping algorithm and the second-order derivative data of the post-casing electromagnetic logging curves in each dimension, respectively calculating the first and second-order derivative similarity distances between the post-casing electromagnetic logging curves in each two dimensions; Based on each of the first conventional similarity distances, each of the first first-order derivative similarity distances and each of the first second-order derivative similarity distances, the first similarity distance between each two-dimensional post-casing electromagnetic logging curve is obtained.

2. The method according to claim 1, characterized in that The step of clustering the N-dimensional nested electromagnetic logging curves using a hierarchical clustering algorithm according to each of the first similarity distances, and obtaining an M-dimensional dimension reduction curve based on the clustering result includes: Determine each initial cluster of the current clustering round and each distance between initial clusters of the current clustering round; wherein, if the current clustering round is the first clustering round, each initial cluster of the current clustering round is the N-dimensional nested electromagnetic logging curve, and each distance between initial clusters of the current clustering round is each of the first similarity distances; According to a predetermined density statistical threshold and each initial inter-cluster distance of the current clustering round, respectively counting the local density of each initial cluster in the current clustering round; wherein the local density of each initial cluster is the number of initial clusters whose initial inter-cluster distance with the initial cluster is less than the density statistical threshold; Based on a preset screening threshold, determining a core cluster of the current clustering round from each initial cluster of the current clustering round; Taking each core cluster of the current clustering round as the clustering center, clustering each initial cluster of the current clustering round using the hierarchical clustering algorithm until a preset clustering end condition is met and a clustering result of the current clustering round is obtained; Determine whether to execute the next clustering round according to the clustering result of the current clustering round; If necessary, the initial clusters of the next clustering round and the distances between the initial clusters of the next clustering round are updated according to the clustering results of the current clustering round, and the next clustering round is entered; If not necessary, the M-dimensional dimensionality reduction curve is obtained.

3. The method according to claim 2, characterized in that The clustering end condition is to merge the K initial clusters of the current clustering round into K / I cluster clusters, where I is a preset positive integer; The step of judging whether to execute the next clustering round according to the clustering result of the current clustering round includes: For each cluster obtained in the current clustering round, if the cluster includes at least two-dimensional post-casing electromagnetic logging curves, then the post-casing electromagnetic logging curves of each dimension included in the cluster are merged into a one-dimensional post-casing electromagnetic logging curve; The dynamic time warping algorithm is used to calculate the second similarity distance between each two-dimensional electromagnetic logging curves in the K / I-dimensional electromagnetic logging curves; If at least one of the second similarity distances is less than a preset similarity threshold, it is determined that the next clustering round needs to be executed; otherwise, it is determined that the next clustering round does not need to be executed.

4. The method according to claim 3, characterized in that The updating of each initial cluster of the next clustering round and each initial inter-cluster distance of the next clustering round according to the clustering result of the current clustering round includes: The K / I-dimensional electromagnetic logging curves are used as initial clusters for the next clustering round, and the second similarity distances are used as initial inter-cluster distances for the next clustering round.

5. The method according to claim 3 or 4, characterized in that: The step of merging the various dimensions of the post-casing electromagnetic logging curves included in the cluster into a one-dimensional post-casing electromagnetic logging curve comprises: The mean value of each dimension of the post-casing electromagnetic logging curve included in the cluster is calculated to obtain a one-dimensional post-casing electromagnetic logging curve.

6. The method according to any one of claims 1 to 4, characterized in that: The preset clustering rule is that at least one of the first similarity distances is smaller than a preset similarity threshold.

7. A dimension reduction device for post-casing electromagnetic logging curves, characterized in that: include: A curve acquisition module is used to acquire an N-dimensional post-casing electromagnetic logging curve; wherein the N-dimensional post-casing electromagnetic logging curve is an electromagnetic response curve obtained by performing electromagnetic logging on the same casing well; A first similarity distance calculation module is used to calculate the first similarity distance between each two-dimensional set of electromagnetic logging curves in the N-dimensional set of electromagnetic logging curves by using a dynamic time warping algorithm; A hierarchical clustering module, configured to cluster the N-dimensional nested electromagnetic logging curves using a hierarchical clustering algorithm according to each of the first similarity distances when a preset clustering rule is satisfied, and obtain an M-dimensional dimension reduction curve based on the clustering result; wherein M is less than N, and both M and N are positive integers; Wherein, the first similarity distance calculation module includes: A derivative unit, used for respectively calculating the first-order derivative data and the second-order derivative data of the post-casing electromagnetic logging curve in each dimension; A first conventional similarity calculation unit is used to calculate the first conventional similarity distance between each two-dimensional post-casing electromagnetic logging curves according to the dynamic time warping algorithm and the N-dimensional post-casing electromagnetic logging curves; A first first-order derivative similarity calculation unit is used to calculate the first first-order derivative similarity distance between each two-dimensional post-casing electromagnetic logging curves according to the dynamic time warping algorithm and the first-order derivative data of each dimension of the post-casing electromagnetic logging curve; A first second-order derivative similarity calculation unit is used to calculate the first second-order derivative similarity distance between each two-dimensional post-casing electromagnetic logging curves according to the dynamic time warping algorithm and the second-order derivative data of each dimension of the post-casing electromagnetic logging curve; The similarity distance calculation unit is used to obtain the first similarity distance between each two-dimensional post-casing electromagnetic logging curve based on each of the first conventional similarity distances, each of the first first-order derivative similarity distances and each of the first second-order derivative similarity distances.

8. A storage medium, characterized in that: The storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the dimension reduction method of the post-casing electromagnetic logging curve as described in any one of claims 1 to 6.

9. A computer device, characterized in that: include: one or more processors, and memory; The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the dimension reduction method of the post-casing electromagnetic logging curve as claimed in any one of claims 1 to 6 are performed.