A method and device for calculating similarity of tight gas wells
By performing importance fusion calculation on the multi-dimensional feature data of tight gas wells, the problem of accuracy in similarity judgment of tight gas wells is solved, and the scientificity and accuracy of production law analysis are improved.
Patent Information
- Application Number
- CN202310389751.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-12
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2043-04-12
AI Technical Summary
Existing technologies make it difficult to accurately calculate the similarity of different tight gas wells, making it difficult to analyze differences in production characteristics.
By obtaining the feature data of N wells to be calculated, dividing them into multiple types, calculating the importance of each type of feature, and performing importance fusion, the similarity of the wells is determined.
It achieves more accurate judgment of the similarity of tight gas wells and improves the scientificity and accuracy of production law analysis.
Smart Images

Figure CN116401558B_ABST
Abstract
Description
Technical Field
[0001] This article relates to the field of well logging technology, and in particular to a method and device for calculating similarity of tight gas wells. Background Art
[0002] With the advancement of digitalization in the oil and gas industry, more and more applications are moving towards information technology and intelligent systems. However, due to limitations such as geological conditions, the production characteristics of different tight gas wells vary significantly. In practical scenarios, it is often necessary to combine and analyze similar wells to identify corresponding production patterns. Therefore, the exploration of similarity calculation methods for tight gas wells has become a necessary issue. Summary of the Invention
[0003] The present application provides a method and device for calculating the similarity of tight gas wells, which integrates the multi-dimensional characteristics of the wells, comprehensively calculates and evaluates the similarity of the wells, and makes the similarity judgment more accurate.
[0004] The present application provides a method for calculating similarity of tight gas wells, the method comprising:
[0005] Acquire characteristic data of N wells to be calculated, and classify the characteristic data into multiple types, where N is a positive integer greater than or equal to 2;
[0006] Calculate the importance of each type of feature of the first well and the second well respectively according to the feature data;
[0007] An importance fusion calculation is performed based on the importance of multiple types of features to determine the similarity between the first well and the second well.
[0008] In an exemplary embodiment, the feature data is divided into multiple types, including:
[0009] The feature data are divided into three types.
[0010] In an exemplary embodiment, calculating the importance of each type of feature of the first well and the second well respectively based on the feature data includes:
[0011] determining a total value of the same feature in the first type based on the first type feature data of the first well and the first type feature data of the second well;
[0012] The determined total value is used as the importance of the first type of feature of the first well and the second well.
[0013] In an exemplary embodiment, calculating the importance of each type of feature of the first well and the second well respectively based on the feature data includes:
[0014] A string fuzzy matching method is used to extract the production layer and the target layer from the second type feature data of the first well and the second type feature data of the second well;
[0015] Calculating the Jaccard similarity coefficient between the production layer and the target layer according to a similarity coefficient formula;
[0016] The jaccard similarity coefficient is used as the importance of the second type of feature between the first well and the second well;
[0017] Wherein, the similarity coefficient formula is:
[0018]
[0019] In an exemplary embodiment, calculating the importance of each type of feature of the first well and the second well respectively based on the feature data includes:
[0020] The 3σ principle was used to filter out the outliers in the third type characteristic data of the first well and the third type characteristic data of the second well respectively;
[0021] Normalize the filtered feature data;
[0022] Calculating the Euclidean distance between the normalized feature data of the first well and the normalized feature data of the second well;
[0023] The obtained Euclidean distance is used as the importance of the third type feature between the first well and the second well.
[0024] In an exemplary embodiment, performing importance fusion calculation based on the importance of multiple types of features to determine the similarity between the first well and the second well includes:
[0025] A weighted average calculation is performed based on the determined importance of each type of feature and a predetermined weight coefficient corresponding to each type of feature to determine the similarity between the first well and the second well.
[0026] In an exemplary embodiment, the second well is any well other than the first well among the N wells;
[0027] After performing importance fusion calculation based on the importance of multiple types of features and determining the similarity between the first well and the second well, the method further includes:
[0028] Arrange the obtained N-1 similarities in descending order;
[0029] Perform similarity correction on the first M similarities in descending order, where M is a positive integer less than N-1.
[0030] In an exemplary embodiment, the similarity correction of the first M similarities in the descending order includes:
[0031] Determine the total number of features for N wells;
[0032] Determining the maximum number of common features between the first well and the second well corresponding to the M similarities;
[0033] According to the total number of features and the number of the largest common features, a correction coefficient formula is used to determine the corrected similarity; wherein the correction coefficient formula is:
[0034]
[0035] In the above formula, a is the total number of features of each type in N wells, and b is the number of the maximum common features between the first well and the second well corresponding to M similarities.
[0036] The present application also provides a device for calculating the similarity of tight gas wells, which includes: a memory and a processor; the memory is used to store a program for calculating the similarity of tight gas wells, and the processor is used to read and execute the program for calculating the similarity of tight gas wells, and perform any method described in the above embodiments.
[0037] The present application also provides a computer storage medium, in which a program for calculating the similarity of tight gas wells is stored. The program is configured to execute the method for calculating the similarity of tight gas wells described in any one of the above embodiments when running.
[0038] Compared to related technologies, this application provides a method and apparatus for calculating the similarity of tight gas wells. The method includes: obtaining feature data for N wells to be calculated and classifying them into multiple categories based on their data features; calculating the importance of each feature category for a first well and a second well based on the well feature data; and performing importance fusion based on the importance of each feature category to determine the similarity between the first and second wells. The technical solution disclosed in this application integrates the multi-dimensional features of the wells to comprehensively calculate and evaluate the similarity of the wells, making similarity determination more accurate.
[0039] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. Other advantages of the present application can be realized and obtained by the solutions described in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The accompanying drawings are used to provide an understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.
[0041] Figure 1 This is a flow chart of a method for calculating the similarity of tight gas wells according to an embodiment of the present application;
[0042] Figure 2 Schematic diagram of a tight gas well similarity calculation device according to an embodiment of the present application. DETAILED DESCRIPTION
[0043] This application describes multiple embodiments, but this description is exemplary rather than restrictive, and it will be apparent to those skilled in the art that there may be more embodiments and implementations within the scope of the embodiments described herein. Although many possible feature combinations are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with any other feature or element in any other embodiment, or may replace any other feature or element in any other embodiment.
[0044] This application includes and contemplates combinations of features and elements known to those of ordinary skill in the art. The embodiments, features, and elements disclosed in this application may also be combined with any conventional features or elements to form a unique inventive solution defined by the claims. Any features or elements of any embodiment may also be combined with features or elements from other inventive solutions to form another unique inventive solution defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in this application may be implemented individually or in any appropriate combination. Therefore, except for the limitations made according to the appended claims and their equivalents, the embodiments are not subject to other limitations. In addition, various modifications and changes may be made within the scope of protection of the appended claims.
[0045] In addition, when describing representative embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not rely on the specific order of the steps described herein, the method or process should not be limited to the steps in the specific order described. As will be understood by those skilled in the art, other orders of steps are also possible. Therefore, the specific order of the steps set forth in the specification should not be interpreted as a limitation to the claims. In addition, the claims for the method and / or process should not be limited to performing their steps in the order written, and those skilled in the art can readily understand that these orders can be changed and still remain within the spirit and scope of the embodiments of the present application.
[0046] In some technologies, the similarity calculation of wells is generally performed by evaluating the similarity of work maps and well logging images. The above methods are limited to the calculation and comparison of well parameters, such as work map similarity and well logging image similarity;
[0047] The dynamometer diagram similarity method optimizes data analysis by identifying the similarity of collected dynamometer diagrams, reducing unnecessary computation and storage, thereby saving analysis time and lowering computer load. This helps speed up data analysis and enhances database storage and query efficiency. However, this method is only applicable to similarity calculations for dynamometer diagrams of wells with pumping capacity.
[0048] The logging image similarity method uses the dynamic image data of each plate of electrical imaging logging to calculate multiple texture feature values, and calculates the Euclidean distance of various texture feature values between electrical imaging images, thereby realizing the similarity of two electrical imaging logging images by computer comparison.
[0049] Existing methods do not involve similarity calculation at the overall well level.
[0050] The embodiment of the present disclosure provides a method for calculating the similarity of tight gas wells, such as Figure 1 As shown, the method includes steps S100-S120, which are specifically as follows:
[0051] S100. Obtain characteristic data of N wells to be calculated, and divide the characteristic data into multiple types, where N is a positive integer greater than or equal to 2;
[0052] S110. Calculate the importance of each type of feature of the first well and the second well respectively based on the feature data;
[0053] S120. Perform importance fusion calculation based on the importance of multiple types of features to determine the similarity between the first well and the second well
[0054] In an exemplary embodiment, characteristic data of N wells whose similarity is to be calculated are obtained, where the number of N wells is determined according to specific circumstances, for example, N can be 50 wells or 100 wells.
[0055] In one exemplary embodiment, well data features are prioritized and their importance calculated using the Analytic Hierarchy Process (AHP). Based on collected data features and experience, feature priorities are categorized into three types: the first type: branch, block, development unit, well type, and well type; the second type: production layer and target layer; and the third type: numerical features (e.g., spontaneous potential, natural gamma, volume density, acoustic transit time, gas saturation, current, voltage, well diameter, etc.). Data features can also be categorized into four or five types, depending on specific circumstances.
[0056] In an exemplary embodiment, the importance of each type of feature of the first well and the second well is calculated based on the feature data, including:
[0057] For the first type:
[0058] Determine the total number of features that are common to the first well and the second well based on the first type of feature well data of the first well and the first type of feature well data of the second well;
[0059] The determined total value is used as the importance of the first type of features of the first well and the second well.
[0060] For example, if the corresponding features are the same, it is counted as 1, and if they are different, it is counted as 0. The maximum value of this type of feature result is 5, and the minimum value is 0.
[0061] For the second type:
[0062] The string fuzzy matching method is used to extract the production layer and the target layer from the second type of characteristic well data of the first well and the second type of characteristic well data of the second well;
[0063] Calculate the jaccard similarity coefficient between the production layer and the target layer;
[0064] The Jaccard similarity coefficient is used as the importance of the second type feature between the first well and the second well.
[0065] For example: 1) First, use string fuzzy matching (regular or similarity metric based on edit distance) to extract the production layer and target layer. For example, for the Shanxi Group 4+5 coal and the Taiyuan Group 8+9 coal, the extraction results are: Shanxi 4 coal, Shanxi 5 coal, Taiyuan 8 coal, Taiyuan 9 coal.
[0066] 2) Calculate the Jaccard similarity coefficient between the production layer and the target layer;
[0067] Wherein, the similarity coefficient formula is:
[0068]
[0069] In this calculation process, the production layer and the target layer are processed according to the text, and the similarity of the text is compared. The larger the jaccard value, the higher the similarity.
[0070] For the third type:
[0071] The 3σ principle was used to filter out the outliers in the third type characteristic data of the first well and the third type characteristic data of the second well;
[0072] Normalize the filtered feature data;
[0073] Calculate the Euclidean distance between the normalized feature data of the first well and the normalized feature data of the second well, and use it as the importance of the third type of feature;
[0074] The calculation formula of Euclidean distance is as follows:
[0075]
[0076] In the above calculation formula, n is the total number of features of the third category, x i and y i are the values of the i-th feature of the first and second wells, respectively. For example, if n = 3, the feature vector of the first well is [0.1, 0.2, 0.4], and the feature vector of the second well is [0.5, 0.5, 0.4], then the Euclidean distance d = 0.5.
[0077] In an exemplary embodiment, importance fusion is performed based on the importance of multiple types of features to determine the similarity between the first well and the second well, including: performing a weighted average calculation based on the determined importance of each type of feature and the predetermined weight coefficient corresponding to each type of feature to obtain the similarity between the first well and the second well.
[0078] In an exemplary embodiment, the second well is any well among the N-1 wells excluding the first well. For example, if N is 20 and the first well is Well A, then Well A is compared one-by-one with the other 19 wells, and its similarity with each well is determined. Thus, for Well A, the first well will have 19 similarity results.
[0079] In one exemplary embodiment, after calculating the similarity between the first well and all other second wells, all similarities are sorted in descending order. The second wells corresponding to the first M sequences in the descending order are selected for similarity correction. For example, if N is 20 and the first well is Well A, there will be 19 similarity results for Well A. These 19 similarity results are sorted, and the first 10 wells in the similarity sequence are selected for similarity correction.
[0080] In an exemplary embodiment, selecting the second well corresponding to the first M sequences in the descending order for similarity correction includes:
[0081] The first step is to determine the total number of various characteristics of N wells;
[0082] Step 2: determine the number of the largest common features between the first well and the second well;
[0083] Step 3: Based on the total number of features of each type and the number of the largest common features, a correction coefficient is used to determine the corrected similarity; wherein the correction coefficient formula is:
[0084]
[0085] In the above formula, a is the total number of features across N wells, and b is the maximum number of common features between the first well and the second well corresponding to M similarities. For example, if the total number of features for 20 wells is 50, and the maximum number of common features between wells A and B is determined to be 10, a correction factor is calculated and used to adjust the similarity between wells A and B.
[0086] An embodiment of the present disclosure also provides a device for calculating the similarity of tight gas wells, which includes: a memory 210 and a processor 220; the memory is used to store a program for calculating the similarity of tight gas wells, and the processor is used to read and execute the program for calculating the similarity of tight gas wells, and perform any method described in the above embodiments.
[0087] The embodiment of the present disclosure further provides a computer storage medium, in which a program for calculating the similarity of tight gas wells is stored. The program is configured to execute the method for calculating the similarity of tight gas wells described in any one of the above embodiments when running.
[0088] Example 1
[0089] The following example illustrates the method for calculating the similarity of tight gas wells. The process is as follows:
[0090] 300. Obtain characteristic data of N wells to be predicted; the value of N here is determined according to different situations, and this example takes N=100 as an example.
[0091] 310. Based on the characteristics of the collected data, the feature priorities are divided into three types;
[0092] Among them, the first type: branch, block, development unit, well type, well type;
[0093] The second type: production layer, destination layer;
[0094] The third type: numerical characteristics (spontaneous potential, natural gamma, volume density, acoustic time difference, gas saturation, wellbore diameter, current, voltage, etc.).
[0095] 320. Calculate the importance of each type of data features according to different types.
[0096] The first type of feature is counted as 1 if the corresponding features are the same, and 0 if they are different. The maximum value of this type of feature result is 5, and the minimum value is 0.
[0097] The second type of feature uses the following process to calculate the importance:
[0098] 1) First, use string fuzzy matching (regular or similarity metric based on edit distance) to extract the production layer and target layer. For example, for the Shanxi Group 4+5 coal and the Taiyuan Group 8+9 coal, the extraction results are: Shanxi 4 coal, Shanxi 5 coal, Taiyuan 8 coal, Taiyuan 9 coal.
[0099] 2) Calculate the jaccard similarity coefficient of the production layer and the target layer. In this step, the production layer and the target layer are processed according to the text, and the similarity of the text is compared. The larger the jaccard value, the higher the similarity. The formula for calculating the jaccard similarity coefficient is as follows:
[0100]
[0101] For the third type of features, after filtering outliers using the 3σ principle for numerical features, the data is normalized to the maximum and minimum, and then the Euclidean distance is calculated as the similarity result.
[0102] 330. The similarity results of the three types of features are weighted averaged according to the empirical weight coefficient to obtain the final similarity result between the well and another well.
[0103] 340. For any well, calculate its similarity with the other wells, sort them in descending order of similarity, and retain only the top m similar wells. The value of m here can be determined based on the specific situation and can be 10, 20, etc. For example, m = 20 is used for the correction calculation.
[0104] 350. Perform similarity correction on the top 20 similar wells.
[0105] Corrections were made for this well and 20 other similar wells. The specific correction process is as follows:
[0106] 3501. Determine the total number of characteristics of each type for N = 100 wells (e.g., 50);
[0107] 3502. Determine the maximum number of common features between well A and the other 20 wells (B1, B2, ..., B20). (For example, the maximum number of common features between A and B1 is 10.)
[0108] 3503. Based on the total number of features of each type and the number of the largest common features, a correction coefficient formula is used to determine the corrected similarity; wherein the correction coefficient formula is:
[0109]
[0110] In the above formula, a is the total number of features of 100 wells, for example, 50, and b is the number of the maximum common features of A and B1, that is, 10. The correction coefficient is determined based on the values of a and b, and the similarity between A and B1 is corrected based on the determined correction coefficient.
[0111] This embodiment of the application comprehensively evaluates the similarity of two wells based on multi-dimensional feature data, rather than being limited to a single feature data set, making the evaluation process more comprehensive. Furthermore, the calculation and fusion methods for different types of features (text, numerical) make the evaluation results more scientific and accurate.
[0112] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
Claims
1. A method for calculating the similarity of tight gas wells, characterized in that: The method comprises: Acquire characteristic data of N wells to be calculated, and classify the characteristic data into multiple types, where N is a positive integer greater than or equal to 2; Calculate the importance of each type of feature of the first well and the second well respectively according to the feature data; Performing importance fusion calculation based on the importance of multiple types of features to determine the similarity between the first well and the second well; The feature data is divided into multiple types, including: Dividing the characteristic data into three types; Calculating the importance of each type of feature of the first well and the second well respectively based on the feature data includes: determining a total value of the same feature in the first type based on the first type feature data of the first well and the first type feature data of the second well; The determined total value is used as the importance of the first type of feature of the first well and the second well; Calculating the importance of each type of feature of the first well and the second well respectively based on the feature data includes: A string fuzzy matching method is used to extract the production layer and the target layer from the second type feature data of the first well and the second type feature data of the second well; Calculating the Jaccard similarity coefficient between the production layer and the target layer according to a similarity coefficient formula; The jaccard similarity coefficient is used as the importance of the second type of feature between the first well and the second well; Wherein, the similarity coefficient formula is: ; Calculating the importance of each type of feature of the first well and the second well respectively based on the feature data includes: The 3σ principle was used to filter out the outliers in the third type characteristic data of the first well and the third type characteristic data of the second well respectively; Normalize the filtered feature data; Calculating the Euclidean distance between the normalized feature data of the first well and the normalized feature data of the second well; The obtained Euclidean distance is used as the importance of the third type feature between the first well and the second well.
2. The method for calculating the similarity of tight gas wells according to claim 1, characterized in that: The performing importance fusion calculation based on the importance of multiple types of features to determine the similarity between the first well and the second well includes: A weighted average calculation is performed based on the determined importance of each type of feature and a predetermined weight coefficient corresponding to each type of feature to determine the similarity between the first well and the second well.
3. The method for calculating the similarity of tight gas wells according to claim 1, characterized in that: The second well is any well other than the first well among the N wells; After performing importance fusion calculation based on the importance of multiple types of features and determining the similarity between the first well and the second well, the method further includes: Arrange the obtained N-1 similarities in descending order; Perform similarity correction on the first M similarities in descending order, where M is a positive integer less than N-1.
4. The method for calculating the similarity of tight gas wells according to claim 3, characterized in that: The similarity correction of the first M similarities in the descending order includes: Determine the total number of features for N wells; Determining the maximum number of common features between the first well and the second well corresponding to the M similarities; According to the total number of features and the number of the largest common features, a correction coefficient formula is used to determine the corrected similarity; wherein the correction coefficient formula is: , In the above formula, a is the total number of features of each type in N wells, and b is the number of the maximum common features between the first well and the second well corresponding to M similarities.
5. A device for calculating similarity of tight gas wells, characterized in that: The device includes: a memory and a processor; the memory is used to store a program for calculating the similarity of tight gas wells, and the processor is used to read and execute the program for calculating the similarity of tight gas wells, and execute the method according to any one of claims 1-4.
6. A computer storage medium, characterized in that The storage medium stores a program for calculating the similarity of tight gas wells, and the program is configured to execute the method for calculating the similarity of tight gas wells according to any one of claims 1 to 4 when running.
Citation Information
Patent Citations
Method and device for determining the position of perforation layer of oil reservoir
CN106504107A
Analysis method for weight of main controlling factor of oil field output
CN106779275A