A gas well working condition diagnosis and analysis method based on real-time production data
By setting up collection points inside the gas well, performing data preprocessing and feature screening, and establishing a gradient boosting decision tree algorithm model, the problem of real-time diagnosis of gas well production conditions was solved, enabling early detection and treatment of gas well liquid accumulation and hydrate blockage, and improving production efficiency.
Patent Information
- Application Number
- CN202411886047.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2044-12-20
AI Technical Summary
Existing technologies lack the ability to diagnose operating conditions by analyzing the changing trends of real-time production data parameters of gas wells. This leads to inaccurate judgments on the timing and extent of liquid accumulation in gas wells, thus affecting gas well production.
Based on real-time production data, data is acquired by setting up collection points inside the gas well, and then preprocessed, feature-filtered, and feature-integrated to establish a gradient boosting decision tree algorithm model to determine whether there is liquid accumulation or hydrate blockage inside the gas well.
It enables early detection and timely management of abnormal operating conditions in gas wells, increases natural gas production, and reduces errors in production data analysis.
Smart Images

Figure CN119357839B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of working condition diagnosis and early warning, and particularly relates to a gas well working condition diagnosis and analysis method based on real-time production data. BACKGROUND
[0002] Most gas wells will produce water in the later production stage, thereby affecting the normal production of the gas well, and accurate determination of the gas well liquid loading time and liquid loading degree is extremely important for the production of the gas well and directly affects the development of the gas well drainage gas recovery measures and the later management.
[0003] Chinese Patent Publication No. CN111075428A discloses a rapid discrimination method for gas well wellbore liquid loading timing and depth, which comprises: calculating the gas well monthly water-gas ratio according to the gas well wellhead daily water production and daily gas production; obtaining a water-gas ratio curve varying with time according to the gas well monthly water-gas ratio; determining the gas well liquid loading time by monitoring the change rule of the curve; and calculating the gas well wellbore liquid loading depth according to the change of water production before and after liquid loading.
[0004] It can be seen that the prior art lacks working condition diagnosis by the change trend of real-time production data parameters of the production working condition. SUMMARY
[0005] Therefore, the present application provides a gas well working condition diagnosis and analysis method based on real-time production data to overcome the problem of working condition diagnosis by the change trend of real-time production data parameters of the production working condition in the prior art.
[0006] To achieve the above-mentioned purpose, the present application provides a gas well working condition diagnosis and analysis method based on real-time production data, which comprises,
[0007] A plurality of collection points are arranged in the target gas well, actual characteristic production data of each collection point is obtained, each actual characteristic production data is integrated to obtain an actual gas well production data set, and the actual gas well production data set is preprocessed to obtain a target gas well production data set;
[0008] Each actual characteristic production data in the target gas well production data set is subjected to feature screening to obtain a plurality of target characteristic production data, and each target characteristic production data is integrated to obtain a target gas well production data string;
[0009] The target gas well production data string is subjected to end-of-life processing, and the target gas well production data string is divided according to the end-of-life processing result to obtain a training set, a test set and a validation set;
[0010] The target gas well production data string is sequentially subjected to gas well production data training, testing and validation to establish a target algorithm model;
[0011] The collected real-time gas well production data string is input into the target algorithm model to determine the actual gas well working condition in the target gas well.
[0012] Further, the process of preprocessing the actual gas well production data set to obtain a target gas well production data set comprises:
[0013] The actual acquisition condition of the actual characteristic production data of each collection point is identified.
[0014] The data missing level is determined according to the actual acquisition condition.
[0015] Based on the data missing level, it is determined whether to start the singular data filtering mode, or to calculate the actual redundancy, or to start the re-collection mode.
[0016] Further, the process of preprocessing the actual gas well production data set to obtain a target gas well production data set comprises:
[0017] Based on starting the singular data filtering mode, singular data determination is performed on each actual characteristic production data, and the actual singular condition of the actual gas well production data set is determined according to the singular data determination result.
[0018] The singular data level is determined according to the actual singular condition.
[0019] Based on the singular level, it is determined whether to start the feature screening mode.
[0020] Further, the process of determining the data missing level according to the actual acquisition condition comprises:
[0021] The actual characteristic production data is marked with missing data according to whether the actual characteristic production data is missing data, and the data missing level is determined according to the number of missing data marks.
[0022] The primary data missing level is one missing data mark.
[0023] The intermediate data missing level is two missing data marks.
[0024] The high-level data missing level is three or more missing data marks.
[0025] Further, the process of determining, based on the data missing level, whether to start the singular filtering mode, or to calculate the actual redundancy, or to start the re-collection mode comprises:
[0026] When the intermediate data missing level is two missing data marks, the actual redundancy is determined according to the actual influence value of the two missing data, and whether to start the singular filtering mode is determined through the actual redundancy calculation result.
[0027] Further, the process of determining the actual singular situation of the actual gas well production data set based on the singular determination of each actual characteristic production data when the singular filtering mode is turned on comprises:
[0028] determining the singular data based on whether each actual characteristic production data is in the preset actual characteristic production data interval.
[0029] Further, the process of determining the actual singular situation of the actual gas well production data set based on the singular determination of each actual characteristic production data when the singular filtering mode is turned on further comprises:
[0030] determining the singular data based on whether each actual characteristic production data is the preset standard actual characteristic production data.
[0031] Further, the process of determining the singular data level based on the actual singular situation comprises:
[0032] marking the actual gas well production data set containing one singular data as a primary singular data level;
[0033] marking the actual gas well production data set containing two singular data as an intermediate singular data level;
[0034] marking the actual gas well production data set containing more than or equal to three singular data as a high-level singular data level.
[0035] Further, the process of determining whether to turn on the characteristic screening mode based on the singular level comprises:
[0036] turning on the characteristic screening mode for the actual gas well production data set marked as the primary singular data level;
[0037] deleting data for the actual gas well production data set marked as the high-level singular data level.
[0038] Further, the process of determining whether to turn on the characteristic screening mode based on the singular level further comprises:
[0039] based on the actual gas well production data set being marked as the intermediate singular data level, determining the actual singular data redundancy by calculating the singular value of the two singular data, and determining whether to turn on the characteristic screening mode based on the actual singular data redundancy.
[0040] Further, the process of performing characteristic screening on each actual characteristic production data in the target gas well production data set to obtain a plurality of target characteristic production data comprises:
[0041] According to the actual correlation degree of each actual feature production data and the actual gas production, the actual correlation degree is calculated;
[0042] According to the first screening standard, the actual feature production data is screened to obtain a plurality of target feature production data;
[0043] The first screening standard is to select the top four features with the highest actual correlation degree as the screening result.
[0044] Further, the target gas well production data string is processed at the end, and the target gas well production data string is divided according to the end processing result; training set, test set and validation set
[0045] The actual missing number of each target feature production data in a preset fixed time is counted;
[0046] The missing number of each feature of the target gas well production data string is counted, and the actual missing number of features A, B, C and D is respectively 10, 20, 30 and 40. The actual missing number of feature A is the least;
[0047] Feature A is set as the dependent variable Y, and the remaining features B, C and D are set as the independent variables X1, X2 and X3 in turn;
[0048] According to whether the data of feature A is missing, the training set and the test set are divided. In the training set, feature A is not missing, which is used as the training set. The data set with missing feature A is used to construct the test set, and the remaining data set is used to construct the validation set.
[0049] Further, the real-time gas well production data string collected is input into the target algorithm model to determine the actual gas well working condition in the target gas well, including:
[0050] When training the target gas well production data string, the target algorithm model is updated by gradient descent to minimize the loss function;
[0051] When testing the target gas well production data string, each target gas well production data string is added to obtain the final test result.
[0052] Further, when the target algorithm model is trained according to the training set, the target feature data input by the trained target algorithm model is used to predict gas well liquid loading or hydrate freezing blockage.
[0053] Further, based on the mechanism formula, liquid loading label and freezing blockage label are added to gas well liquid loading and hydrate freezing blockage, and corresponding processing measures are taken according to the corresponding labels.
[0054] Compared with the prior art, the beneficial effects of the present application are that by preprocessing, feature processing and end processing of the gas well production data, the gas well production data is divided into training set, test set and validation set, the gradient boosting decision tree algorithm model is established, and based on the gas well temperature, gas well oil pressure and gas well casing pressure of the real-time gas well production data, it is judged whether the working condition in the gas well is liquid loading or hydrate frozen blockage. The method realizes early discovery and timely treatment of abnormal working conditions of the gas well, lays a foundation for reasonable planning of field production resources and improvement of natural gas production.
[0055] Further, by preprocessing the gas well production data, the gas well production data can be filled, filtered and repeated data deleted, and the preprocessed gas well production data can better restore the real production data in the gas well, avoiding errors in subsequent analysis of the gas well production data.
[0056] Further, by setting the missing level of the gas well production data to filter each gas well production data string, the gas well production data can be quickly preprocessed. When the number of missing data in the actual gas well production data string is less than or equal to 1, the missing level of the gas well production data is the primary missing level, and the actual gas well production data string is not cleared. Since one data is missing, it has little effect on the actual gas well production data string and will not affect the subsequent processing of the gas well production data string. When the number of missing data in the actual gas well production data string is less than or equal to 2, the missing level of the gas well production data is the intermediate missing level, and the actual gas well production data string is not cleared. According to the correlation degree of the missing data and the gas production of the gas well, it is determined whether to mark the actual gas well production data string with singular data. The actual gas well production data string with high correlation degree of missing data and gas production of the gas well can be cleared, avoiding the subsequent calculation process of the actual gas well production data string with high correlation degree of missing data and gas production of the gas well, and affecting the diagnosis result. BRIEF DESCRIPTION OF DRAWINGS
[0057] Figure 1 Process flowchart of the gas well working condition diagnosis and analysis method based on real-time production data in the embodiment;
[0058] Figure 2 Process flowchart of determining data missing level of the gas well working condition diagnosis and analysis method based on real-time production data in the embodiment;
[0059] Figure 3 Process flowchart of determining data singular level of the gas well working condition diagnosis and analysis method based on real-time production data in the embodiment;
[0060] Figure 4 Process flowchart of obtaining a plurality of target feature production data of the gas well working condition diagnosis and analysis method based on real-time production data in the embodiment. DETAILED DESCRIPTION
[0061] In order to make the objects and advantages of the present application clearer, the following further describes the present application with reference to examples; it should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.
[0062] The preferred embodiments of the present application are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that the embodiments are only used to explain the technical principles of the present application and are not intended to limit the protection scope of the present application.
[0063] It should be noted that, in the description of the present application, the terms indicating the direction or positional relationship such as "upper", "lower", "left", "right", "inner", "outer" and the like are based on the direction or positional relationship shown in the drawings, which is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present application.
[0064] In addition, it should also be noted that, in the description of the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connection" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integrally connected; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, or the internal communication of two elements. Those skilled in the art can understand the specific meaning of the above terms in the present application according to the specific circumstances.
[0065] Please refer to Figure 1 - Figure 4 shown, Figure 1 is a process flow chart of the gas well working condition diagnosis and analysis method based on real-time production data in the example; Figure 2 is a process flow chart of determining data missing level of the gas well working condition diagnosis and analysis method based on real-time production data in the example; Figure 3 is a process flow chart of determining data singularity level of the gas well working condition diagnosis and analysis method based on real-time production data in the example; Figure 4 is a process flow chart of obtaining a plurality of target characteristic production data of the gas well working condition diagnosis and analysis method based on real-time production data in the example.
[0066] The present embodiment provides a gas well working condition diagnosis and analysis method based on real-time production data, comprising the following steps,
[0067] S1, a plurality of collection points are arranged in a target gas well, actual characteristic production data of each collection point is obtained, each actual characteristic production data is integrated to obtain an actual gas well production data set, and the actual gas well production data set is preprocessed to obtain a target gas well production data set;
[0068] S2, performing feature screening on each actual feature production data in the target gas well production data set to obtain a plurality of target feature production data, and integrating each target feature production data to obtain a target gas well production data string;
[0069] S3, performing end-of-life processing on the target gas well production data string, and dividing the target gas well production data string according to the end-of-life processing result to obtain a training set, a test set and a validation set;
[0070] S4, sequentially performing gas well production data training, testing and validation on the target gas well production data string to establish a target algorithm model;
[0071] S5, inputting the collected real-time gas well production data string into the target algorithm model to determine the actual gas well working condition in the target gas well.
[0072] In the embodiment, the target algorithm model is a GBDT algorithm model, which can realize high prediction accuracy, and other algorithm models capable of realizing this function are not specifically limited.
[0073] Specifically, the process of preprocessing the actual gas well production data set to obtain a target gas well production data set includes:
[0074] Identifying the actual acquisition situation of the actual feature production data of each collection point;
[0075] Determining the data missing level according to the actual acquisition situation;
[0076] Based on the data missing level, determining to start the singular data filtering mode, or calculating the actual redundancy, or starting the re-collection mode.
[0077] Specifically, the actual feature production data is marked according to whether the actual feature production data is missing data, and the data missing level is determined according to the number of missing data marks,
[0078] The primary data missing level is one missing data mark;
[0079] The intermediate data missing level is two missing data marks;
[0080] The advanced data missing level is three or more missing data marks.
[0081] A plurality of collection points are arranged in the target gas well, and various sensors capable of collecting gas well production data are arranged at each collection point. The various sensors include oil pressure sensors, casing pressure sensors, gas well temperature sensors, gas production sensors, water production sensors, flow pressure sensors, and water flow velocity sensors. The gas production collected by the gas production sensor is the dependent variable.
[0082] The actual gas well production data strings collected by each collection point are numbered. The first actual gas well production data string is Aa1, the second actual gas well production data string is Aa2, the third actual gas well production data string is Aa3, and the nth actual gas well production data string is Aan.
[0083] Any actual gas well production data string Aai (i = 1, 2, 3,..., n) is selected, and the number of missing data is determined.
[0084] When the number of missing data in the actual gas well production data string Aai is less than or equal to one data, the actual gas well production data is at a primary missing level, and the next group of actual gas well production data strings is determined.
[0085] When the number of missing data in the actual gas well production data string Aai is equal to two data, the actual gas well production data is at an intermediate missing level, and the actual missing redundancy is determined according to the actual influence value of the two missing data. Whether to start the singular filtering mode is determined by the actual missing redundancy calculation result.
[0086] The first missing data weight is Ba, and the second missing data weight is Bb.
[0087] The actual missing influence value HBa of the first missing data Ba is calculated, HBa = Ba / (Aa1+Aa2+Aa3...+Aan)×q1, where q1 is the influence compensation parameter of the proportion of the first missing data weight in the actual gas well production data string to the actual missing influence value of the first missing data.
[0088] The actual missing influence value HBb of the second missing data Bb is calculated, HBb = Bb / (Aa1+Aa2+Aa3...+Aan)×q2, where q2 is the influence compensation parameter of the proportion of the second missing data weight in the actual gas well production data string to the actual missing influence value of the first missing data.
[0089] The actual redundancy Rq of the two missing data is (HBa+HBb)×a1, where a1 is the influence compensation parameter of the sum of the first missing data and the second missing data to the actual missing redundancy of the two missing data.
[0090] The standard missing redundancy Rqb is set.
[0091] If Rq≤Rqb, the actual gas well production data missing level is determined as a primary missing level;
[0092] If Rq>Rqb, the actual gas well production data missing level is determined as a high missing level.
[0093] Specifically, the process of determining the actual singularity of the actual gas well production data set based on the singularity determination of each actual characteristic production data when the singularity filtering mode is turned on comprises:
[0094] The singularity data is determined based on whether each actual characteristic production data is in a preset actual characteristic production data interval.
[0095] Specifically, the process of determining the actual singularity of the actual gas well production data set based on the singularity determination of each actual characteristic production data when the singularity filtering mode is turned on further comprises:
[0096] The singularity data is determined based on whether each actual characteristic production data is a preset standard actual characteristic production data.
[0097] The singularity data comes from an abnormal acquisition device, and the singularity data filtering rule is set as
[0098] The casing pressure is less than the oil pressure, and the singularity data is determined;
[0099] The casing pressure divided by the oil pressure is not less than five, and the singularity data is determined;
[0100] The oil pressure and the casing pressure are greater than 0 MPa and less than 40 MPa;
[0101] The temperature is greater than -50℃ and less than 50℃;
[0102] Specifically, the process of determining the singularity data level based on the actual singularity comprises:
[0103] The actual gas well production data set containing one singularity data is marked as a primary singularity data level;
[0104] The actual gas well production data set containing two singularity data is marked as an intermediate singularity data level;
[0105] The actual gas well production data set containing greater than or equal to three singularity data is marked as a high singularity data level.
[0106] Specifically, the process of determining whether to turn on the characteristic screening mode based on the singularity level comprises:
[0107] The characteristic screening mode is turned on for the actual gas well production data set marked as the primary singularity data level;
[0108] The actual gas well production data set marked as the high-level singular data level is subjected to the missing data processing.
[0109] Specifically, the process of determining whether to start the feature screening mode based on the singular level further comprises:
[0110] Based on the case that the actual gas well production data set is marked as the intermediate singular data level, the actual singular data redundancy is determined by calculating the singular values of the two singular data, and whether to start the feature screening mode is determined according to the actual singular data redundancy.
[0111] Any actual gas well production data string Aai, i = 1, 2, 3,..., n is selected, and the number of missing data thereof is judged,
[0112] When the number of singular data in the actual gas well production data string Aai is less than or equal to 1 data, the actual gas well production data singular level is the primary singular level, and the next group of actual gas well production data string is judged;
[0113] When the number of singular data in the actual gas well production data string Aai is equal to 2 data, the actual gas well production data singular level is the intermediate singular level, the actual singular redundancy is determined according to the actual influence values of the two singular data, and whether to start the feature screening mode is determined by the actual singular redundancy calculation result
[0114] The first singular data weight is set as Ca, and the second singular data weight is set as Cb,
[0115] The actual singular influence value GCa of the first singular data Ca is calculated, GCa = Ca / (Aa1 + Aa2 + Aa3... + Aan) × j1, wherein j1 is an influence compensation parameter of the proportion of the first singular data weight in the actual gas well production data string to the actual influence value of the first singular data;
[0116] The actual singular influence value GCb of the second missing data Cb is calculated, GCb = Cb / (Aa1 + Aa2 + Aa3... + Aan) × j2, wherein j2 is an influence compensation parameter of the proportion of the second singular data weight in the actual gas well production data string to the actual influence value of the first singular data;
[0117] The actual redundancy Cq of the two singular data is calculated, Cq = (GCa + GCb) × a2, wherein a2 is an influence compensation parameter of the sum of the first singular data and the second singular data to the actual singular redundancy of the two singular data;
[0118] The standard singular redundancy Cqb is set,
[0119] If Cq ≤ Cqb, the actual gas well production data singular level is determined as the primary singular data level;
[0120] If Cq>Cqb, the actual gas well production data singularity level is determined as high singularity data level.
[0121] Specifically, the process of feature screening each actual feature production data in the target gas well production data set to obtain a plurality of target feature production data includes:
[0122] According to each actual feature production data and actual gas production, calculate the actual correlation degree;
[0123] Based on the first screening standard, screen each actual feature production data to obtain a plurality of target feature production data;
[0124] The first screening standard is to select the top four features in the actual correlation degree as the screening result.
[0125] In this embodiment, the gas well oil pressure Aa, the gas well casing pressure Ba, the gas well temperature C, the actual gas production D, the water production E, the flow pressure F and the water flow velocity G actual feature production data are screened,
[0126] The first actual correlation degree of the gas well oil pressure A and the actual gas production D is r1;
[0127] r1= , wherein Aa i is the ith acquisition value of the gas well oil pressure; is the average value of the gas well oil pressure =(Aa1+Aa2+...+Aan) / n;D i is the ith acquisition value of the actual gas production; is the average value of the actual gas production =(D1+D2+...+Dn) / n;
[0128] The second actual correlation degree of the gas well casing pressure and the actual gas production is r2;
[0129] r2= , wherein Ba i is the ith acquisition value of the gas well casing pressure; is the average value of the gas well casing pressure =(Ba1+Ba2+...+Ban) / n; is the ith acquisition value of the actual gas production; is the average value of the actual gas production =(D1+D2+...+Dn) / n;
[0130] The third actual correlation degree of the gas well temperature and the actual gas production is r3;
[0131] r3= wherein, C i is the ith obtained value of casing pressure of the gas well; is the average value of casing pressure of the gas well = (C1+C2+...+Cn) / n;D i is the ith obtained value of actual gas production; is the average value of actual gas production = (D1+D2+...+Dn) / n;
[0132] The fourth actual correlation degree between water production and actual gas production is r4;
[0133] r4= wherein, E i is the ith obtained value of water production; is the average value of water production = (E1+E2+...+En) / n; is the ith obtained value of actual gas production; is the average value of actual gas production = (D1+D2+...+Dn) / n; ...
[0134] The fifth actual correlation degree between flowing pressure and actual gas production is r5;
[0135] The sixth actual correlation degree between water flow velocity and actual gas production is r6;
[0136] The actual correlation degrees are sorted according to descending order, and the top four features of the actual correlation degrees are selected as the screening results according to the sorting results, and the screening results include the gas well oil pressure Aa, the gas well casing pressure Ba, the gas well temperature C, and the actual gas production D.
[0137] Specifically, the target gas well production data string is processed at the end, and the target gas well production data string is divided according to the end processing result;
[0138] The actual missing number of each target feature production data within a preset fixed time is counted;
[0139] The missing level of the gas well production data is set to filter each gas well production data string, so that the gas well production data can be quickly preprocessed. When the number of missing data in the actual gas well production data string is less than or equal to 1, the missing level of the actual gas well production data is a primary missing level, and the actual gas well production data string is not cleared. Since one missing data has little effect on the actual gas well production data string, it will not affect the subsequent processing of the actual gas well production data string. When the number of missing data in the actual gas well production data string is less than or equal to 2, the missing level of the actual gas well production data is a middle missing level, and the actual gas well production data string is not cleared. Whether to mark the actual gas well production data string as singular data is determined according to the correlation degree of the missing data and the gas production of the well. The actual gas well production data string with a high correlation degree of missing data and gas production of the well can be cleared, so that the actual gas well production data string with a high correlation degree of missing data and gas production of the well is avoided in the subsequent calculation process, and the diagnosis result is affected.
[0140] The collection device is abnormal to cause singular gas well production data. The filtering rule for filtering singular data according to the self-defined algorithm is as follows,
[0141] The casing pressure is less than the oil pressure; the casing pressure divided by the oil pressure is greater than or equal to five; the oil pressure and the casing pressure are greater than 0 MPa and less than 40 MPa; the temperature is greater than -50℃ and less than 50℃. When any one of the above conditions is met in the real-time gas well production data string, the data is determined as singular data, and the real-time gas well production data string is marked as singular data,
[0142] The number of singular data in the real-time gas well production data string Aai marked as singular data is obtained,
[0143] In this embodiment, the existing data set contains four features, namely, gas well oil pressure Aa, gas well casing pressure Ba, gas well temperature C and gas well gas production D. Please refer to Figure 3 for the data set schematic diagram. The method for filling missing data by the random forest algorithm comprises the following steps:
[0144] In step S11, the number of missing values of each feature of the existing data set is counted. It is assumed that the number of missing values of the features of the gas well oil pressure Aa, the gas well casing pressure Ba, the gas well temperature C and the gas well gas production D is 10, 20, 30 and 40 respectively. Among them, the number of missing values of the feature of the gas well oil pressure Aa is the least;
[0145] In step S12, the feature of the gas well oil pressure Aa is set as the dependent variable Y, and the remaining features of the gas well casing pressure Ba, the gas well temperature C and the gas well gas production D are sequentially set as the independent variables X1, X2 and X3;
[0146] Step S13, dividing the training set and the test set according to whether the data of the characteristic gas well oil pressure Aa is missing, the training set without missing characteristic Aa is used as the training set, the data set with missing characteristic Aa is used to construct the test set, and the remaining data set is used to construct the validation set;
[0147] Step S14, using the random forest algorithm to construct a model on the data set 1, F(B, C, D)=Aa, and filling the missing values of the characteristic gas well oil pressure Aa in the test set;
[0148] Step S15, re-dividing the input features and output features according to the actual number of missing characteristic gas well oil pressure Aa, and repeating steps S13-S15 until all data is filled.
[0149] Specifically, the determination method for filtering abnormal gas well production data based on the self-defined algorithm includes,
[0150] When the gas well casing pressure is less than the gas well oil pressure, or the gas well casing pressure divided by the gas well oil pressure is greater than or equal to five, or the gas well oil pressure and the gas well casing pressure are greater than 0 MPa and less than 40 MPa respectively, or the gas well temperature is greater than-50℃ and less than 50℃,
[0151] When the above conditions are met, the gas well production data is abnormal gas well production data, and the abnormal gas well production data comes from the abnormal acquisition equipment.
[0152] The maximum correlation feature generation method in this embodiment adopts the following steps:
[0153] Step S21, according to the actual data and operation experience, three parameters of oil pressure, casing pressure and temperature are selected as initial input features;
[0154] Step S22, constructing an input parameter set of the analysis model with the initial input features, for identifying and diagnosing abnormal working conditions of the gas well.
[0155] The minimum redundant feature deletion determination method is that when the initial input features are only gas well oil pressure, gas well casing pressure and gas well temperature, there is no homogeneity relationship based on the three features, and no feature is deleted.
[0156] According to the gas well liquid loading mechanism formula determination method, it is determined whether liquid loading occurs in the gas well, and according to the gas well hydrate freezing plugging mechanism formula determination method, it is determined whether hydrate occurs in the gas well, including,
[0157] When the unload flow is greater than the actual daily gas production, it is determined that liquid loading occurs in the gas well,
[0158] The unload flow calculation formula is:
[0159]
[0160]
[0161] wherein q cr is the unloaded flow 10 4 m 3 / d;u cr is the unloaded flow rate m / s; A is the cross-sectional area of the tubing m 2 ; p is the pressure at a certain point of the tubing MPa; T is the temperature at the same point of the tubing K; Z is the gas deviation factor; P g is the gas density kg / m 3 ; p l is the liquid density kg / m 3 ,
[0162] The hydrate freezing plugging mechanism formula judgment method of the gas well comprises,
[0163] When the actual pressure of the tubing of the gas well is greater than the critical pressure calculated by the hydrate freezing plugging mechanism formula, it is determined that hydrate is generated in the gas well, and the critical pressure calculation formula is:
[0164]
[0165]
[0166] wherein Ta is the tubing temperature at the test point K; p is the hydrate generation critical pressure corresponding to the tubing temperature Ta at the test point MPa; B and B1 are parameters, which are determined according to the value of the relative density of the natural gas .
[0167] When the gas well production data is trained, the decision tree algorithm model parameters are updated by the gradient descent method to minimize the loss function;
[0168] When the gas well production data is tested, the gradient boosting decision tree algorithm model adds the test results of each basic model to obtain the final test result.
[0169] In the running process of the gradient boosting decision tree algorithm model, the initialization stage basic parameters are determined, including: random seed, number of voltage boosting to be executed, learning rate, loss function to be optimized, sample score for fitting each basic learner, function for measuring segmentation quality, minimum number of samples required for splitting internal nodes, minimum number of samples required for each node and maximum depth of a single regression estimator.
[0170] The feature latitude of the input variables of the gradient boosting decision tree algorithm model is determined as 3, including the gas well temperature A, gas well oil pressure B and gas well casing pressure C of the real-time gas well production data, and the output feature latitude is 1, that is, the output gas well working condition is whether to accumulate liquid or whether to freeze hydrate.
[0171] The gradient boosting decision tree algorithm model judgment process first checks the input data, then outputs the initial judgment category according to the input features, learns the relationship between the input features and the category, and then judges whether the gas well is liquid loading or hydrate frozen blockage through the input data.
[0172] The gradient boosting decision tree algorithm model algorithm judgment process includes,
[0173] Step S510, checking the input gas well production data;
[0174] Step S511, outputting the initial judgment category according to the input features;
[0175] Step S512, learning the relationship between the input features and the category;
[0176] Step S513, judging whether the gas well is liquid loading or hydrate frozen blockage through the input gas well production data.
[0177] The gradient boosting decision tree algorithm model method provided in the embodiment is a gas well working condition diagnosis and analysis method based on real-time production data, and includes,
[0178] Step S531, given a training set { (x1, y1), (x2, y2),..., (x n ,y n )}, minimize the loss function L (y, f) corresponding to the model f (x) to solve the optimal model:
[0179]
[0180] Step S532, in the boosting decision tree, assume is expressed as the sum of a series of decision trees:
[0181]
[0182] Wherein, the function corresponding to the jth decision tree is denoted as h j (x), and the model obtained in the h j (x) step is denoted as
[0183] Step S533, in the boosting tree algorithm, sequentially build the model h j (x), when building the model h j (x) in the h j (x) step, the negative gradient is used as the new target value of x i to train the model;
[0184] Step S534, in the gradient boosting decision tree algorithm, the negative derivative As a new target to build the model, the derivative is calculated according to the data in the training data set, and the loss function on the training set is minimized.
[0185] The gradient boosting decision tree algorithm model building step includes,
[0186] Step S521, an initial model function f0(x) is built, ;
[0187] Step S522, the negative gradient r ij , ;
[0188] Step S523, as a new target, a new model h j (x) is built on the training set;
[0189] Step S524, the step size s j is solved, ;
[0190] Step S525, the function is updated to ;
[0191] Step S526, steps S522-S525 are repeated;
[0192] Step S527, the final model is obtained as f(x)=f m (x).
[0193] The gradient boosting decision tree algorithm model is iterated for multiple rounds, wherein each round fits a new decision tree model to correct the residual of all previous decision trees. In each iteration, the gradient boosting decision tree algorithm model first calculates the residual between the predicted value of the current model and the true value of the training data. Based on the residual, a new decision tree model is trained, which attempts to minimize the residual, i.e. fit the negative gradient of the current residual. In the next round, the gradient boosting decision tree algorithm model trains a new decision tree to fit the residual, and the new decision tree is added to the existing model to reduce the residual generated in the last round of training. The core part of the training code is the fit function.
[0194] Within 6 hours, the average value of the first 10 minutes and the average value of the last 10 minutes meet the conditions of a 10% drop in gas well oil pressure and a 10% rise in gas well casing pressure simultaneously, and the gas well is determined to have liquid accumulation; within 2 hours, the average value of the first 10 minutes and the average value of the last 10 minutes meet the conditions of a 20% or more rise or fall in gas well oil pressure and a 10% or more rise in gas well casing pressure simultaneously, and the gas well is determined to have hydrate frozen blockage.
[0195] So far, the technical solutions of the present application have been described in combination with the preferred embodiments shown in the drawings, but it is easy for those skilled in the art to understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the present application, and the technical solutions after the changes or replacements will all fall within the protection scope of the present application.
[0196] The above only describes the preferred embodiments of the present application and is not intended to limit the present application; the present application can have various changes and variations for those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A gas well working condition diagnosis analysis method based on real-time production data, characterized in that, Comprising, A plurality of collection points are arranged in the target gas well, and actual characteristic production data of each collection point is obtained. The actual characteristic production data is integrated to obtain an actual gas well production data set. The process of obtaining the target gas well production data set includes identifying the actual acquisition condition of the actual characteristic production data of each collection point; determining the data missing level according to the actual acquisition condition; determining the actual missing redundancy based on the two missing data label numbers according to the actual influence value of the two missing data when the intermediate data missing level is two; and determining whether to start the singular filtering mode through the actual missing redundancy calculation result. The determination process sets the first missing data weight as Ba and the second missing data weight as Bb. The first missing data Ba actual missing influence value HBa is calculated, HBa=Ba / (Aa1+Aa2+Aa3...+Aan) ×q1, wherein q1 is the influence compensation parameter of the proportion of the first missing data weight in the actual gas well production data string to the first missing data actual missing influence value, Aa1 is the first actual gas well production data string, Aa2 is the second actual gas well production data string, Aa3 is the third actual gas well production data string,..., and Aan is the nth actual gas well production data string. The second missing data Bb actual missing influence value HBb is calculated, HBb=Bb / (Aa1+Aa2+Aa3...+Aan) ×q2, wherein q2 is the influence compensation parameter of the proportion of the second missing data weight in the actual gas well production data string to the first missing data actual missing influence value. The two missing data actual redundancy Rq=(HBa+HBb) ×a1, wherein a1 is the influence compensation parameter of the sum of the first missing data and the second missing data to the actual missing redundancy of the two missing data. The standard missing redundancy Rqb is set. If Rq≤Rqb, the actual gas well production data missing level is determined as the primary missing level. If Rq>Rqb, the actual gas well production data missing level is determined as the high-level missing level. When the data missing level is the primary data missing level, the singular data filtering mode is started. When the data missing level is the high-level data missing level, the re-collection mode is started. The process of determining the data missing level according to the actual acquisition condition includes marking the actual characteristic production data according to whether the actual characteristic production data is missing data, determining the data missing level according to the number of missing data labels, the primary data missing level is one missing data label number; the intermediate data missing level is two missing data label numbers; and the high-level data missing level is three or more missing data label numbers. The actual gas well production data set is preprocessed. When the singular data filtering mode is started, singular data is determined for each actual characteristic production data based on the singular data determination result, and the actual singular situation of the actual gas well production data set is determined according to the actual singular situation. The process of determining the singular data level according to the actual singular situation includes marking the actual gas well production data set containing one singular data as the primary singular data level. marking the actual gas well production data set containing two singular data as a primary singular data level; marking the actual gas well production data set containing more than or equal to three singular data as a high-level singular data level; determining whether to start a feature screening mode based on the singular data level; determining singular data based on whether each actual feature production data is in a preset actual feature production data interval; performing feature screening on each actual feature production data in the target gas well production data set to obtain a plurality of target feature production data; starting the feature screening mode for the actual gas well production data set marked as the primary singular data level; performing deletion data processing on the actual gas well production data set marked as the high-level singular data level; based on the actual gas well production data set being marked as the intermediate singular data level, determining the actual singular data redundancy by calculating the singular values of the two singular data, and determining whether to start the feature screening mode according to the actual singular data redundancy; in the feature screening process, setting the first singular data weight as Ca, the second singular data weight as Cb, calculating the actual singular influence value GCa of the first singular data Ca, GCa=Ca / (Aa1+Aa2+Aa3...+Aan) ×j1, wherein j1 is an influence compensation parameter of the proportion of the first singular data weight in the actual gas well production data string to the actual influence value of the first singular data; calculating the actual singular influence value GCb of the second missing data Cb, GCb=Cb / (Aa1+Aa2+Aa3...+Aan) ×j2, wherein j2 is an influence compensation parameter of the proportion of the second singular data weight in the actual gas well production data string to the actual influence value of the first singular data; the actual redundancy Cq of the two singular data= (GCa+GCb) ×a2, wherein a2 is an influence compensation parameter of the sum of the first singular data and the second singular data to the actual singular redundancy of the two singular data; setting a standard singular redundancy Cqb; if Cq≤Cqb, the actual gas well production data singular level is determined as the primary singular data level; if Cq>Cqb, the actual gas well production data singular level is determined as the high-level singular data level; calculating an actual correlation degree according to each actual feature production data and actual gas production; screening each actual feature production data based on a first screening standard to obtain a plurality of target feature production data; wherein the first screening standard is to select the top four features in the actual correlation degree as the screening result; the actual feature production data includes gas well oil pressure Aa, gas well casing pressure Ba, gas well temperature C, actual gas production D, water production E, flow pressure F, and water flow velocity G; integrating each target feature production data to obtain a target gas well production data string; performing end-of-life processing on the target gas well production data string, and dividing the target gas well production data string according to the end-of-life processing result to obtain a training set, a test set, and a validation set; counting the actual missing number of each target feature production data within a preset fixed time; The number of missing features of the statistical target gas well production data string is counted. It is assumed that the actual number of missing features of Aa, Ba, Ca` and Da is 10, 20, 30 and 40 respectively, and the actual number of missing features of Aa is the least; Aa is set as the dependent variable Y, and the remaining features Ba, Ca` and Da are set as the independent variables X1, X2 and X3 in turn; The training set and the test set are divided according to whether the data of the feature Aa is missing. The feature Aa is not missing in the training set, which is used as the training set. The data set is missing the feature Aa, which is used to construct the test set; Wherein, Aa is the gas well oil pressure, Ba is the gas well casing pressure; The target algorithm model is established by sequentially performing gas well production data training, testing and verification on the target gas well production data string. The gas well production data is divided into training set, test set and verification set, and the gradient boosting decision tree algorithm model is established. Based on the real-time gas well production data of gas well temperature, gas well oil pressure and gas well casing pressure, it is judged whether the working condition of the gas well is liquid loading or hydrate frozen blockage; The collected real-time gas well production data string is input into the target algorithm model to determine the actual gas well working condition in the target gas well.
2. The real-time production data based gas well condition diagnostic analysis method of claim 1, wherein, The process of determining the actual singular situation of the actual gas well production data set based on the singular determination of each actual feature production data when the singular filtering mode is turned on further includes: Determine the singular data based on whether each actual feature production data is a preset standard actual feature production data.
3. The real-time production data based gas well condition diagnostic analysis method of claim 2, wherein, Inputting the collected real-time gas well production data string into the target algorithm model to determine the actual gas well working condition in the target gas well includes: When training the target gas well production data string, the decision tree algorithm model parameters are updated by gradient descent to minimize the loss function; When testing the target gas well production data string, the gradient boosting decision tree algorithm model adds the test results of each base model to obtain the final test result.
4. The real-time production data based gas well condition diagnostic analysis method of claim 3, wherein, When the target algorithm model is trained according to the training set, each target feature data input by the trained target algorithm model is used to predict gas well liquid loading or hydrate frozen blockage.
5. The real-time production data based gas well performance diagnostic analysis method of claim 4, wherein, Based on the mechanism formula, liquid loading and hydrate frozen blockage are added with liquid loading label and frozen blockage label, and corresponding processing measures are taken according to the corresponding label.
Citation Information
Patent Citations
Gas well effusion prediction method based on ensemble learning
CN110163442A
Method for quickly judging liquid accumulation time and depth of gas well shaft
CN111075428A