A physical constraint and multi-source data double-driven gas well liquid loading diagnosis method
The gas well fluid accumulation diagnosis method, driven by both physical constraints and multi-source data, combines multi-source time-series data and physical constraints to solve the problems of insufficient theoretical assumptions and lack of physical constraints in pure data-driven models in traditional methods, thus achieving accurate diagnosis and high-precision identification of gas well fluid accumulation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTHWEST PETROLEUM UNIV
- Filing Date
- 2025-12-23
- Publication Date
- 2026-07-21
AI Technical Summary
Traditional gas well fluid accumulation diagnosis methods rely on overly idealistic theoretical assumptions, neglect the reservoir-wellbore system coupling mechanism, fail to adequately characterize transient evolution processes, and lack physical constraints in purely data-driven models, resulting in low diagnostic accuracy and poor generalization.
A gas well fluid accumulation diagnosis method driven by both physical constraints and multi-source data is adopted. By integrating MultiROCKET and Transformer networks and PINN embedded physical constraints, a gas well fluid accumulation diagnosis model is established. Combined with deep fusion of multi-source time series data and long-range dependency modeling, physical constraints are applied using the Beggs-Brill multiphase flow and reservoir Darcy flow equations to improve diagnostic accuracy and robustness.
It enables accurate diagnosis of liquid accumulation in gas wells, improves the model's recognition accuracy and physical consistency under complex and unsteady conditions, reduces information loss, and enhances the interpretability and robustness of the diagnosis.
Smart Images

Figure CN121744045B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oil and gas engineering technology, specifically to a method for diagnosing fluid accumulation in gas wells driven by both physical constraints and multi-source data. Background Technology
[0002] As my country's natural gas development enters a new stage of simultaneous large-scale and high-efficiency development, gas fields centered on basins such as Sichuan and Ordos continue to contribute to production capacity, and the cumulative output of gas wells nationwide shows a steady upward trend. However, while production scale continues to expand, the deep challenges faced in the development process are becoming increasingly prominent. Currently, most water-bearing gas reservoirs in China have entered the mid-to-late stage of production. Decreased formation pressure leads to a reduction in the gas phase velocity in the wellbore, thereby weakening the fluid-carrying capacity and causing fluid accumulation in the gas well. Fluid accumulation not only causes problems such as reduced gas well production, decreased recovery rate, and equipment corrosion, but in severe cases, it can even lead to wellhead shutdown and production depletion. Therefore, how to diagnose fluid accumulation during gas well production has become a key technical bottleneck that urgently needs to be overcome in the current process of stabilizing gas field production and improving recovery rate.
[0003] Traditional methods for diagnosing fluid accumulation in gas wells mainly include empirical analysis, pressure gradient analysis, nodal analysis, and critical velocity method. These methods provide a basis for judging the fluid carrying capacity of the wellbore in theory, but they generally have the following limitations: (1) The theoretical assumptions are too idealistic, usually based on steady-state, homogeneous fluid or single-phase flow conditions, which do not match the actual situation of strong unsteady and multiphase flow in gas wells; (2) The coupling mechanism of the reservoir-wellbore system is not sufficiently considered, and the comprehensive influence of the dynamic interaction between the formation and the wellbore on the formation of fluid accumulation is ignored; (3) The transient evolution process is not sufficiently characterized, and the model parameters or judgment criteria are mostly statically set, making it difficult to achieve automated iteration and real-time updates with the dynamic production of gas wells.
[0004] In recent years, artificial intelligence methods have been gradually introduced into the field of gas well fluid accumulation diagnosis. By mining the temporal patterns in dynamic production data, data-driven models have achieved real-time intelligent diagnosis of fluid accumulation, which has improved the diagnostic accuracy to a certain extent. However, such pure data-driven models have obvious limitations: (1) Pure data-driven models only learn data correlations without incorporating core physical mechanisms, and cannot capture the physical essence behind the phenomena. Their "black box" characteristics make their decisions difficult to verify and prone to outputting unreliable results under complex working conditions; (2) Most existing data-driven models often separate the correlation between static geological engineering attributes and dynamic production data, making it difficult to fully characterize the complex coupling mechanism of fluid accumulation, resulting in diagnostic results that are sensitive to noise and have poor generalization. Therefore, it is evident that there is an urgent need in this field for a diagnostic method that can simultaneously utilize multi-source data and physical constraint information. Summary of the Invention
[0005] In view of this, the present invention proposes a gas well fluid accumulation diagnosis method driven by both physical constraints and multi-source data. By integrating data-driven feature extraction capabilities with the consistency requirements of physical constraints, the method achieves real-time and reliable diagnosis of the gas well fluid accumulation process. It can fully explore the potential patterns of production data, while the constraint model output satisfies the physical laws of reservoir-wellbore coupling, thereby effectively improving the diagnostic accuracy, robustness and generalization ability, and providing technical support for stable gas well production development.
[0006] To solve at least one of the above-mentioned technical problems, the present invention provides a method for diagnosing fluid accumulation in gas wells driven by both physical constraints and multi-source data, comprising the following steps: Step S1: Collect multi-source dynamic production data and static parameters of gas wells as the initial dataset and perform preprocessing; Step S2: Set a time window, extract time series data from each target time point in the multi-source dynamic production data, assign liquid accumulation level labels to the time series data, obtain a sample set, and divide the sample set into a training set, a validation set, and a test set; Step S3: Establish a gas well fluid accumulation diagnosis model driven by both physical constraints and multi-source data; Step S4: Train the gas well fluid accumulation diagnostic model using the training set and validation set, and optimize the model using the test set; Step S5: Deploy the optimized model to the field for real-time diagnosis of gas well fluid accumulation.
[0007] The technical effects achieved by this invention are: 1. This invention establishes a gas well fluid accumulation diagnosis mechanism driven by both physical constraints and multi-source data. Through collaborative modeling, it overcomes the problems of traditional methods that rely solely on data-driven approaches and lack physical consistency and interpretability, thereby enabling accurate diagnosis of fluid accumulation in gas wells. 2. This invention utilizes MultiROCKET and Transformer networks and PINN embedded physical constraints to achieve deep fusion of multi-source time series data from gas wells, long-range dependency modeling, and explicit expression of gas-liquid two-phase flow laws, thereby improving the model's accuracy, physical consistency, and interpretability in identifying liquid accumulation under complex unsteady conditions. 3. By performing class equalization processing on the training set, this invention can significantly optimize the class imbalance problem, improve the recognition accuracy and model robustness of minority class fluid accumulation states, achieve high-precision diagnosis, and effectively reduce information loss. Attached Figure Description
[0008] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained from these drawings without creative effort.
[0009] Figure 1 This is a flowchart illustrating the method described in this invention; Figure 2 This is a schematic diagram of the gas well fluid accumulation diagnostic model described in this invention; Figure 3 This is the confusion matrix of the model used in this embodiment of the invention on the test set; Figure 4 This is a visualization of the real-time diagnostic results of actual gas well fluid accumulation after the model is deployed in an embodiment of the present invention. Detailed Implementation
[0010] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings.
[0011] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention.
[0012] See Figure 1 A method for diagnosing fluid accumulation in gas wells, driven by both physical constraints and multi-source data, includes the following steps: Step S1: Collect multi-source dynamic production data and static parameters of gas wells as the initial dataset and perform preprocessing; The scale of the gas wells for data collection can be determined based on the sample set to be used in subsequent models. The main data collected are multi-source dynamic production data and static parameters. Multi-source dynamic production data includes gas production, water production, oil pressure, casing pressure, transmission pressure, and wellhead temperature. Static parameters include wellbore geometric parameters, fluid properties, and reservoir properties. Wellbore geometric parameters include wellbore inclination angle and wellbore diameter. Fluid properties include liquid phase density, gas phase density, and fluid viscosity. Reservoir properties include absolute permeability.
[0013] The preprocessing method includes the following steps: Step S11: Detect and remove outliers in the initial dataset using the 3-sigma criterion; Step S12: Use cubic spline interpolation to fill in missing values to ensure the integrity of the data trend; for static parameters with strong spatial distribution correlation, use nearest neighbor interpolation based on the similarity of neighboring wells and the distribution of well locations to fill in missing values. Step S13: Use wavelet transform to filter out high-frequency noise in the completed multi-source dynamic production data, use Lowess smoothing to correct local outliers in the completed multi-source dynamic production data, stabilize the overall trend, filter out high-frequency noise, and retain key abrupt changes such as well opening / closing. Step S14: Perform min-max normalization on the multi-source dynamic production data to map the data to the [0,1] interval; perform Z-score standardization on the static parameters to transform them into a distribution with a mean of 0 and a standard deviation of 1.
[0014] Step S2: Set a time window, extract time series data from each target time point in the multi-source dynamic production data, assign liquid accumulation level labels to the time series data, obtain a sample set, and divide the sample set into a training set, a validation set, and a test set; Based on expert experience in this field and combined with dynamic production data preprocessed by S1, a sliding window technique is used to extract the temporal context data for each target time point. By analyzing the continuous evolution trend of the effusion state, an effusion level label is assigned to the data at each time point. The method for classifying effusion levels in this invention is shown in Table 1. The label system covers four levels: Level 0 corresponds to "no effusion," Level 1 corresponds to "mild effusion," Level 2 corresponds to "moderate effusion," and Level 3 corresponds to "severe effusion." In subsequent supervised learning, the label level is directly used as the ground truth.
[0015] Table 1. Classification Method for Fluid Accumulation Grades Meanwhile, without overlap between time periods, the sample set is divided into training, validation, and test sets in a 7:2:1 ratio according to the time sequence.
[0016] To address the class imbalance problem in the training set, a combination of SMOTE oversampling and random undersampling is used for class balancing. The steps include: generating synthetic samples from "severe effusion" and "moderate effusion" samples using the SMOTE algorithm to supplement the sample count; and randomly sampling "no effusion" and "mild effusion" samples to ensure that the difference between their number and the number of severe and moderate effusion samples after supplementation does not exceed 10%, thus obtaining a class-balanced training set.
[0017] No sampling is performed on the validation and test sets to preserve the original label distribution, thereby ensuring that the model evaluation results can accurately reflect the liquid accumulation status of the gas wells in the field.
[0018] Step S3: Establish a gas well fluid accumulation diagnosis model driven by both physical constraints and multi-source data; See Figure 2 The gas well fluid accumulation diagnostic model includes: The input layer is used to receive preprocessed multi-source dynamic production data and static parameters, perform tensor fusion, construct input features that the model can recognize, and simultaneously input the input features into the multi-source data-driven branch and the physical constraint-driven branch. The formula for the input layer is shown below: In the formula, X represents the final input model fusion; Concat() is the tensor concatenation operation; For multi-source dynamic production data tensors; T For time step; C For dynamic parameter dimensions; It is a static parameter tensor; K For static parameter dimensions; Let X be a two-dimensional tensor in the real number space, with dimensions of T rows × (C+K) columns; The output layer is used to output the model prediction results as liquid level labels. Specifically, it outputs the category index corresponding to the maximum probability in the probability distribution obtained by the softmax layer as the liquid level label. The formula for the output layer is shown below: In the formula, Label is the final predicted effusion level label; argmax is the category index corresponding to the maximum value of the array; c represents the true value corresponding to the effusion category, and the value of c is 0, 1, 2, and 3, representing no effusion, mild effusion, moderate effusion, and severe effusion, respectively; Y pred [c] represents the probability of the fluid accumulation category; Multi-source data-driven branches are used to extract time-series features; Among them, the multi-source data-driven branch includes a multi-source temporal feature extraction submodule composed of a MultiROCKET network, an anomaly feature recognition submodule composed of a Transformer network, a fully connected layer, and a Softmax layer; The multi-source temporal feature extraction submodule is used to perform high-dimensional feature extraction on the input features constructed in the input layer, and to obtain multi-source temporal features including latent sequence patterns and preliminary anomaly information. The formula representation of the multi-source temporal feature extraction submodule is shown in the following formula: F rocket The output of MultiROCKET is a high-dimensional feature vector; Stack() is a stacking operation that stacks the pooling results of convolutional kernels of all scales into a high-dimensional feature vector; max() is a global max pooling operation; t is the current time step; W s Let X be the s-th random convolution kernel, where s = 1, 2, 3, ..., S, and S is the total number of multi-scale convolution kernels; * represents the convolution operation; X dyn,t For dynamic data at time step t, a feature slice is generated. b s M is the convolution bias term; M is the feature dimension; F represents rocket It is an M-dimensional real vector; The multi-source temporal feature extraction submodule uses the MultiROCKET network to extract high-dimensional features from multi-source dynamic production data through multi-scale random convolution kernels, in order to capture the potential patterns and preliminary abnormal information of the sequence and provide input for subsequent abnormal feature identification.
[0019] The anomaly feature recognition submodule is used to perform deep similarity calculations on the original sequence features in the input layer and the multi-source temporal features extracted by the multi-source temporal feature extraction submodule, thereby accurately identifying potential anomalies. The formula for the anomaly feature recognition submodule is shown in the following equation: In the formula, W q F is the projection weight matrix; sim The similarity enhancement feature tensor; CrossAttn() is the cross-attention operation; Q is the query matrix; Softmax() is the normalization operation; The scaling factor; MultiHead() is a multi-head attention operation; h is the number of attention heads; Attn h For self-attention output of a single head; W O For multi-head attention, a weight matrix is constructed; F enc The temporal dependency augmentation feature tensor output by the Transformer network; FFN() is a feedforward network; F attn The feature tensor output by the multi-head attention operation; F represents sim It is T time steps × d k A two-dimensional tensor with dimensional features; The anomaly feature recognition submodule uses a Transformer network to perform deep similarity calculations on the original sequence features and multi-source temporal features, achieve cross-validation, enhance the expression of anomalous signals, and model the temporal dependence of sequences to achieve fine recognition of potential anomalies.
[0020] Fully connected layers are used for feature fusion and dimension mapping. The main process involves converting the high-dimensional features F output by the MultiRocket network into dimensional features. rocket Temporal augmentation features F from the output of the Transformer network enc The two are concatenated and fused into a more representative feature vector F using a weight matrix W1 and the activation function ReLU. fusion Meanwhile, on the multi-source data-driven branch, the feature dimensions are mapped to dimensions that adapt to the subsequent Softmax layer.
[0021] The Softmax layer is used for probability normalization, and its main process is to normalize the feature vector F output by the fully connected layer. fusion This is transformed into a probability distribution corresponding to the "liquidity level category" (ensuring that the sum of the probabilities of all categories is 1), so that the subsequent output layer can select the category with the highest probability as the prediction result.
[0022] The formulas for the fully connected layer and the Softmax layer are shown below: In the formula, ReLU() is the activation function; W1 and W2 are the weight matrices of the fully connected layer; b 1. b 2 represents the bias term for the fully connected layer; Softmax() is the normalization function; F fusion The output of Y is a high-dimensional feature vector that fuses MultiRocket and Transformer features in a fully connected layer; pred Y is the probability distribution vector of the fluid level category output by the Softmax layer; pred It is a vector in 4-dimensional real space, and the vector contains 4 real elements that correspond to the prediction probabilities of the four types of liquid accumulation, respectively. By fusing the high-dimensional features extracted by the MultiROCKET network with the feature representations output by the Transformer network, the sequence features are mapped into liquid level vectors through fully connected layers and Softmax layers, providing input for the subsequent physical constraint embedding module.
[0023] Physical constraint-driven branching utilizes the input features constructed from the input layer and the feature vector F obtained from the fully connected layer. fusion Calculate the physical constraint loss of the model.
[0024] Step S4: Train the gas well fluid accumulation diagnostic model using the training set and validation set, and optimize the model using the test set; The training method for the gas well fluid accumulation diagnosis model is a two-stage training. Specifically, firstly, a multi-source data-driven branch is used to obtain stable temporal features as intermediate representations of the fluid accumulation level label. Then, a physical constraint-driven branch is introduced for joint training, with the goal of minimizing the total loss function. This achieves bidirectional constraints of data features and physical laws. During the training process, an adaptive optimizer and gradient clipping are used to avoid gradient anomalies, and an early stopping mechanism is set to prevent overfitting. The total loss function during joint training is shown in the following equation: L total = λ data (t)L data +λ phys (t)L phys In the formula, L total Total loss; L data For data-driven loss; λ data (t) represents the dynamic weights of the data-driven loss; L phys For physical constraint loss; λ phys (t) represents the dynamic weight of the physical constraint loss; Data-driven loss L data The FocalLoss loss function is used to reduce the impact of class imbalance on model training, as shown in the following formula: In the formula, L focal Here, FocalLoss is the loss function; N is the total number of samples; C is the total number of classes; i represents the current sample; α c The balance weight coefficient for the fluid accumulation sample category corresponding to the true value c; Let be the probability that the i-th sample is predicted to be the fluid sample category corresponding to the true value c; y i,c Let i be the true label of the i-th sample; Dynamic weights λ of data-driven loss data (t) satisfies the following equation: λ data (t) = 1-λ phys (t); The physical constraint-driven branch is based on Physics-Informed Neural Networks (PINN), which uses the feature vector F obtained from the fully connected layer. fusion Multiple sets of tensor mappings are used to substitute specific values into the constraint equations to calculate physical losses. The constraint equations include the Beggs-Brill multiphase flow equation, which reflects gas-liquid two-phase flow, and the reservoir Darcy flow equation, which is introduced at the bottom boundary conditions to reflect the formation gas supply capacity. The Beggs-Brill multiphase flow equation for gas-liquid two-phase flow is shown below: In the formula, The pressure drop gradient is calculated based on the Beggs-Brill method; P is the wellbore pressure; z is the wellbore axial coordinate. ρ TP The density of the two-phase mixture; g It is the acceleration due to weight; The wellbore inclination angle; The coefficient of friction between the two phases; D The diameter of the wellbore; V m Mixing speed; The rate of change of mixing velocity along the wellbore axis; Two-phase mixing density ρ TP Satisfy the following formula: ρ TP = ρ L H L + ρ g (1- H L ) In the formula, ρ L The density of the liquid phase fluid; H L Liquid holdup; ρ g The density of the gas phase fluid; The Darcy equation for reservoir flow is shown below: In the formula, v is the seepage velocity; The absolute permeability of the reservoir; μ For fluid viscosity; p ρ is the reservoir pressure; r is the radial coordinate. Physical constraint loss L phys It consists of the Beggs-Brill multiphase flow constraint loss function and the reservoir Darcy flow constraint loss function, as shown in the following equation: L phys = L BB + L Darcy In the formula, L BB L is the Beggs-Brill multiphase flow constraint loss function; Darcy The Darcy flow constraint loss function; The Beggs-Brill multiphase flow constraint loss function is shown in the following equation: In the formula, N bb This represents the number of spatial sampling points; , , The neural network at the 1st The wellbore pressure, liquid holdup, and mixing velocity predicted at each spatial location are represented by the feature vector F output from the fully connected layer. fusion The tensor mapping in the middle is obtained; The Darcy flow constraint loss function is shown in the following formula: In the formula, N d This represents the number of time sampling points; , The first prediction of the neural network j The seepage velocity and reservoir pressure at each time point, and the feature vector F output by the fully connected layer. fusion The tensor mapping in is obtained, μ For fluid viscosity, r Radial coordinates, The absolute permeability of the reservoir; Dynamic weights λ of physical loss phys (t) satisfies the following formula: In the formula, clip() is a range cutoff function used to restrict the input value to [λ]. min , λ max Within the range; β is the steepness coefficient of the sigmoid function, used to control the sensitivity of the physical loss weights to dynamic changes; To differentiate between the stable and dynamic periods, the balance between "accuracy of effusion level prediction" and "physical equation residuals" under different thresholds was observed during model training. Ultimately, the threshold that optimizes the model's generalization ability was selected. min λ represents the lower bound constraint value for the physical loss weight. max This represents the upper limit constraint value for the physical loss weight; α ema t represents the EMA decay coefficient; t represents the current time. q g ( t ) represents the wellhead gas volumetric flow rate at time t; Δ t For time step.
[0025] The specific method for optimizing the model is as follows: After training, the model is independently evaluated using a test set, calculating precision, recall, and F1 score to verify the model's generalization ability and stability. If the calculation results do not meet the preset performance threshold, the gas well fluid accumulation diagnosis model is retrained until it meets the threshold. The performance threshold can be set independently based on business needs, data baseline, and validation results. Simultaneously, during retraining, if it is found that the performance failure to meet the threshold is due to the obtained total loss function, the data-driven loss and physical constraint loss need to be optimized to rebalance their contributions to the total loss function. If the calculation results meet the performance threshold, the model weights with the best performance are saved, and the model at this time is output as the final gas well fluid accumulation diagnosis model.
[0026] Step S5: Deploy the optimized model to the field for real-time diagnosis of fluid accumulation in gas wells, mainly including the following steps: (1) The gas well liquid accumulation diagnosis model that has been trained is subjected to lightweight processing. Knowledge distillation and quantization pruning methods are used to reduce the model parameter scale and computational complexity so that it can adapt to the computing resource limitations of the field edge computing equipment. (2) Deploy the lightweight model to a field control terminal with embedded computing capabilities to realize real-time input of sensor data from any wellhead within the gas field and online diagnosis of the liquid accumulation status; (3) During the model operation, sensor data streams are collected in real time, including dynamic parameters such as gas production, water production, wellhead pressure and wellhead temperature, and input features are constructed in combination with static parameters to maintain consistency with the data distribution during training. (4) The control terminal performs real-time calculations on the input data through the diagnostic module, outputs the liquid accumulation level at the corresponding time, and sends the diagnostic results back to the ground monitoring system to assist in the adjustment of the drainage strategy; (5) Perform model calibration and database updates regularly, maintain the static parameter table, and if the system detects a decrease in model diagnostic accuracy or an over-limit of on-site computing resources, it will automatically return to step S4 to retrain and optimize the model in order to ensure the accuracy of diagnostic results and operational stability.
[0027] Example: Using 20 shale gas wells from a shale gas field in Sichuan Province, my country as the research object, the effectiveness of the method of the present invention in real-time diagnosis of fluid accumulation in gas wells was verified.
[0028] First, multi-source dynamic production data (including gas production, water production, oil pressure, casing pressure, transmission pressure, and wellhead temperature) from 20 wells in the gas field were collected from June 2023 to June 2024 at a sampling frequency of once per minute, as well as the required static parameters (including wellbore geometry parameters, fluid properties parameters, and reservoir properties parameters). The initial dataset was then preprocessed according to the process described in step S1 as follows: outliers were removed using the 3-sigma criterion; missing values were specifically filled using cubic spline interpolation and nearest neighbor interpolation; wavelet transform and LOWESS smoothing were combined to smooth and denoise the time-series dynamic production data; min-max normalization was performed on the smoothed data; and Z-score standardization was applied to the static parameters to complete the data preprocessing.
[0029] Next, leveraging the extensive experience of domain experts and combining it with dynamic production data preprocessed using S1, a sliding window technique was employed, with a 30-minute window size, to assign corresponding effusion level labels to the data at each time point. The labels were categorized into four levels: no effusion, mild effusion, moderate effusion, and severe effusion. Through this process, a dataset containing 186,405 valid samples was constructed and divided into training, validation, and test sets in a 7:2:1 ratio. For the training set, the SMOTE algorithm was used for oversampling and random undersampling to achieve class balance, while the validation and test sets retained their original distributions.
[0030] Subsequently, a dual-branch fusion model of "physical constraint-driven + multi-source data-driven" was constructed. The multi-source data-driven branch consists of a MultiRocket network, a Transformer network, and fully connected and Softmax layers. The physical constraint-driven branch embeds the Beggs-Brill multiphase flow equation and the reservoir Darcy flow equation. A dynamic weight balancing method is used to integrate the dual-branch losses to obtain the total loss function, which is the weighted sum of the data loss and the physical loss. The model was trained and optimized using the training and validation sets, employing the Adam optimizer with an initial learning rate of 0.001, decaying by 10% every 50 epochs, for 200 epochs, a batch size of 64, and early stopping enabled. The model performance was evaluated on the test set, and the experimental results are as follows: Figure 3 As shown in the confusion matrix, the model performs excellently in the gas well fluid accumulation status identification task, with an overall accuracy of 84.94%. Particularly noteworthy is the model's outstanding performance in identifying both heavily fluidized and non-fluidized conditions. The recall rate for heavily fluidized conditions reached 93.33%, and the F1 score was as high as 89.34%, indicating that the model can effectively capture severe fluid accumulation situations, providing reliable assurance for engineering safety. Meanwhile, the accuracy for non-fluidized conditions reached 100%, meaning that all wellbores judged as non-fluidized were indeed in a truly non-fluidized state, effectively avoiding false alarms.
[0031] Finally, the final model obtained from the experiment was packaged into a lightweight algorithm module and deployed to the field. Real-time diagnosis was then performed on a gas well in the gas field, and the real-time diagnosis results are as follows: Figure 4 As shown in the figure, the model can continuously and stably output predictions of fluid accumulation status, promptly capturing the transition from "no fluid accumulation" to "slight fluid accumulation." When the wellbore shows signs of fluid accumulation risk, the system issues an early warning, providing a valuable time window for on-site personnel to take drainage and gas production measures. Practical application demonstrates that the diagnostic system responds quickly, and the prediction results highly match the actual on-site conditions, effectively supporting the refined management of gas wells.
[0032] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the embodiments of the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for diagnosing fluid accumulation in gas wells driven by both physical constraints and multi-source data, characterized in that, Includes the following steps: Step S1: Collect multi-source dynamic production data and static parameters of gas wells as the initial dataset and perform preprocessing; Step S2: Set a time window, extract time series data from each target time point in the multi-source dynamic production data, assign liquid accumulation level labels to the time series data, obtain a sample set, and divide the sample set into a training set, a validation set, and a test set; Step S3: Establish a gas well fluid accumulation diagnosis model driven by both physical constraints and multi-source data; Step S4: Train the gas well fluid accumulation diagnostic model using the training set and validation set, and optimize the model using the test set; Step S5: Deploy the optimized model to the field for real-time diagnosis of fluid accumulation in gas wells; The multi-source dynamic production data includes gas production, water production, oil pressure, casing pressure, transmission pressure, and wellhead temperature; the static parameters include wellbore geometric parameters, fluid properties, and reservoir properties, wherein the wellbore geometric parameters include wellbore inclination angle and wellbore diameter; the fluid properties include liquid phase density, gas phase density, and fluid viscosity; and the reservoir properties include absolute permeability. The gas well fluid accumulation diagnostic model mentioned in step S3 includes: The input layer is used to receive preprocessed multi-source dynamic production data and static parameters, perform tensor fusion, construct input features that the model can recognize, and simultaneously input the input features into the multi-source data-driven branch and the physical constraint-driven branch. The formula for the input layer is shown below: In the formula, X represents the final input model fusion; Concat() is the tensor concatenation operation; For multi-source dynamic production data tensors; T For time step; C For dynamic parameter dimensions; It is a static parameter tensor; K For static parameter dimensions; Let X be a two-dimensional tensor in the real number space, with dimensions of T rows × (C+K) columns; The output layer is used to output the model prediction results as liquid level labels; The formula for the output layer is shown below: In the formula, Label is the final predicted effusion level label; argmax is the category index corresponding to the maximum value of the array; c represents the true value corresponding to the effusion category, and the value of c is 0, 1, 2, and 3, representing no effusion, mild effusion, moderate effusion, and severe effusion, respectively; Y pred [c] represents the probability of the fluid accumulation category; Multi-source data-driven branches are used to extract time-series features; Among them, the multi-source data-driven branch includes a multi-source temporal feature extraction submodule composed of a MultiROCKET network, an anomaly feature recognition submodule composed of a Transformer network, a fully connected layer, and a Softmax layer; The multi-source temporal feature extraction submodule is used to perform high-dimensional feature extraction on the input features constructed in the input layer, and to obtain multi-source temporal features including latent sequence patterns and preliminary anomaly information. The formula representation of the multi-source temporal feature extraction submodule is shown in the following formula: F rocket The output of MultiROCKET is a high-dimensional feature vector; Stack() is a stacking operation that stacks the pooling results of convolutional kernels of all scales into a high-dimensional feature vector; max() is a global max pooling operation; t is the current time step; W s Let X be the s-th random convolution kernel, where s = 1, 2, 3, ..., S, and S is the total number of multi-scale convolution kernels; * represents the convolution operation; X dyn,t For dynamic data at time step t, a feature slice is generated. b s M is the convolution bias term; M is the feature dimension; F represents rocket It is an M-dimensional real vector; The anomaly feature recognition submodule is used to perform deep similarity calculations on the original sequence features in the input layer and the multi-source temporal features extracted by the multi-source temporal feature extraction submodule, thereby accurately identifying potential anomalies. The formula for the anomaly feature recognition submodule is shown in the following equation: In the formula, W q F is the projection weight matrix; sim The similarity enhancement feature tensor; CrossAttn() is the cross-attention operation; Q is the query matrix; Softmax() is the normalization operation; The scaling factor; MultiHead() is a multi-head attention operation; h is the number of attention heads; Attn h For self-attention output of a single head; W O For multi-head attention, a weight matrix is constructed; F enc The temporal dependency augmentation feature tensor output by the Transformer network; FFN() is a feedforward network; F attn The feature tensor output by the multi-head attention operation; F represents sim It is T time steps × d k A two-dimensional tensor with dimensional features; Fully connected layers are used for feature fusion and dimension mapping, converting F... rocket With F enc The features are concatenated and merged into a more representative feature vector F. fusion ; The Softmax layer is used for probability normalization, transforming the feature vector F output by the fully connected layer... fusion The probability distribution is converted into the corresponding effusion level category; The formulas for the fully connected layer and the Softmax layer are shown below: In the formula, ReLU() is the activation function; W1 and W2 are the weight matrices of the fully connected layer; b 1. b 2 represents the bias term for the fully connected layer; Softmax() is the normalization function; F fusion The output of Y is a high-dimensional feature vector that fuses MultiRocket and Transformer features in a fully connected layer; pred Y is the probability distribution vector of the fluid level category output by the Softmax layer; pred It is a vector in 4-dimensional real space, and the vector contains 4 real elements that correspond to the prediction probabilities of the four types of liquid accumulation, respectively. Physical constraint-driven branching utilizes the input features constructed from the input layer and the feature vector F obtained from the fully connected layer. fusion Calculate the physical constraint loss of the model.
2. The gas well fluid accumulation diagnosis method driven by both physical constraints and multi-source data according to claim 1, characterized in that: The preprocessing method described in step S1 includes the following steps: Step S11: Detect and remove outliers in the initial dataset using the 3-sigma criterion; Step S12: Use cubic spline interpolation to complete the multi-source dynamic production data, and use nearest neighbor interpolation to complete the static parameters; Step S13: Use wavelet transform to filter out high-frequency noise in the completed multi-source dynamic production data, and use Lowess smoothing to correct local outliers and stabilize the overall trend in the completed multi-source dynamic production data. Step S14: Perform min-max normalization on the multi-source dynamic production data to map the data to the [0,1] interval; perform Z-score standardization on the static parameters to transform them into a distribution with a mean of 0 and a standard deviation of 1.
3. The gas well fluid accumulation diagnosis method driven by both physical constraints and multi-source data according to claim 1, characterized in that: The fluid accumulation level label mentioned in step S2 includes four levels, specifically: Level 0 corresponds to no fluid accumulation; Grade 1 corresponds to mild effusion; Level 2 corresponds to moderate effusion; Level 3 corresponds to severe effusion.
4. The gas well fluid accumulation diagnosis method driven by both physical constraints and multi-source data according to claim 1, characterized in that: The method for dividing the training set, validation set, and test set in step S2 is as follows: under the condition that the time periods do not overlap, the sample set is divided into the training set, validation set, and test set in a 7:2:1 ratio according to the time sequence.
5. The gas well fluid accumulation diagnosis method driven by both physical constraints and multi-source data according to claim 3, characterized in that: The training set mentioned in step S2 is subjected to a combination of SMOTE oversampling and random undersampling for equalization processing, including the following steps: using the SMOTE algorithm to generate synthetic samples for severe and moderate effusion samples to supplement the sample quantity; randomly sampling samples without effusion and with mild effusion samples, so that the difference between their quantity and the quantity of severe and moderate effusion samples after supplementing the sample quantity does not exceed 10%.
6. The gas well fluid accumulation diagnosis method driven by both physical constraints and multi-source data according to claim 1, characterized in that: The training method for the gas well fluid accumulation diagnostic model in step S4 is a two-stage training. Specifically, the method is as follows: first, multi-source data is used to drive branches to obtain stable time-series features, and then physical constraints are introduced to drive branches for joint training. During the training process, an adaptive optimizer and gradient clipping are used to avoid gradient anomalies, and an early stopping mechanism is set to prevent overfitting. The total loss function during joint training is shown in the following equation: L total = λ data (t)L data +λ phys (t)L phys In the formula, L total Total loss; L data For data-driven loss; λ data (t) represents the dynamic weights of the data-driven loss; L phys For physical constraint loss; λ phys (t) represents the dynamic weight of the physical constraint loss; Data-driven loss L data The loss is calculated using the FocalLoss function, as shown in the following formula: In the formula, L focal Here, FocalLoss is the loss function; N is the total number of samples; C is the total number of classes; i represents the current sample; α c The balance weight coefficient for the fluid accumulation sample category corresponding to the true value c; Let be the probability that the i-th sample is predicted to be the fluid sample category corresponding to the true value c; y i,c Let i be the true label of the i-th sample; Dynamic weights λ of data-driven loss data (t) satisfies the following equation: l data (t) = 1-λ phys (t); The physical constraint-driven branch uses the Beggs-Brill multiphase flow equation for gas-liquid two-phase flow as the main constraint equation, and introduces the reservoir Darcy flow equation at the bottom hole boundary condition. The Beggs-Brill multiphase flow equation for gas-liquid two-phase flow is shown below: In the formula, The pressure drop gradient is calculated based on the Beggs-Brill method; P is the wellbore pressure; z is the wellbore axial coordinate. ρ TP The density of the two-phase mixture; g It is the acceleration due to weight; The wellbore inclination angle; The coefficient of friction between the two phases; D The diameter of the wellbore; V m Mixing speed; The rate of change of mixing velocity along the wellbore axis; Two-phase mixing density ρ TP Satisfy the following formula: ρ TP = ρ L H L + ρ g (1- H L ) In the formula, ρ L The density of the liquid phase fluid; H L Liquid holdup; ρ g The density of the gas phase fluid; Physical constraint loss L phys It consists of the Beggs-Brill multiphase flow constraint loss function and the reservoir Darcy flow constraint loss function, as shown in the following equation: L phys = L BB + L Darcy In the formula, L BB L is the Beggs-Brill multiphase flow constraint loss function; Darcy The Darcy flow constraint loss function; The Beggs-Brill multiphase flow constraint loss function is shown in the following equation: In the formula, N bb This represents the number of spatial sampling points; , , The neural network at the 1st The wellbore pressure, liquid holdup, and mixing velocity predicted at each spatial location are represented by the feature vector F output from the fully connected layer. fusion The tensor mapping in the middle is obtained; The Darcy flow constraint loss function is shown in the following formula: In the formula, N d This represents the number of time sampling points; , The first prediction of the neural network j The seepage velocity and reservoir pressure at each time point, and the feature vector F output by the fully connected layer. fusion The tensor mapping in is obtained, μ For fluid viscosity, r Radial coordinates, This represents the absolute permeability of the reservoir. Dynamic weights λ of physical loss phys (t) satisfies the following formula: In the formula, clip() is a range cutoff function used to restrict the input value to [λ]. min , λ max Within the range; β is the steepness coefficient of the sigmoid function, used to control the sensitivity of the physical loss weights to dynamic changes; A threshold to distinguish between stable periods and periods of dynamic change; λ min λ represents the lower bound constraint value for the physical loss weight. max This represents the upper limit constraint value for the physical loss weight; α ema t represents the EMA decay coefficient; t represents the current time. q g ( t ) represents the wellhead gas volumetric flow rate at time t; Δ t For time step.
7. The gas well fluid accumulation diagnosis method driven by both physical constraints and multi-source data according to claim 1, characterized in that: The specific method for optimizing the model in step S4 is as follows: After training, the model is independently evaluated using the test set, and the accuracy, recall, and F1 score are calculated to verify the generalization ability and stability of the model. If the calculation results do not meet the preset performance threshold, the gas well fluid accumulation diagnosis model is retrained until it meets the threshold. If the calculation results meet the performance threshold, the weights of the model with the best performance are saved, and the model at this time is output as the final gas well fluid accumulation diagnosis model.