Underground oil and gas well fault prediction method based on multi-modal space-time diagram neural network

Through the multimodal spatiotemporal graph neural network, the problems of insufficient coupling of multimodal feature and weak spatial and temporal correlation modeling in traditional oil and gas downhole fault prediction methods are solved, and efficient and accurate fault prediction and early warning of downhole equipment are achieved.

CN120387000APending Publication Date: 2025-07-29CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 0 Cites 20 Cited by

Patent Information

Application Number
CN202510530059.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Traditional oil and gas underground fault prediction methods have problems such as insufficient coupling of multimodal features, weak spatial and temporal correlation modeling, excessive dependence on manual experience and poor real-time performance, making it difficult to achieve efficient and accurate fault prediction.

Method used

The multimodal spatiotemporal graph neural network is used to extract multi-frequency sensor data and maintain log semantic features through the time convolution network (TCN) and BERT model. Combined with dynamic spatiotemporal graph convolution and hierarchical prediction mechanism, the dynamic physical relationships and fault propagation paths of downhole equipment are explicitly modeled to realize the spatiotemporal alignment of sensor signals and text logs and the fusion of multimodal features.

Benefits of technology

It significantly improves the accuracy of downhole fault feature extraction and spatiotemporal correlation modeling capabilities, realizes high-precision early fault detection and residual effective time (RUT) prediction, and improves the real-time and accuracy of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005376388210000041
    Figure BDA0005376388210000041
  • Figure BDA0005376388210000044
    Figure BDA0005376388210000044
  • Figure BDA0005376388210000053
    Figure BDA0005376388210000053
Patent Text Reader

Abstract

The invention discloses an underground oil and gas well fault prediction method based on a multi-modal space-time diagram neural network, and aims to solve the problems of insufficient multi-modal feature coupling, weak space-time correlation modeling, poor real-time performance and the like of a traditional method. The method is characterized in that multi-frequency sensor data features and maintenance log semantic knowledge are respectively extracted through a time expansion convolutional network (TCN) and a BERT-BiLSTM model, and a cross-modal gating attention mechanism is designed to realize heterogeneous data dynamic fusion; an equipment space topological graph is constructed based on Delaunay triangulation, dynamic causal association between nodes is quantized in combination with transfer entropy to generate a time graph, and a fault propagation path is jointly modeled through residual space-time graph convolution; a hierarchical prediction module is constructed by adopting a bidirectional LSTM and a graph attention network (GAT), short-term fault classification and long-term equipment residual life prediction are respectively realized, and the fault prediction precision and industrial landing feasibility under complex working conditions are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - technical field of oil and gas exploration and development technology and artificial intelligence, and particularly relates to a method for predicting oil and gas downhole faults based on a multi - modal spatio - temporal graph neural network. Background Art

[0002] With the continuous development of oil and gas production technology, the operating state of downhole equipment in oil and gas fields is crucial for production efficiency and safety. However, the downhole working conditions are complex and variable, and equipment failures occur frequently. These failures not only lead to production interruptions and economic losses but may also cause safety accidents and environmental pollution. Traditional oil and gas downhole fault prediction methods face three major technical bottlenecks: First, the problem of data heterogeneity. Downhole multi - source data includes time - series sensing data such as high - frequency vibration, low - frequency temperature, pressure, flow, and vibration, as well as unstructured logs of equipment operating states and maintenance records. Due to limited representation capabilities, traditional statistical methods and shallow machine learning are difficult to effectively fuse cross - modal spatio - temporal features, resulting in missed detection of key fault patterns. Second, the lack of spatio - temporal coupling. Although existing LSTM and CNN models can separately process information in the time or space dimension, they cannot model the combined action mechanism of complex spatial topological relationships and time - dynamic propagation effects between equipment, resulting in large prediction errors for progressive faults. Third, the lag in early warning. The rule - based engine based on threshold discrimination has insufficient sensitivity to weak fault features and is difficult to adapt to dynamic changes in working conditions, with a long early - warning response delay.

[0003] In recent years, the graph neural network (GNN) that has been adopted can model the spatial dependence relationship of equipment, but ignores the dynamic characteristics in the time dimension; introducing the attention mechanism to achieve multi - modal fusion, but the spatio - temporal alignment processing for heterogeneous sampling frequencies is not sufficient; constructing a spatio - temporal graph model can jointly analyze the spatio - temporal dimension, but it is not optimized for the physical hierarchical relationship of downhole multi - source heterogeneous data and the differences in sensor noise, resulting in poor model generalization and low edge - deployment efficiency, and it is difficult to meet the real - time requirements. Therefore, there is an urgent need to develop a new fault prediction method that can break through the above - mentioned technical bottlenecks to achieve efficient and accurate prediction of oil and gas downhole equipment failures and ensure the safety and economy of oil and gas production. Summary of the Invention

[0004] In order to overcome the problems of insufficient coupling of multi - modal features, weak spatio - temporal correlation modeling, over - reliance on manual experience, and poor real - time performance in traditional oil and gas downhole fault prediction methods, the present invention provides a method for predicting oil and gas downhole faults based on a multi - modal spatio - temporal graph neural network. By constructing a multi - modal spatio - temporal graph neural network, it explicitly models the dynamic physical relationship and fault propagation path of downhole equipment, and realizes collaborative perception and accurate prediction of fault features.

[0005] To achieve the above object, the technical solution of the present invention mainly includes the following steps:

[0006] A. Multi-modal Feature Coupling and Alignment:

[0007] (1) Extract multi-scale temporal features of multi-source sensor data such as pressure, temperature, and vibration acceleration through a Temporal Convolutional Network (TCN), adopt dilated convolutional layers and causal convolution constraints, and use cubic spline interpolation to align low-frequency signals.

[0008] (2) Extract semantic features of maintenance log texts based on the BERT pre-trained model, combine the BiLSTM-CRF model to identify device entities and fault types, construct a knowledge graph, and generate graph embedding vectors.

[0009] (3) Design an attention gating mechanism to align sensor features and text features, and achieve spatio-temporal adaptive fusion of multi-modal features through dynamic weight allocation.

[0010] B. Construction of Dynamic Spatio-temporal Graph Neural Network:

[0011] (1) Construct a spatial adjacency matrix based on the physical connections of devices and 3D Delaunay triangulation, and use a normalized distance decay function for weight calculation.

[0012] (2) Quantify the information flow between nodes based on Transfer Entropy, generate a dynamic temporal adjacency matrix through sigmoid activation, and update it every 5 minutes.

[0013] (3) Design a graph convolution formula with residual connections, fuse the topological relationships of spatial graphs and temporal graphs, and use the GELU activation function to enhance the non-linear expression ability.

[0014] C. Optimization of Fault Prediction Decoder:

[0015] (1) Extract local temporal patterns through bidirectional LSTM and temporal attention pooling, and output the probability distribution of 8 types of faults in the short term.

[0016] (2) Aggregate global spatio-temporal features based on the Graph Attention Network (GAT), and predict the Remaining Useful Time (RUT) of the device and its confidence interval.

[0017] (3) Combine cross-entropy loss, Dynamic Time Warping loss (DTW), and temporal smoothing regularization terms, and use a staged curriculum learning strategy to train the model.

[0018] The beneficial effects of the present invention are as follows: By fusing downhole sensor data and maintenance log knowledge through a multi-modal spatio-temporal graph neural network, combined with dynamic graph convolution and a hierarchical prediction mechanism, the problems of insufficient fault feature extraction and weak spatio-temporal correlation modeling in traditional methods under complex working conditions are effectively solved. Among them, the collaborative perception of multi-modal uses cross-modal gated fusion of a temporal convolutional network (TCN) and a BERT model to achieve spatio-temporal alignment of sensor signals and text logs, significantly reducing the feature coupling error; using dynamic causal reasoning, based on the dynamic update mechanism of Delaunay spatial graphs and transfer entropy temporal graphs, explicitly models the propagation path of faults in the device topology, greatly improving the causal reasoning ability; in addition, through a hierarchical decoder of bidirectional LSTM short-term classification and GAT long-term life prediction, high-precision early fault detection and remaining useful time (RUT) prediction are achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a flowchart of the method of the present invention DETAILED DESCRIPTION OF THE INVENTION

[0020] The following is a more detailed description of the present invention in conjunction with Figure 1 :

[0021] A. Multi-modal Feature Coupling and Alignment:

[0022] (1) Collect the time-series data collected by downhole sensors, including: pressure P t ∈R T , temperature T t ∈R T , vibration acceleration A t ∈R T . Set the time window T = 3600 and the sampling frequency 1Hz, and perform structured data encoding and feature alignment.

[0023] (2) Adopt a 4-layer dilated convolution structure with dilation coefficients [1, 2, 4, 8] in sequence. The convolution kernel size k = 5 for each layer, and the output channels C = 64. After each layer, batch normalization (BatchNorm) and GELU activation function are connected to generate multi-scale time-series features F s ∈R T×64 . At the same time, only historical data is allowed to participate in the current moment calculation through zero-padding to avoid future information leakage.

[0024] (3) For the low-frequency signal: the flow rate Q t with a sampling rate of 0.2Hz, use the cubic spline interpolation function:

[0025]

[0026] where B i (t) is the cubic B-spline basis function, ti As the original sampling points, align them to 1 Hz to ensure the consistency in the time dimension.

[0027] (4) For the vibration signal A t Perform 5-layer decomposition using the Daubechies-4 wavelet function, retain the detail coefficients of the 2nd - 4th layers, and the corresponding frequency range is [0.1, 50] Hz, so as to remove high-frequency noise.

[0028] (5) Encode and perform knowledge modeling on the unstructured data. For the input device maintenance log text, first perform text feature extraction, and extract the sentence vector v t ∈R 768 , for the text sequence {w1, w2, …, w n} perform average pooling operation: At the same time, use the BiLSTM-CRF model to identify entities, and identify device components "pump body", "valve", "bearing" and failure modes "temperature overrun", "leakage", "wear".

[0029] (6) Subsequently, carry out the construction of the knowledge graph, define the triple containing the device, failure mode, and frequency. The calculation formula of the frequency is: Where e1 refers to the starting entity device, e2 refers to the ending entity failure mode, e′ refers to all possible ending entities, and count(e1→e2) represents the number of instances where the device e1 has the failure mode e2. Then store it as a graph structure G k =(V k , E k ), where the node V k is the device or failure type, and the edge E k represents the association strength.

[0030] (7) Fuse the cross-modal features, and use the gated alignment mechanism to calculate the cross-modal attention for the sensor feature F s ∈R T×64 and the text feature v t ∈R 768 . First, generate query matrices for F s and v t respectively through linear transformation: Q = F s W q ∈R T×64 , key matrix: K = v t W k ∈R 1×64 , where W q , W k ∈R 64×64 are learnable parameters. The specific formula for calculating the attention weight is as follows:

[0031]

[0032] Among them, d = 64 is the scaling factor, which aggregates text features weighted by the time dimension, and finally outputs the alignment feature F a = α t v t W v ∈ R T×64 , W v ∈ R 768×64 is a learnable parameter matrix used to linearly transform the text feature v t from its original 768 dimensions to a new feature space.

[0033] (8) Next, perform dynamic weight allocation, and calculate the modality weight through a fully connected layer: α = σ(W m [F s ; F a + b m )R T ×1 , where W m ∈ R 128×1 , b m ∈ R is the bias term, σ is the sigmoid function, [;] represents the concatenation operation, and finally the fused feature h m = β ⊙ F s + (1 - β) ⊙ F a ∈ R T×64 , where F s is the sensor feature, F a is the alignment feature, h m is the fused feature, and ⊙ represents element-wise multiplication.

[0034] B. Construction of dynamic spatio-temporal graph neural network:

[0035] (1) Generate a spatial graph. Define a total of N = 120 nodes for underground equipment and sensors, and the three-dimensional coordinates of the nodes are P i = (x i , y i , z i ) ∈ R 3 .

[0036] (2) Construct the spatial adjacency matrix A s : If there is a direct connection of pipes or cables between nodes i and j, then the spatial adjacency matrix A s (i, j) = 1; for non-connected nodes, it is necessary to calculate the three-dimensional Delaunay triangulation to generate topological connection edges, and the weight calculation formula is: where σ s = 50 is the spatial attenuation coefficient, which controls the attenuation rate of the weight with distance.

[0037] (3) Perform symmetric normalization on the spatial adjacency matrix: where D s is the degree matrix, and D s (i,i) = ∑ j A s (i,j).

[0038] (4) Generate a temporal graph, and use transfer entropy to quantify the information transfer amount from node i to node j based on historical data. The specific formula is as follows:

[0039]

[0040] The temporal edge weights are activated through the sigmoid function: where γ = 0.5 is the activation threshold, and the temporal adjacency matrix A t is updated every 5 minutes.

[0041] (5) In the spatio-temporal convolution operation, adopt the dynamic graph convolution formula:

[0042]

[0043] where is the normalized spatial graph, is the temporal graph, W s , W t R d×d (d = 128 is the feature dimension) are learnable parameters, and H (l) represents the feature matrix of the l-th layer. By multiplying the normalized adjacency matrices of the spatial graph and the temporal graph with the feature matrix respectively and performing a linear transformation, and then passing through the GELU activation function, the feature matrix H (l+1) of the next layer is obtained.

[0044] (6) Perform a residual connection by stacking the original input at each layer output: H (l+1) = H (l) + Dropout(H (l) ), where the dropout rate is set to 0.2 to prevent overfitting.

[0045] C. Fault prediction decoder optimization:

[0046] (1) In the case of short-term prediction (0 - 6 hours), use the multi-modal fusion feature H t ∈ R 3600×128 of the most recent 1 hour as the input, and utilize a bidirectional LSTM. Set the hidden layer dimension h = 256 and the time step T = 3600 to output the hidden state sequence to capture local temporal patterns and calculate the attention weights at each time step:

[0047]

[0048] where \(W\) a \(\in\mathbb{R}\) 256×256 , \(v\) a \(\in\mathbb{R}\) 256 are learnable parameters, and the aggregated feature is output Then, through a fully connected layer and the softmax function, the fault probability distribution \(p\) short \(\in[0,1]\) K is generated, where \(K = 8\) is the number of fault types

[0049] (2) In the case of long-term prediction (24 - 72 hours), using the spatio-temporal graph node feature \(H\) g \(\in\mathbb{R}\) N×128 as the input, a graph attention network (GAT) is used to aggregate node information. First, the calculation of multi-head attention is performed, and the number of heads \(h = 4\). The specific formula is as follows

[0050]

[0051] where \(W\) g \(\in\mathbb{R}\) 128×32 , \(a\in\mathbb{R}\) 64 are learnable parameters, and \(\|\) represents the concatenation operation. Then, the attention coefficients are normalized using the softmax function

[0052] \(a\) ij \(=\text{softmax}(e\) ij )

[0053] Finally, the aggregated feature is output where \(\sigma\) is the ELU activation function

[0054] Then, the global dependence of the time series features is modeled through multi-head self-attention to generate the remaining useful time (RUT) prediction value and the confidence interval \(\sigma\) RUT .

[0055] (3) Loss function and training strategy A composite loss function is adopted, and the relevant calculation formula is as follows

[0056] \(L = 0.6\cdot L\) CE + 0.4\cdot L\) DTW + 0.1\cdot L\) Reg

[0057]

[0058] where the cross-entropy loss \(L\) CE is used for the classification task, \(\lambda\) is the regularization coefficient, and the dynamic time warping loss \(L\) DTWFor constraining the alignment of the predicted curve with the true fault evolution time series, π ij is the optimal alignment path matrix, L Reg is the time series smoothing regularization term.

[0059] (4) The training strategy is divided into two stages: The first stage (0 - 100 epochs): Only train the spatial graph branch, set the learning rate to 3×10 -4 , freeze the time graph parameters; The second stage (101 - 200 epochs): Jointly optimize the spatio-temporal graph, linearly decay the learning rate to 1×10 -4 , unfreeze the time graph parameters; The third stage (201 - 300 epochs): Fine-tune the decoder parameters, fix the learning rate to 5×10 -5 .

[0060] The above embodiments are only used to illustrate the technical solutions of the present invention. Any person skilled in the art may modify or change the above-described technical solutions into equivalent examples of equivalent changes. Any simple modification, change, or modification made to the above embodiments based on the technical solutions of the invention without departing from the content of the technical solutions of the invention shall fall within the protection scope of the technical solutions of the invention.

Claims

1. A method for predicting oil and gas downhole faults based on a multi-modal spatio-temporal graph neural network, characterized in that, It includes the following steps: A. Multimodal feature fusion processing: Collect the time-series data collected by downhole sensors, including pressure, temperature, and vibration acceleration. Extract the multi-scale time-series features of multi-source sensor data through a Temporal Convolutional Network (TCN), adopt dilated convolutional layers and causal convolution constraints, and align low-frequency signals through cubic spline interpolation. Extract the semantic features of maintenance log texts based on the BERT pre-trained model, identify device entities and fault types in combination with the BiLSTM-CRF model, construct a knowledge graph and generate graph embedding vectors, design a cross-modal attention gating mechanism, dynamically allocate the fusion weights of sensor features and text features, and achieve spatio-temporal adaptive alignment of heterogeneous data; B. Construction of a dynamic spatio-temporal graph neural network: Based on the three-dimensional physical coordinates of downhole equipment, use Delaunay triangulation to generate a spatial adjacency matrix, and the weight calculation uses a normalized distance attenuation function. Quantify the information transfer intensity between nodes based on Transfer Entropy, generate a dynamic time adjacency matrix through sigmoid activation, update it every 5 minutes, construct a spatio-temporal graph convolutional module with residual connections, combine the topological relationships of the spatial graph and the time graph, and adopt the GELU activation function to enhance the non-linear expression ability; C. Optimization of the fault prediction decoder: Extract short-term fault time-series features through a bidirectional LSTM network, combine with a time attention pooling layer to output the probability distribution of 8 types of faults in the case of short-term prediction, aggregate global spatio-temporal features based on the Graph Attention Network (GAT), predict the remaining useful time (RUT) of the equipment and its 95% confidence interval in the long term, combine the cross-entropy loss, the dynamic time warping loss (DTW), and the time-series smoothing regularization term, and train the model using a staged curriculum learning strategy.

Citation Information

Cited By

  • Mine high-temperature heat damage monitoring system based on multi-mode Internet of Things data

    CN120560049A

  • Mine high temperature heat disaster monitoring system based on multi-modal internet of things data

    CN120560049B

  • Equipment fault prediction method fusing physical constraint and adversarial network

    CN120744639A

  • Device fault prediction method fusing physical constraints and adversarial network

    CN120744639B

  • Fault diagnosis agent data preprocessing method and system based on multi-modal alignment

    CN120892239A