Natural gas pipeline leakage detection method based on ResNet50 and SMOTE

By using a method based on ResNet50 and SMOTE, and utilizing infrared image preprocessing and feature extraction, the accuracy problem of natural gas pipeline leak detection was solved, and efficient natural gas leak detection was achieved.

CN120655602APending Publication Date: 2025-09-16CHINA JILIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510753354.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively detect natural gas pipeline leaks, especially in complex transportation environments where there are huge safety risks.

Method used

A natural gas pipeline leak detection method based on ResNet50 and SMOTE is adopted. Infrared images are obtained for preprocessing, and the RepVGG module is used for feature extraction. The SMOTE algorithm is combined to process unbalanced data sets to improve the feature extraction and learning capabilities of the model.

Benefits of technology

The accuracy of natural gas pipeline leak detection has been improved. Experimental results reached 99.83% and 99.80% on the GasVid and IOD-Video datasets, respectively, outperforming traditional models and existing detection algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655602A_ABST
    Figure CN120655602A_ABST
Patent Text Reader

Abstract

The invention provides a natural gas pipeline leakage detection method based on ResNet50 and SMOTE. The natural gas pipeline leakage detection method comprises the following steps: acquiring an infrared image of a natural gas pipeline; preprocessing the infrared image, including denoising the infrared image, and specifically including: performing gray processing on the infrared image to obtain a grayed image; carrying out filtering processing on the grayed image; and detecting the preprocessed image, performing feature extraction on the preprocessed image, and determining whether the natural gas pipeline leaks or not based on the extracted features. According to the natural gas pipeline leakage detection method based on ResNet50 and SMOTE provided by the invention, the natural gas leakage feature extraction, learning and fusion capabilities of the model are improved. Experimental results show that the improved model is excellent in performance on GasVid and IOD-Video data sets, the accuracy rates of the improved model reach 99.83% and 99.80% respectively, and the improved model is superior to a traditional model and an existing detection algorithm in multiple evaluation indexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural gas pipeline detection, and specifically relates to a natural gas pipeline leakage detection method based on ResNet50 and SMOTE. Background Art

[0002] Natural gas, a clean energy source primarily composed of hydrocarbons, is widely used in industries such as industry, transportation, residential heating, and power generation due to its high efficiency and environmental friendliness. Compared to coal and oil, natural gas emits only low levels of carbon dioxide and pollutants when burned, making it considered a key pillar of the low-carbon energy transition. Against the backdrop of the global push for carbon peak and carbon neutrality, natural gas, due to its unique role as a bridge energy source, is becoming a crucial component of the global energy landscape.

[0003] In recent years, China has become not only a major energy producer but also a major energy consumer. Natural gas supply in my country has been in short supply, and natural gas consumption continues to rise. According to the China Natural Gas Development Report (2024), China's natural gas consumption is projected to reach 394.5 billion cubic meters in 2023, a year-on-year increase of 7.6%. Natural gas's share of total primary energy consumption will increase to 8.5%. Urban gas consumption is growing fastest, accounting for 33%, while industrial fuel and power generation use are also experiencing rapid growth. This data reflects my country's enormous demand for clean energy. However, as domestic natural gas reserves and production are insufficient to meet rapidly growing consumption, my country needs to import large quantities of natural gas annually. Imports will reach 165.6 billion cubic meters in 2023, a year-on-year increase of 9.9%, making it a major net importer in the global natural gas market.

[0004] Currently, my country's natural gas imports primarily come from countries such as Russia, Turkmenistan, Australia, and Myanmar. This interconnection is achieved through long-distance pipelines such as the China-Russia East Line, the Central Asia Pipeline, and the China-Myanmar Pipeline. On December 2, 2024, the China-Russia East Line Pipeline was fully operational. This milestone project, totaling 5,111 kilometers, provides a solid energy guarantee for my country's eastern region. Improving natural gas reserves and supply capacity has become a key task in ensuring national energy security. According to industry statistics, my country's annual natural gas production target is approximately 230 billion cubic meters by 2025. Meanwhile, the construction of domestic natural gas pipeline networks and long-distance pipelines continues to accelerate, and the overall scale of the national oil and gas pipeline network is expected to grow to 210,000 kilometers. By 2023, the total mileage of long-distance natural gas pipelines nationwide reached 124,000 kilometers. While pipeline transportation offers significant advantages such as efficiency, safety, and low cost, the numerous regions, complex natural conditions, and long transportation distances involved in natural gas pipeline transportation pose significant safety risks that cannot be ignored. Summary of the Invention

[0005] The present invention provides a natural gas pipeline leakage detection method based on ResNet50 and SMOTE, which can effectively solve the above problems.

[0006] The present invention is achieved through the following technical solutions:

[0007] The present invention provides a natural gas pipeline leakage detection method based on ResNet50 and SMOTE, comprising:

[0008] Acquire infrared images of natural gas pipelines;

[0009] Preprocess the infrared image, including denoising the infrared image, specifically including:

[0010] Perform grayscale processing on the infrared image to obtain a grayscale image;

[0011] Perform filtering on the grayscale image;

[0012] The preprocessed image is detected and features are extracted from the preprocessed image, and whether the natural gas pipeline is leaking is determined based on the extracted features.

[0013] In some embodiments, the module used to extract features from the preprocessed image is a RepVGG module, and the RepVGG module uses a ReLU activation function.

[0014] In some embodiments, in the RepVGG module, the fusion process of the BN layer and the two-dimensional convolution is as follows:

[0015]

[0016] Among them, X i is the input of the i-th channel, μ i is the mean of the ith channel, σ i 2 is the variance of the ith channel, γ and β are learned through training, where ε is a constant to prevent the denominator from being zero, W is the weight of the ith channel, and b is the bias.

[0017] In some embodiments, in the RepVGG module, the branch fusion process is as follows:

[0018] C=C1+C2+C3,B=B1+B2+B3

[0019]

[0020] Among them, C i is the convolution on different branches after fusion, I and O are the input and output of the inference model respectively, and B iis the bias on different branches, C is the convolution after the convolution on each branch is added, and B is the bias after the addition of each branch.

[0021] In some embodiments, performing filtering processing on the grayscale image includes performing mean filtering processing on the grayscale image.

[0022] In some embodiments, performing filtering processing on the grayscale image includes performing median filtering processing on the grayscale image.

[0023] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0024] The natural gas pipeline leak detection method proposed in this paper, based on ResNet50 and SMOTE, improves the model's ability to extract, learn, and integrate natural gas leak features. Experimental results show that the improved model performs well on the GasVid and IOD-Video datasets, achieving accuracies of 99.83% and 99.80%, respectively, outperforming traditional models and existing detection algorithms across multiple evaluation metrics. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings in the embodiments will be briefly introduced below. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0026] Figure 1 A schematic diagram of the ResNet50 structure provided for some embodiments of the present invention;

[0027] Figure 2 Schematic diagram of the residual structure of the ResNet neural network model provided for some embodiments of the present invention.

[0028] Figure 3 A graph showing the natural gas pipeline leak detection results based on ResNet50 provided in some embodiments of the present invention. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments.

[0030] In the description of the present invention, it should be noted that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside", etc. used to indicate the orientation or position relationship are based on the orientation or position relationship shown in the accompanying drawings, or are the orientation or position relationship in which the product of the invention is usually placed when in use. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they should not be understood as limiting the present invention.

[0031] Furthermore, the use of terms such as "horizontal" and "vertical" in the description of the present invention does not necessarily imply that the component must be absolutely horizontal or suspended, but rather that it can be slightly tilted. For example, "horizontal" simply means that its direction is more horizontal than "vertical" and does not mean that the structure must be completely horizontal, but rather that it can be slightly tilted.

[0032] It should also be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention in specific contexts.

[0033] The terms "comprise," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or modules is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other steps or modules inherent to the process, method, product, or apparatus.

[0034] The basic structure of convolutional neural networks

[0035] Convolutional neural networks are mainly composed of convolutional layers, pooling layers and fully connected layers.

[0036] (1) Convolutional layer

[0037] The convolutional layer is the most important component of a convolutional neural network. It consists of multiple convolution kernels, which extract image features. First, the convolution kernel extracts features from the input image, and then uses these features as input to the next convolution layer. This process consumes a significant amount of computation time. The computational process can be expressed as Equation (2-1).

[0038]

[0039] Where X1(t) and X2(t) are convolution variables, and * represents the convolution operation. Each convolution kernel slides over the input array in a specific order, performing a convolution operation on the input array region to extract features of the target region.

[0040] (2) Pooling layer

[0041] The pooling layer, also known as the downsampling layer, primarily reduces high-dimensional data to low-dimensional data. Pooling reduces the size of feature maps, reducing the number of model parameters and thereby improving the model's generalization capabilities. Common pooling methods include maximum pooling and average pooling, with maximum pooling being the most frequently used and often performing better than average pooling. The difference between these two pooling methods is that maximum pooling uses the maximum value within a window as the pooling result, while average pooling uses the average value within a window as the pooling result.

[0042] (3) Fully connected layer

[0043] Fully connected layers (MLPs) are a type of multi-layer perceptron (MLP) model that flattens the high-dimensional features obtained by convolution and pooling operations into low-dimensional features and links the feature vector of the previous node with the feature vector of the next node. In convolutional neural networks, fully connected layers are equivalent to a classifier. Compared to the parameters of convolutional and pooling layers, fully connected layers have a large number of parameters and are computationally intensive, requiring more powerful computer hardware.

[0044] ResNet50 neural network

[0045] In order to improve the recognition performance of the model, traditional neural networks often capture more complex features by increasing the number of network layers. However, the continuous deepening of the network often leads to problems such as gradient explosion and gradient disappearance, which makes the accuracy of the neural network model continue to decline. ResNet effectively alleviates the gradient problem through its unique design, making the training of deep networks more stable and efficient. ResNet neural networks are usually composed of multiple residual blocks, each residual block is usually composed of two or three convolutional layers, and each convolutional layer is followed by batch normalization (Batch Normalization) and activation functions (such as ReLU). The residual module of ResNet skips the input through one or more layers and adds it to the output information, so that the information jumps and passes in the network, so that subsequent layers can learn the features of the output of the previous layers. This design helps to alleviate the problem of gradient disappearance, makes the network have deeper layers, and is also easier to train. Considering the lightweight of the model, this application uses the ResNet50 model as the research model, and its model structure is as follows Figure 1 shown.

[0046] The residual structures of ResNet neural network models of different depths are different. In ResNet18 and ResNet34, the residual structure consists of two sets of 3×3 convolutional layers, while in ResNet50 and above models, the second 3×3 convolutional layer is replaced by a more complex residual module. This module usually contains a series of residual connections and batch normalization layers, which helps to further enhance the expressive power of the model. This change enables the model to better solve the gradient disappearance and representation bottleneck problems as the depth increases, thereby improving the performance and stability of the model. Specifically, Figure 2 shown.

[0047] Transformers network model

[0048] The Transformer model uses a self-attention mechanism to establish correlations between elements in the input sequence, thereby replacing the traditional recurrent neural network and convolutional neural network structures. This structural change has achieved great success in natural language processing tasks and has later been proven to be equally effective in other tasks such as image processing.

[0049] The Transformer model primarily consists of an encoder and a decoder. Stacking multiple encoders and decoders forms a complete Transformer structure. Another core component of the Transformer model is the multi-head self-attention mechanism. In this self-attention mechanism, each input element interacts with all other elements, dynamically determining the weights of the relationships between different positions. This attention mechanism allows the model to establish long-range dependencies between different positions, helping to capture global information in a sequence. In the field of fine-grained image recognition, the input sequence and self-attention concepts of the Transformer have been borrowed and have achieved promising results.

[0050] Unbalanced data processing methods

[0051] Oversampling method

[0052] (1) SMOTE oversampling method

[0053] The SMOTE (Synthetic Minority Oversampling Technique) oversampling technique is a synthetic minority oversampling method proposed by Chawla et al. in 2002. Experiments have shown that this method performs well on most imbalanced datasets. The core idea of ​​the SMOTE oversampling method is as follows: For each minority sample x, the Euclidean distance to the remaining minority samples is calculated using formula (2-2). The k closest samples are obtained. A sampling factor N is set based on the sample imbalance ratio. Then, random linear interpolation is performed on the k adjacent samples of the minority sample using formula (2-3) to synthesize many new, non-repeating samples.

[0054]

[0055] X new =X i +rand(0,1)×(X j -X i ) (2-3)

[0056] Where: X i represents the minority class sample, X j Represents the remaining minority class samples, X new Synthesized minority class samples. The SMOTE oversampling method uses random linear interpolation to synthesize samples rather than simply copying the original samples. This makes oversampling of minority class samples more reasonable and overcomes the overfitting problem of random oversampling methods. However, it also has disadvantages. Due to its relatively simple synthesis method, the newly synthesized samples may contain noise data and sample overlap.

[0057] (2) Borderline-SMOTE oversampling method

[0058] Borderline-SMOTE is an oversampling method for imbalanced datasets. It balances the dataset by interpolating minority class samples to generate new synthetic samples. Borderline-SMOTE divides minority class samples into safe, danger, and noise classes. Here, more than half of the samples surrounding a safe sample are minority class samples, more than half of the samples surrounding a dangerous sample are majority class samples, and all samples surrounding a noise sample are majority class samples. Borderline-SMOTE primarily oversamples dangerous samples, generating new samples by randomly selecting minority class samples from the K-nearest neighbors and interpolating them.

[0059] Undersampling method

[0060] (1) Tomek Link undersampling method

[0061] The Tomek Link undersampling method was proposed by Tomek in 1976. Its principle is very simple. The core idea is as follows: Assume that X i and X j Represent the minority class sample points and the majority class sample points respectively, and use formula (2-4) to calculate the Euclidean distance d(X i , X j ), if there is no such sample point X in all samples k , satisfying d(X k , X i ) is less than d(X i , X j ) form a Tomek Link pair. This means that the two sample points are very close, both on the boundary, or one of the sample points is noise. Tomek Link pairs are typically processed by deleting the majority class sample point or both sample points. Before the Tomek Link undersampling method, there is no clear dividing line between the majority and minority classes. With this method, by deleting the majority class samples in the Tomek Link pair, samples that are closer together are assigned to the same class, resulting in a clear dividing line between the majority and minority classes.

[0062]

[0063] CatBoost algorithm

[0064] The CatBoost algorithm is a machine learning algorithm developed by Yandex and is also an improved implementation of the GBDT algorithm. Its unique advantages are that it automatically and efficiently processes categorical variables, generates cross-features using a greedy strategy, and solves problems such as gradient deviation and prediction offset. A simple way to deal with categorical features is the Greedy Target-based Statistics (abbreviated as Greedy TBS) method, which uses the average value of the target variable label corresponding to the category as the replacement value. However, the Greedy TBS method has the following defects. First, if the number of samples in a category is too small, replacing it with the mean will generate noise; second, the feature contains more information, and directly replacing it with the label average will lose a lot of information. Adding a smoothing prior on this basis, the formula is:

[0065]

[0066] Where a is the weight and p is the prior. While this is a solution, it can lead to data leakage. To address this, CatBoost has designed an efficient encoding method for ordered target variable statistics. This method first randomly shuffles the samples, then for each sample, only calculates the numerical encoding for the samples that precede it. To mitigate prediction drift, CatBoost has designed an ordered boosting algorithm that similarly randomly permutes the dataset and then trains the model based solely on the preceding data. This algorithm effectively reduces gradient bias. Furthermore, CatBoost uses symmetric trees as base learners, which can speed up scoring.

[0067] GBDT algorithm

[0068] GBDT belongs to the Boosting family of algorithms. Its core idea is to iteratively train decision trees based on weak learners and combine them to form a powerful strong learner. Unlike traditional decision tree algorithms, GBDT is an additive model that continuously adds new decision trees to the existing model to gradually reduce the model's prediction error.

[0069] The boosted tree uses the decision tree as the base classifier and adopts an additive model, which can be expressed as:

[0070]

[0071] Where: T(X; Θ m ) is the mth decision tree, Θ m is the parameter of the decision tree, and M is the total number of trees.

[0072] GBDT integrates multiple weak learners and, by continuously optimizing the loss function, effectively captures complex relationships in the data, resulting in high prediction accuracy. Due to its additive model and stepwise iteration approach, GBDT is less prone to overfitting. This is especially true when combined with appropriate regularization methods, such as limiting the tree depth and controlling the number of leaf nodes, which enhances the model's generalization capabilities.

[0073] XGBoost algorithm

[0074] XGBoost is a framework based on gradient boosting decision trees. Its core is to iteratively train a series of decision trees to build a powerful ensemble model. It continuously adds new decision trees to the existing model to gradually reduce the model's prediction error. The objective function in XGBoost consists of two parts: the loss function y and the regularization term k.

[0075]

[0076] Among them: the loss function is used to measure the model prediction value y i and the true value Y i The difference between them, the regularization term, is used to control the complexity of the model and prevent overfitting.

[0077] During training, XGBoost calculates the first and second derivatives of the loss function with respect to the predicted value and uses this information to determine the growth direction and splitting point for each tree. To improve computational efficiency, XGBoost also uses approximate algorithms for large-scale data, such as the quantile approximation algorithm, which quickly finds the approximate optimal splitting point by bucketing the eigenvalues.

[0078] Natural gas pipeline leak detection based on ResNet50

[0079] Preprocessing of infrared images of natural gas pipeline leakage

[0080] Image quality is one of the important factors affecting the effectiveness of deep learning training. Whether the image quality can support training and whether the detected object is clear will directly affect the training effect and detection accuracy of subsequent models.

[0081] When capturing images with an infrared camera, the quality of the image captured is affected by a variety of factors. Imaging distance is one of these factors. When the imaging distance is too far, the target object's pixel ratio in the image becomes smaller, and many details may not be clearly captured, resulting in a blurry image and difficulty extracting effective features. Light intensity is also important. Both excessive and insufficient light can negatively impact infrared images. Excessive light may cause the image to appear overexposed and lose some details, while insufficient light can cause the image to appear darker, increase noise, and reduce contrast. Shooting orientation is also important. Different shooting orientations may obscure key areas of the target object, thus affecting the model's ability to learn the target's complete features. Furthermore, electronic noise in the imaging device can negatively impact image quality. For example, shot noise and white noise are introduced during signal amplification, conversion, and transmission. This noise manifests itself in various forms, such as graininess and streaking, seriously interfering with image clarity and legibility.

[0082] Given the impact of these factors on infrared image quality, preprocessing of the resulting images is necessary. Image filtering algorithms can effectively improve image quality, enhance image clarity and contrast, and remove noise and interference, thereby providing a better data foundation for subsequent deep learning training, improving model training effectiveness and detection accuracy, and ensuring that the entire deep learning task can be completed more efficiently and accurately.

[0083] This application performs denoising and smoothing processing on the salt and pepper noise and white noise present in infrared images. During imaging, due to reasons such as the small temperature difference between the target and the environment, infrared images often have low contrast and cannot clearly distinguish the contours of the target objects in the image. Therefore, this application performs enhancement processing on the problem of low contrast in infrared images so that these processed images can meet the needs of subsequent contour feature extraction and target detection.

[0084] Improved ResNet50 network model overall framework

[0085] To achieve high-accuracy natural gas leak detection, a deep learning-based natural gas pipeline leak detection method uses ResNet50 as the baseline model. This model is a relatively deep neural network that captures multi-layered features, converges quickly, and performs well on medium-sized datasets. In conventional detection scenarios, the difference between gas and background in natural gas infrared images is subtle, as the temperature difference between the background and gas in the image is small. Furthermore, gas is easily affected by the natural environment, resulting in complex motion and shape. Therefore, high requirements are placed on the model's feature extraction, feature fusion, and spatial modeling capabilities. To improve the accuracy of natural gas leak detection, this application improves ResNet50. To learn richer features and enhance the model's recognition performance, the original 7×7 convolution and max pooling layers are replaced with an improved RepVGG module. A newly constructed ConWit module is introduced in the middle of the network to integrate local and global features at a higher level. Finally, the fully connected layers are replaced with convolutional layers, which are translation-invariant, reduce model complexity, and optimize detection performance.

[0086] Improved RepVGG module

[0087] Due to light scattering caused by atmospheric turbulence and disturbances, methane can be difficult to distinguish from the background in methane infrared images. The original ResNet50 feature extraction network has a large receptive field, resulting in the loss of some geometric information. Furthermore, detailed feature information is lost after passing through the max pooling layer. To enhance the model's feature extraction capabilities and enrich geometric information, this application will restructure the feature extraction network.

[0088] The RepVGG network uses a multi-branch structure during training to enhance the model's representational capabilities; during inference, it adopts a single-path structure, increasing parallelism for faster inference and more flexible models. This application improves the RepVGGBlock by replacing the ReLU activation function with the SiLU activation function. The ReLU activation function is computationally efficient, but when the value of the input leakage region is less than zero, the ReLU gradient will remain zero, resulting in information loss. In contrast, the SiLU activation function has a non-zero gradient across the entire real range, providing complete gradient information, thereby improving the model's representational capabilities and performance. The training model draws on the residual link concept of ResNet, forming a multi-branch structure consisting of 3×3 convolutions, 1×1 convolutions, batch normalization layers, and residual links, accelerating model convergence. The inference model transforms each branch in the training model into a fusion of 3×3 convolutional layers and batch normalization layers. This results in each branch being fused into a new 3×3 convolution, and finally, multiple branches being fused into a single 3×3 convolution. This inference model simplifies the computational process, improves parallel processing capabilities, and thus improves inference speed and efficiency. The following describes the fusion calculation process of the BN layer and the two-dimensional convolution. The BN of the i-th channel of the input feature map, y iCalculation is as shown in formula (3-1):

[0089]

[0090] Where: X i is the input of the i-th channel, μ i is the mean of the ith channel, σ i 2 is the variance of the i-th channel, and γ and β are learned through training. ε is a very small constant to prevent the denominator from being zero. From this formula, we can derive formula (3-2), which can then be simplified to obtain formula (3-3) for the weight W and bias b of the i-th channel.

[0091] For the fusion of 1×1 convolution and BN layer, the 1×1 convolution needs to be converted into 3×3 convolution and then fused with the BN layer. For only BN layer, a 3×3 convolution needs to be constructed to perform identity mapping on the input and then fused with the BN layer. Finally, the multi-branch is converted into a single 3×3 convolution C. The branch fusion calculation formula is as follows:

[0092] C=C1+C2+C3,B=B1+B2+B3 (3-4)

[0093]

[0094] Where: C i is the convolution on different branches after fusion, I and O are the input and output of the inference model respectively, and B i is the bias on different branches, C is the convolution after the convolution on each branch is added, and B is the bias after the addition of each branch.

[0095] Improved Conwit module

[0096] The model feature extraction network has been enhanced. Although it enriches the geometric information, due to its small receptive field, the global feature extraction of the gas will be weakened. It pays too much attention to the local and ignores the global, which will affect the model detection accuracy. In order to more effectively utilize the feature information extracted by the feature extraction network, this application introduces a multi-head self-attention mechanism to construct the ConWit module; Transformer can generate more advanced semantic features and integrate local information with global information to improve model performance; the feature map output by the latter convolution has integrated local and global features, and after further screening and enhancement, it has the expression of a variety of feature combinations, so that the output features are richer and the generalization ability of the model is improved. The Transformer in the ConWit module consists of a normalization layer (LayerNorm), a multi-head self-attention mechanism (Multi-HeadAttention), a feedforward neural network (MLP), and a dropout layer (Dropout).

[0097] The multi-head self-attention mechanism plays an important role in the Transformer Encoder. It is derived from the scaled dot-product attention mechanism. The formula of the dot-product attention mechanism is as follows:

[0098]

[0099] Where: Q, K, V are query vector, key vector, and value vector respectively, which are obtained by transforming the input features through three learnable parameter matrices, and V is the dimension of K. The correlation matrix is ​​obtained by multiplying Q by the transpose of K. To prevent the value after the dot product from being too large, the correlation matrix is ​​divided by the square root of d K , and then multiplied by V after softmax normalization.

[0100] Multi-head self-attention was proposed based on this. Its essence is to process input features from different positions with different weights to obtain the final result. The formula of the multi-head self-attention mechanism is as follows:

[0101] MultiHead(Q,K,V)=Concat(head1,...,head h ) (3-7)

[0102] wherehead i =Attention(QW i Q ,KW i K ,VW i V ) (3-8)

[0103] Where: W Q , W K and W V is the parameter matrix for implementing projection, i represents the i-th self-attention head, and Concat is the concatenation operation. The multi-head attention mechanism can obtain feature information from different locations, which is more advantageous than a single attention mechanism.

[0104] Improved output layer

[0105] The output layer of the ResNet50 model consists of an adaptive average pooling layer (Avg) and fully connected layers (FC). The number of parameters in the fully connected layer increases exponentially with the increase in the input feature dimension, while the convolutional layer can share weights when processing images, so the number of parameters is relatively small. The fully connected layer performs a large amount of computation on the feature map, while the computational complexity of the convolutional layer is mainly affected by the size and number of convolution kernels, which is relatively small. Therefore, the fully connected layer is replaced with a convolutional layer, which consists of a 1×1 convolution, a ReLU activation function, a 1×1 convolution, and a global average pooling layer.

[0106] Experiment and result analysis

[0107] Improved ResNet50 network leakage detection results

[0108] This application uses the Python-based OpenCV library to segment videos. 49,890 infrared images were selected from the GasVid dataset, including 29,934 training images, 9,978 validation images, and 9,978 test images. 30,000 images were selected from the IOD-Video dataset, with 18,000 training images, 6,000 validation images, and 6,000 test images, respectively.

[0109] Analysis of experimental results on GasVid

[0110] The experimental results on the GasVid dataset are as follows Figure 3 As shown in the figure, the accuracy of the training set stabilized at 99.76% after 40 iterations, with a loss value of approximately 0.001. The accuracy of the validation set gradually stabilized at approximately 99.83% after 30 iterations, with a loss value of approximately 0.0005. The model training effect was good, and the loss curve eventually converged, indicating that the model did not suffer from underfitting or overfitting.

[0111] In order to further verify the effectiveness of the different modules proposed in this application in improving the model's natural gas leak detection, this application conducted an ablation experiment on the model on the GasVid dataset. The experimental results are shown in Table 3.1.

[0112] Table 3.1 Ablation experiment results

[0113]

[0114] In the ablation experiment, ResNet50 was used as the baseline model for Group 1. Table 2 compares the experiments of Group 1 and Group 2. Based on the model of Group 1, Group 2 replaced the fully connected layer of the output layer with a convolutional layer, reducing the number of parameters from the original 23.5M to 27.7M, and slightly improving the model performance. The results show that the convolutional layer not only reduces the parameters of the model but also improves the performance of the model. Comparing Group 2 and Group 3, Group 3 replaced the feature extraction network with the RepVGG module based on the model of Group 2. Its detection accuracy, precision, recall rate, and F1 score increased by 1.11%, 1.08%, 0.99%, and 1.03%, respectively. The data show that the RepVGG module improves the model's feature extraction ability, thereby improving the model's accuracy in detecting natural gas leaks. Comparing Group 3 with Group 4, Group 4 added the ConWit module to Group 3's model, improving detection accuracy, precision, recall, and F1 score by 0.15%, 0.14%, 0.26%, and 0.20%, respectively. Experiments show that the ConWit module enhances the model's ability to fuse local and global features and the positional connections between features, helping the model more accurately detect natural gas leak areas.

[0115] Comparative experiment

[0116] To ensure the effectiveness of the model, this application uses four evaluation indicators: Accuracy (A), Precision (Precision), Recall (Recall), and F1 score (F1-score) to evaluate the performance of the model on the dataset.

[0117]

[0118]

[0119] Among them: TP and TN refer to the number of positive and negative samples with correct reasoning, respectively, and FP and FN refer to the number of positive and negative samples with incorrect reasoning.

[0120] The proposed model is compared with the traditional models MobileNet-V2, DenseNet121, RegNetX-200MF, and ResNet50 on the IOD-Video dataset, and the results are shown in Table 3.2.

[0121] Table 3.2 Comparison results with other network models

[0122]

[0123] As shown in Table 3.2, the proposed model achieved an accuracy of 99.80% on the IOD-Video dataset, outperforming other models in accuracy, precision, recall, and F1-score. This result is primarily due to two factors: firstly, the proposed RepVGG module provides more discriminative features to improve gas leak detection performance; secondly, the proposed ConWit module uses a multi-head self-attention mechanism to effectively integrate local and global information, enhancing the model's detection capabilities.

[0124] Table 3.3 shows the comparison results of different models (MobileNet-V2, DenseNet121, RegNetX-200MF, ResNet50) and the improved ResNet50 model in this application on the GasVid dataset. Experiments were conducted on different datasets to evaluate the generalization ability and reliability of the models, verify the universality of the experimental results, and ultimately determine the optimal model.

[0125] Table 3.3 Comparison results with other network models

[0126]

[0127] It can be seen from Table 3.3 that the evaluation indicators of the models proposed in this application are higher than those of other models. The accuracy of MobileNetV2 is 97.53%. This model performs well in efficient computing due to the lightweight design of depthwise separable convolution, but it performs poorly in modeling global features due to the limitation of receptive field. The accuracy of DenseNet121 is 98.35%. It achieves efficient feature reuse by densely connecting all layers, but as the number of layers increases, the dimension of the feature map also increases, making the network difficult to optimize. The accuracy of RegNetX-200MF is 98.47%. RegNetX-200MF adopts a design that gradually increases the width and depth of the network, making the model highly efficient in computing and with a small number of model parameters. Its limitations are reflected in fine feature extraction and global context modeling.

[0128] Natural gas leak detection based on SMOTE non-equilibrium data processing and ensemble learning

[0129] Algorithm Principle

[0130] The DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm is a density-based clustering algorithm that categorizes sample points in a dataset into core points, boundary points, and noise points based on the density of surrounding sample points. DBSCAN does not require a predefined number of clusters and can discover clusters of any shape, making it highly adaptable. The algorithm flow is as follows:

[0131] Step 1: Determine an unvisited sample point.

[0132] Step 2: Determine the neighborhood radius r and the minimum point count threshold MinPts.

[0133] Step 3: Traverse the sample data set. For each unvisited point P in the data set:

[0134] (1) Mark point P as visited.

[0135] (2) Calculate the neighborhood point set N(P) of point P. According to the neighborhood radius, find all points in the data set whose distance to point P is less than or equal to r to form N(P).

[0136] (3) Determine the type of point P: If N(P) is greater than or equal to MinPts, then point P is a core point, otherwise it is a noise point.

[0137] (4) Expanded clustering: For each unvisited point Q in the neighborhood point set N(P), if Q is a core point, all points in the neighborhood are recursively added to the cluster and the cluster is continued to be expanded until all reachable points are visited.

[0138] Step 4: Repeat the above steps until all points in the dataset have been visited.

[0139] The improved SMOTE algorithm combines the concepts of DBSCAN and triangle incenters. It first uses the DBSCAN algorithm to perform cluster analysis on minority class samples, identifying noise points, boundary points, and safe points. It then calculates the number of new samples required based on the sample proportions of each cluster. When generating new samples, it selects the nearest boundary point of a core point in the cluster and two randomly selected points from other boundary points in the same cluster to form a triangle. The incenter of this triangle is calculated, and a point on the line connecting the two randomly selected boundary points and the triangle incenter is selected as the new synthesized sample, thereby balancing the dataset classes.

[0140] The DBSCAN algorithm parameter, radius r, determines the neighborhood range of a point. If r is too small, many points may be misclassified as noise points, preventing effective clustering. If r is too large, clusters may over-merge, losing local features of the data. Setting it to 0.5 provides a good balance between the number of points in the neighborhood and clustering accuracy. Regarding the minimum point count threshold, MinPts, if MinPts is set too low, some boundary points may be misclassified as core points, resulting in unstable clustering results. Setting MinPts too high may result in too few core points, with many points being classified as noise. Setting it to 5 ensures that core points are representative while avoiding excessive screening of useful points.

[0141] Randomly select a point in the sample and randomly select two sample points in the point family to form a triangle. Calculate the incenter a of the triangle and generate new sample points z on the two angle bisectors.

[0142] The improved SMOTE algorithm focuses on sample points at the data boundaries that are easily misclassified. When generating new samples, it limits the range of sample generation, significantly reducing the likelihood that the newly synthesized samples will fall into the majority class sample space. Newly generated samples cluster toward the triangle's centroid, preventing further marginalization of the sample distribution while keeping them relatively close to the cluster center, resulting in a more balanced sample distribution. Furthermore, new samples are not directly copied from the original data points but are generated through a specific computational method. This feature effectively reduces the risk of overfitting the model, thereby improving its generalization and stability.

[0143] Algorithm steps

[0144] The improved SMOTE algorithm steps are as follows:

[0145] Step 1: Divide the unbalanced dataset H into the minority class H min and the majority set H maj , the corresponding sample size is Y min and Y maj , let the ratio of negative class to positive class after balance be r. Use DBSCAN algorithm to analyze Y min Perform cluster analysis and identify noise points, boundary points, and safe points according to the definition of the DBSCAN algorithm. The total number of noise points is recorded as Ynoise.

[0146] Step 2: Y min The remaining boundary points and safe points are divided into k clusters using the DBSCAN algorithm, m i Cluster C i The cluster centers of (i=1, 2, ...k).

[0147] Step 3: Calculate the number of new samples that need to be generated.

[0148]

[0149] Step 4: For the i-th cluster (i=1, 2, ...k): calculate the number of new samples to be generated for the cluster based on the proportion of the number of samples.

[0150]

[0151] Step 5: In each cluster, select the boundary point closest to the cluster center, and then randomly select two points X from other boundary points in the same cluster. S and X t , forming a triangle. By calculating the intersection of the perpendicular bisectors of the three sides of the triangle, we can get the incenter X of the triangle. B .

[0152]

[0153] Among them: the lengths of the three sides of a triangle in n-dimensional space are a, b, c (j=1, 2, ..., n).

[0154] Step 6: Generate new minority class samples randomly on the line connecting the incenter of the triangle and the boundary points, and add them to the original dataset. The calculation formula is as follows:

[0155]

[0156] Y new2 =X t +rand(0,1)×(X B -X t ) (3-18)

[0157] Step 7: Repeat steps 5 and 6 until the number of samples meets the requirement.

[0158] Preliminary experimental verification

[0159] To further verify the effectiveness of the improved algorithm in processing imbalanced data, seven binary classification imbalanced data sets were selected from the KEEL data platform for experimental comparison of different data sets.

[0160] Before the experiment, each dataset was preprocessed. Missing values ​​in the dataset were filled by using the mean of numerical features and the most frequently occurring category for categorical features. The dataset was divided into a training set and a test set in an 8:2 ratio.

[0161] The training set was balanced using the improved algorithm, the original SMOTE algorithm, and the Borderline-SMOTE method. A random forest classification model was used to perform classification predictions on the balanced dataset. The model was trained using the same random seed to ensure reproducibility, and default parameters were used for training.

[0162] In the evaluation of imbalanced data processing, F1, AUC, and G-mean evaluation indicators are selected for comparison.

[0163]

[0164] Among them: rank insi Represents the serial number of the i-th sample, M is the number of positive samples, N is the number of negative samples, and the sum is the sum of the serial numbers of positive samples; TP and TN refer to the number of positive and negative samples with correct inference, respectively, and FP and FN refer to the number of positive and negative samples with incorrect inference, respectively.

[0165] A random forest model constructed using the improved SMOTE algorithm achieved the best evaluation results across various datasets in most cases. Whether it's the overall model performance reflected by the F1 score, the ability to distinguish between positive and negative samples reflected by the AUC score, or the overall performance on imbalanced data measured by the G-mean score, the improved SMOTE algorithm significantly outperformed both the traditional SMOTE algorithm and the Borderline-SMOTE method. This series of experimental results clearly demonstrates the improved SMOTE algorithm's superiority when processing imbalanced datasets, preliminarily confirming its feasibility and efficiency.

[0166] Building a classification model

[0167] The basic operation of S-fold cross-validation is to divide the original dataset into S non-overlapping data subsets, use S-1 subsets as training sets, and the remaining subset as validation set, resulting in S models. Then, based on the specified evaluation criteria, the model that achieves the best mean across S training runs is selected. Cross-validation can effectively avoid underfitting and overfitting, improve the model's generalization ability, and ultimately produce a more convincing model. This application uses 5-fold cross-validation to adjust the parameters.

[0168] Different models require different parameters to be adjusted, and the detection results are also different. For a single model, the detection effect is also related to the parameter selection. When establishing a model for detection, if you want to obtain a model with higher prediction accuracy, you need to constantly try and adjust it. This is mainly done by selecting the most appropriate parameters through cross-validation on the processed training set to establish the model. If you only adjust the parameters on the test set, it will cause data leakage and easily produce overfitting. The established model will only show good results on the test set, but the generalization ability is not strong, and the prediction of other data is inaccurate. Grid search is a brute force optimization method. Its approach is to divide the parameters of the model into small intervals, and use cross-validation to find the optimal solution for the parameters through a traversal method. Grid search can ensure the best possible global optimization and avoid major errors. When establishing a natural gas leak detection model, this application uses a grid search method to adjust the parameters of each model, compares different models through the evaluation criteria of the classification model, and finally obtains a model with better prediction effect through model fusion based on the voting method.

[0169] In the processing of natural gas pipeline leakage pressure data, different models have different requirements for parameters, and the prediction effects will also be different. In order to obtain a more accurate prediction model, this application uses cross-validation to screen suitable parameters in the processed training set to build a model. If the parameters are adjusted only for the test set, it is easy to cause data leakage, resulting in overfitting of the model and reduced prediction accuracy for new data. Grid search can maximize global optimization and reduce errors by traversing the parameter interval and combining cross-validation to find the optimal solution. This application uses the grid search method to adjust the parameters of each model, and fuses the models through the voting method to improve prediction performance.

[0170] Random Forest Algorithm

[0171] Random forests use CART trees as their base learners and are composed of multiple CART trees. They randomly sample both samples and features when training the decision trees. The RandomForestClassifier function in the Python sklearn library is used to implement the classification function of the random forest model. The main adjustment parameters include n_estimators (the number of decision trees), max_depth (the maximum depth of the decision tree), and min_samples_split (the minimum number of samples required for internal node splitting).

[0172] First, adjust the n_estimators parameter. The number of decision trees significantly affects model performance. Too few may lead to underfitting, while too many may cause overfitting. When the number of decision trees is small, the model's AUC value increases rapidly with increasing numbers, indicating that more trees improve the model's classification ability. However, when the number of decision trees exceeds 100, further increasing the number of trees results in very little change in the AUC value, indicating that the performance improvement has reached saturation. Considering that a large number of decision trees may lead to overfitting, we initially set the search range for the number of decision trees to 60-100 with a step size of 10, using a grid search with 5-fold cross-validation. This method initially determined the optimal number of decision trees to be 90. To further optimize this result, we refined the search range to 80-100 with a step size of 1. After further searching, we determined the optimal number of decision trees to be 92.

[0173] Next, adjust the max_depth parameter. The maximum depth of the decision tree also has a significant impact on model performance. A decision tree with too shallow a depth may not fully learn the complex features of the data, resulting in underfitting; a decision tree with too deep a depth may overfit the training data, reducing the model's generalization ability. Experiments have found that when the maximum depth of the decision tree is greater than 30, the model's AUC value remains almost constant, indicating that the decision tree is deep enough at this point. Further increasing the depth will have little effect on improving model performance. The search range for max_depth is set to 10-30.

[0174] GBDT algorithm

[0175] The GBDT algorithm uses the negative gradient of the loss function as an approximation of the residual, thereby continuously reducing the overall model loss. When processing natural gas pipeline leakage pressure data, the GBDT classification algorithm is implemented by calling the GradientBoostingClassifier function in the Python sklearn library. Its parameters are divided into boosting framework parameters and base learner parameters.

[0176] Tuning parameters for the boosting framework include n_estimators, learning_rate (the learning rate), which controls the weight reduction coefficient of each base learner and coordinates the number of iterations, and subsample (the subsampling ratio), which represents the degree of subsampling of training examples. The search range for n_estimators is set to 70-110; a smaller learning_rate enhances model stability and is set based on experience; subsample is generally set between 0.6 and 0.8.

[0177] The tuning parameters for the base learner are max_depth, min_samples_split, and max_features, which represents the maximum number of features used for splitting. Options include None, log2, and sqrt. The search range for max_depth is set to 4-7. For max_features, assuming there are M features in the dataset, None means all features are considered for splitting, log2 means a maximum of log2M features are considered, and sqrt means a maximum of M features are considered.

[0178] XGBoost algorithm

[0179] The XGBoost algorithm is an efficient machine learning algorithm under the gradient boosting framework. It adds a regularization term and performs a second-order Taylor expansion on the loss function. It has a fast training speed and can effectively prevent model overfitting. In the processing of natural gas pipeline leakage pressure data, the XGBClassifer function in the sklearn package is called for modeling. In the general parameters, booster represents the base learner, and the default parameter gbtree is used, that is, the tree model. In the Booster parameters, n_estimators represents the number of trees, and the search range is set to 70-110; the subsample parameter represents the random sampling ratio of each tree; the colsample_bytree parameter is used to control the ratio of randomly selected features, and the search range of both is set to 0.6-0.9; max_depth, the search range is set to 4-7; gamma specifies the minimum loss function drop value required for node splitting, and the setting range is 0-0.5; learning_rate improves model stability by reducing the weight each time, the search range is 0.01-0.15, and the initial value is set to 0.1. Comparison of the improved SMOTE algorithm and the classic sampling method

[0180] The training set was imbalanced using the SMOTE and Borderline-SMOTE methods, and then four models, random forest, GBDT, XGBoost, and CatBoost, were established respectively. The parameter ranges and methods of adjustment of each model were consistent in the comparison.

[0181] Across four machine learning models, the improved SMOTE algorithm achieved significant improvements across all metrics compared to the traditional SMOTE algorithm. Accuracy increased by an average of 11.0%, G-mean by 9.1%, F1 by 10.4%, and AUC by 8.2%. G-mean saw the greatest improvement in the CatBoost model, increasing by 10.47%. The GBDT algorithm saw the largest improvement in F1, increasing by 17.95%. The XGBoost algorithm saw the largest improvement in AUC, increasing by 10.11%.

[0182] Compared with the Borderline-SMOTE method, the improved SMOTE algorithm demonstrates significant advantages. Accuracy increased by an average of 7.9%, G-mean increased by 7.0%, F1 score increased by 7.1%, and AUC increased by 4.2%. Across all models, the improved SMOTE algorithm achieved significant improvements in all three metrics, demonstrating its significant advantages in improving model performance.

[0183] In summary, experiments on a natural gas pipeline leakage pressure dataset show that the improved SMOTE algorithm significantly outperforms the original SMOTE and Borderline-SMOTE algorithms in terms of accuracy, G-mean, F1 value, and AUC. This clearly demonstrates the outstanding effectiveness of the improved SMOTE algorithm. Using different classifiers, the algorithm significantly improves evaluation metrics, strongly demonstrating its superior performance in processing unbalanced natural gas pipeline leakage pressure data and effectively improving the model's performance in natural gas pipeline leak detection.

[0184] The above description is only a preferred embodiment of the present invention and does not constitute any limitation to the present invention. Any simple modification or equivalent change made to the above embodiment based on the technical essence of the present invention shall fall within the scope of protection of the present invention.

Claims

1. A natural gas pipeline leakage detection method based on ResNet50 and SMOTE is characterized by: include: Acquire infrared images of natural gas pipelines; Preprocess the infrared image, including denoising the infrared image, specifically including: Perform grayscale processing on the infrared image to obtain a grayscale image; Perform filtering on the grayscale image; The preprocessed image is detected and features are extracted from the preprocessed image, and whether the natural gas pipeline is leaking is determined based on the extracted features.

2. The natural gas pipeline leakage detection method based on ResNet50 and SMOTE according to claim 1 is characterized in that: The module used for feature extraction of the preprocessed image is the RepVGG module, and the RepVGG module uses the ReLU activation function.

3. The natural gas pipeline leakage detection method based on ResNet50 and SMOTE according to claim 1 is characterized in that: In the RepVGG module, the fusion process of the BN layer and the two-dimensional convolution is as follows: Among them, X i is the input of the i-th channel, μ i is the mean of the ith channel, σ i 2 is the variance of the ith channel, γ and β are learned through training, where ε is a constant to prevent the denominator from being zero, W is the weight of the ith channel, and b is the bias.

4. The natural gas pipeline leakage detection method based on ResNet50 and SMOTE according to claim 1 is characterized in that: In the RepVGG module, the branch fusion process is as follows: C=C1+C2+C3,B=B1+B2+B3 Among them, C i is the convolution on different branches after fusion, I and O are the input and output of the inference model respectively, and B i is the bias on different branches, C is the convolution after the convolution on each branch is added, and B is the bias after the addition of each branch.

5. The natural gas pipeline leakage detection method based on ResNet50 and SMOTE according to claim 1 is characterized in that: The filtering process for the grayscale image includes performing mean filtering process on the grayscale image.

6. The natural gas pipeline leakage detection method based on ResNet50 and SMOTE according to claim 1 is characterized in that: The filtering process performed on the grayscale image includes performing a median filtering process on the grayscale image.